Эта вакансия в архиве

Посмотреть похожие вакансии ↓
Company hidden
обновлено 16 минут назад

AI Systems Performance Specialist

130 000 - 180 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US

Описание вакансии

Текст:
/
TL;DR
AI Systems Performance Specialist (AI Infrastructure, GPU Optimization, and LLM Inference): Optimizing AI training and inference workloads for throughput, latency, scalability, reliability, and cost efficiency with an accent on GPU utilization, distributed computing, and production AI systems. Focus on profiling compute and memory bottlenecks, designing advanced inference optimizations, building benchmarking and regression frameworks, and leading performance initiatives across enterprise-scale AI platforms.

Location: 100% remote within the Continental United States

Salary: $130,000–$180,000 annually

Company

hirify.global is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Optimize AI training and inference pipelines for throughput, low latency, scalability, reliability, and infrastructure efficiency.
  • Improve GPU utilization, memory management, kernel execution, and multi-GPU performance across production workloads.
  • Implement quantization, pruning, mixed precision, batching, caching, speculative decoding, and model parallelism.
  • Profile applications across compute, memory, networking, and storage to identify and resolve bottlenecks.
  • Optimize distributed training and inference using frameworks such as NCCL, DeepSpeed, PyTorch Distributed, Ray, and MPI.
  • Build benchmarking frameworks, performance dashboards, monitoring solutions, and regression testing pipelines while mentoring engineering teams.

Requirements

  • 10+ years of professional experience in performance engineering, AI infrastructure, machine learning systems, HPC, or distributed computing.
  • Bachelor’s or Master’s degree in computer science, computer engineering, electrical engineering, artificial intelligence, or a related discipline.
  • Expert-level Python and C++ programming skills with extensive CUDA and GPU-accelerated AI workload optimization experience.
  • Strong knowledge of LLMs, deep learning frameworks, model serving, production AI inference, distributed systems, networking, and storage optimization.
  • Experience with NVIDIA Nsight Systems, Nsight Compute, PyTorch Profiler, TensorBoard, or similar tools, and with AWS, Azure, or GCP.
  • Applicants must be U.S. citizens, Green Card holders, EAD holders, or H-1B transfer candidates; new H-1B sponsorship is not available.

Nice to have

  • Production-scale LLM inference and large foundation model serving experience.
  • Experience with vLLM, TensorRT-LLM, DeepSpeed, Triton Inference Server, CUTLASS, or FasterTransformer.
  • Knowledge of model compression, KV cache optimization, speculative decoding, and advanced inference techniques.
  • FinOps experience, AI systems research, open-source contributions, patents, or technical publications.
  • Familiarity with AMD ROCm, Intel oneAPI, or custom AI accelerator technologies.

Culture & Benefits

  • Full-time direct W-2 employment.
  • Opportunity to work on enterprise-scale cloud, AI, data, and infrastructure solutions.
  • Collaboration with AI researchers, ML engineers, platform engineers, and infrastructure teams.
  • Career growth opportunities within an established technology organization.