Назад
Company hidden
4 часа назад

Member of Technical Staff, Inference & Serving (AI)

200 000 - 350 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff, Inference & Serving (AI): Building and optimizing high-performance serving systems for low-latency diffusion LLM inference with an accent on distributed orchestration, model endpoint traffic management, and production reliability. Focus on GPU-aware performance optimization, autoscaling, canary deployments, observability, and translating new model architectures and quantization techniques into production serving systems.

Location: Bay Area, United States; in-office

Salary: $200,000–$350,000 annual base salary, plus equity and benefits

Company

hirify.global develops diffusion-based large language models, including Mercury, for faster and more efficient AI inference.

What you will do

  • Build and optimize high-performance model serving systems for low-latency diffusion LLM inference.
  • Extend Kubernetes, Ray, and SLURM orchestration for distributed inference, evaluation, and large-batch serving.
  • Implement load balancing, autoscaling, traffic routing, model versioning, canary deployments, and zero-downtime rollouts.
  • Develop monitoring, alerting, and observability tooling to support SLAs and incident response.
  • Collaborate with ML researchers to productionize new architectures, quantization techniques, and batching strategies.

Requirements

  • BS, MS, PhD, or equivalent experience in Computer Science, Engineering, or a related field.
  • Knowledge of SGLang, vLLM, Triton Inference Server, or TensorRT-LLM.
  • Systems-level understanding of PyTorch and TensorFlow.
  • Familiarity with high-performance computing and GPU programming, including CUDA.
  • Experience with Docker, Kubernetes, CI/CD pipelines, and ML systems performance optimization and profiling.

Nice to have

  • Experience serving large-scale language models with tens of billions of parameters or more.
  • Distributed systems and cloud experience with AWS, GCP, or Azure.
  • Experience with Kubeflow, Airflow, quantization, distillation, speculative decoding, continuous batching, checkpointing, or resource scheduling.

Culture & Benefits

  • Work with AI researchers and engineers who pioneered diffusion models and related technologies.
  • Competitive salary, equity, flexible vacation, and paid time off.
  • Health, dental, and vision insurance, plus a 401(k) match.
  • Catered breakfast, lunch, and dinner, with commuter subsidies.
  • Collaborative and inclusive culture.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →