4 часа назад
Member of Technical Staff, Inference & Serving (AI)
200 000 - 350 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff, Inference & Serving (AI): Building and optimizing high-performance serving systems for low-latency diffusion LLM inference with an accent on distributed orchestration, model endpoint traffic management, and production reliability. Focus on GPU-aware performance optimization, autoscaling, canary deployments, observability, and translating new model architectures and quantization techniques into production serving systems.
Location: Bay Area, United States; in-office
Salary: $200,000–$350,000 annual base salary, plus equity and benefits
Company
develops diffusion-based large language models, including Mercury, for faster and more efficient AI inference.
What you will do
- Build and optimize high-performance model serving systems for low-latency diffusion LLM inference.
- Extend Kubernetes, Ray, and SLURM orchestration for distributed inference, evaluation, and large-batch serving.
- Implement load balancing, autoscaling, traffic routing, model versioning, canary deployments, and zero-downtime rollouts.
- Develop monitoring, alerting, and observability tooling to support SLAs and incident response.
- Collaborate with ML researchers to productionize new architectures, quantization techniques, and batching strategies.
Requirements
- BS, MS, PhD, or equivalent experience in Computer Science, Engineering, or a related field.
- Knowledge of SGLang, vLLM, Triton Inference Server, or TensorRT-LLM.
- Systems-level understanding of PyTorch and TensorFlow.
- Familiarity with high-performance computing and GPU programming, including CUDA.
- Experience with Docker, Kubernetes, CI/CD pipelines, and ML systems performance optimization and profiling.
Nice to have
- Experience serving large-scale language models with tens of billions of parameters or more.
- Distributed systems and cloud experience with AWS, GCP, or Azure.
- Experience with Kubeflow, Airflow, quantization, distillation, speculative decoding, continuous batching, checkpointing, or resource scheduling.
Culture & Benefits
- Work with AI researchers and engineers who pioneered diffusion models and related technologies.
- Competitive salary, equity, flexible vacation, and paid time off.
- Health, dental, and vision insurance, plus a 401(k) match.
- Catered breakfast, lunch, and dinner, with commuter subsidies.
- Collaborative and inclusive culture.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 часа назад
Member of Technical Staff — Inference-Core Engine (AI)
200 000 - 400 000$
4 часа назад
Member of Technical Staff — Inference-Kernel, Compiler & Communication (AI)
200 000 - 400 000$
4 часа назад
Member of Technical Staff — Inference-Multi-Hardware (AI Infrastructure)
200 000 - 400 000$
4 часа назад
Kernel Engineer (AI)
150 000 - 350 000$
4 часа назад
Member of Technical Staff (Applied AI)
150 000 - 350 000$
5 часов назад
Member of Technical Staff (Research Infrastructure)
200 000 - 400 000$