9 часов назад
Generalist Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Generalist Engineer (AI) (vLLM/inference infrastructure): Building and optimizing the full AI inference stack, from GPU kernels and model execution to distributed systems and cloud orchestration, with an accent on performance, scale, and autonomous end-to-end delivery. Focus on designing low-level CUDA or equivalent kernels, serving models across thousands of accelerators, and building reliable Kubernetes-based infrastructure for global AI inference.
Location: Fully remote, worldwide; timezone-flexible with regular overlap with Pacific Time for critical syncs.
Company
develops and advances vLLM as an AI inference engine, focusing on making model inference faster and more cost-efficient.
What you will do
- Optimize LLM and diffusion model serving within the vLLM inference runtime.
- Develop CUDA, Triton, TileLang, Pallas, or equivalent kernels for diverse accelerator architectures.
- Build distributed systems that serve models across thousands of accelerators with minimal latency.
- Develop cloud orchestration, cluster management, deployment automation, and production monitoring infrastructure.
- Work across the vLLM stack, including GPU kernels, distributed systems, model architectures, and ML infrastructure.
Requirements
- Bachelor’s degree or equivalent experience in computer science, engineering, or a related field.
- Demonstrated autonomy and ability to drive complex projects from vague problem statements to shipped code.
- Strong asynchronous communication skills and ability to collaborate across time zones.
- Strong record of delivering high-impact work in complex technical environments.
- Deep expertise in systems programming, GPU or accelerator programming, distributed systems, or ML infrastructure.
- Strong proficiency in at least two areas including GPU architecture and kernels, Rust/Go/C++ distributed systems, Python with PyTorch and LLM inference, Kubernetes, or transformer-based model serving.
Nice to have
- Contributions to vLLM or major open-source ML or systems projects.
- Experience with NVIDIA, AMD, TPU, or Intel accelerators.
- Knowledge of quantization, ML kernel optimization, or compiler technologies.
- Experience improving reliability and performance at scale.
- Technical writing or impactful ML infrastructure side projects.
Culture & Benefits
- Asynchronous, autonomous work with regular Pacific Time overlap for critical coordination.
- Collaboration with the San Francisco headquarters in a globally remote environment.
- Competitive compensation comprising salary and equity, adjusted to local market conditions.
- Location-appropriate benefits, including health coverage where applicable.
- Visa sponsorship available on a case-by-case basis.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
14 часов назад
Senior Backend Engineer (AI Infrastructure)
4 дня назад
Senior AI Systems Engineer (Inference)
Anthropic
2 дня назад
Staff + Senior Software Engineer, Inference (AI)
Baseten
1 день назад
Forward Deployed Engineers (AI)
200 000 - 400 000$
13 часов назад
Applied AI Engineer (AI)
2 дня назад
Principal Software Engineer (AI)
160 200 - 425 000$