Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Forward Deployed Engineers (AI) (AI inference and infrastructure): Building and operating production AI inference, post-training, evaluation, and deployment systems for mission-critical customer workloads with an accent on distributed systems, GPU optimization, and customer-facing technical ownership. Focus on designing benchmarks, debugging complex production failures, improving model serving and post-training workflows, and shipping infrastructure and product changes across multiple accounts.
Location: Hybrid in San Francisco or New York, United States
Salary: $200,000–$400,000 per year, plus equity
Company
Baseten provides inference infrastructure, applied AI research, and developer tooling that help AI companies deploy and operate cutting-edge models in production.
What you will do
- Own the technical outcomes of multiple customer accounts, acting as the primary technical owner for workloads deployed on Baseten.
- Turn ambiguous customer objectives into specifications, proofs of concept, success criteria, and production deployments.
- Design evaluations and benchmarks, optimize inference, improve models through post-training, and close quality or performance gaps.
- Respond to mission-critical incidents, perform triage, own fixes, and remain accountable through resolution.
- Build evaluation and deployment automation, internal tooling, recipes, and reference implementations.
- Influence the product roadmap and ship fixes and features in the Baseten codebase while coordinating customers and internal stakeholders.
Requirements
- 1–2 years of software engineering experience shipping and maintaining code in large production systems.
- Experience debugging complex production issues using logs, metrics, and traces.
- Ability to own ambiguous technical problems, make decisions under uncertainty, and involve the appropriate system owners.
- Strong communication skills with customer engineers, technical leaders, and internal stakeholders.
- Interest in AI inference, training, and the infrastructure that supports them.
- Willingness to support customers outside regular working hours and participate in an on-call rotation.
Nice to have
- Depth in infrastructure domains such as storage, networking, InfiniBand, or RoCE.
- Experience operating Kubernetes, Slurm, or Ray for GPU workloads.
- Knowledge of LLM architectures and inference engines such as vLLM, TensorRT-LLM, or SGLang.
- Experience profiling and optimizing GPU workloads, using post-training techniques such as SFT or RL, or working with PyTorch or JAX.
- Operational experience with on-call, incident response, and distributed systems debugging.
Culture & Benefits
- Competitive compensation with meaningful equity.
- U.S. employees and dependents receive full medical, dental, and vision coverage.
- Flexible PTO and a company-wide winter break.
- Paid parental leave and a fertility and family-building stipend.
- U.S. employees receive access to a company-facilitated 401(k).
- Exposure to ML startups and mission-critical AI workloads.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
11 дней назад
Applied ML Engineer (AI)
200 000 - 235 000$
Anthropic
8 дней назад
Performance Engineer (AI)
280 000 - 850 000$
11 дней назад
Software Engineer (AI)
210 000 - 265 000$
Anthropic
8 дней назад
Research Engineer Scientist (AI)
350 000 - 500 000$
Databricks
14 дней назад
AI Engineer – Forward Deployed Engineering (AI FDE)
152 900 - 210 155$
13 дней назад
Forward Deployed AI Engineer (Multiple Levels)
110 000 - 240 000$