Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Forward Deployed Engineers (AI) (AI inference and infrastructure): Building and operating production AI inference, post-training, evaluation, and deployment systems for mission-critical customer workloads with an accent on distributed systems, GPU optimization, and customer-facing technical ownership. Focus on designing benchmarks, debugging complex production failures, improving model serving and post-training workflows, and shipping infrastructure and product changes across multiple accounts.
Location: San Francisco, United States; hybrid
Salary: $200,000–$400,000 per year plus equity
Company
Baseten provides inference infrastructure and developer tooling that help AI companies bring advanced models into production.
What you will do
- Own technical outcomes for customer workloads, from problem framing and specifications through proof of concept and production deployment.
- Design evaluations and benchmarks, optimize inference, improve models through post-training, and refine evaluation workflows.
- Respond to mission-critical incidents, diagnose root causes, implement fixes, and coordinate with owning teams until resolution.
- Build tooling, automation, deployment infrastructure, recipes, and reference implementations to make engagements more efficient and self-serve.
- Translate customer needs into product roadmap improvements and ship fixes and features in the Baseten codebase.
- Manage multiple customer engagements while aligning customers and internal stakeholders on status and risk.
Requirements
- 1–2 years of software engineering experience shipping and maintaining code in large production systems.
- Experience debugging complex production issues using logs, metrics, and traces.
- Ability to own ambiguous technical problems, make decisions under uncertainty, and involve system owners when needed.
- Strong communication skills with customer engineers and leadership.
- Interest in AI inference, training, and the infrastructure that supports them.
- Willingness to respond to customers outside regular working hours and participate in an on-call rotation.
Nice to have
- Expertise in storage, networking, InfiniBand, RoCE, or other infrastructure domains.
- Experience with Kubernetes, Slurm, or Ray for GPU workloads.
- Knowledge of LLM architectures and inference engines such as vLLM, TensorRT-LLM, or SGLang.
- GPU workload profiling and optimization experience.
- Experience with SFT, reinforcement learning, PyTorch, JAX, or deep learning.
- Operational experience with on-call, incident response, and distributed systems debugging.
Culture & Benefits
- Competitive compensation with meaningful equity.
- Medical, dental, and vision insurance fully covered for employees and dependents.
- Flexible PTO and company-wide Winter Break from Christmas Eve through New Year's Day.
- Paid parental leave and a fertility and family-building stipend.
- Company-facilitated 401(k).
- Exposure to ML startups and mission-critical AI workloads.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
1 день назад
Forward Deployed Engineer (Post-Sales) (AI)
230 000 - 300 000$
6 часов назад
AI Engineer
159 000 - 195 000$
6 часов назад
Forward Deployed Engineer (AI/ML)
200 000 - 400 000$
3 часа назад
Technical Staff
200 000 - 350 000$
2 часа назад
AI Engineer (Fintech)
180 000 - 260 000$
4 часа назад
Forward Deployed Engineer (AI)
100 000 - 300 000$