3 часа назад
Software Engineer (AI)
230 000 - 390 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineer (AI): Building and operating inference systems for self-hosted models and third-party providers with an accent on distributed infrastructure, low latency, reliability, and cost efficiency. Focus on designing routing and serving architecture, managing GPU capacity and quotas, and optimizing large-scale production inference.
Location: San Francisco, CA, United States; work arrangement: on-site
Salary: $230,000–$390,000 annually, plus equity
Company
develops customer-facing AI agents for major brands and operates primarily in person from San Francisco, with offices across North America, Europe, and Asia.
What you will do
- Design ’s inference architecture across self-hosted models and third-party inference providers.
- Build routing, failover, capacity-management, quota, serving, and proxy systems for large-scale production workloads.
- Operate self-hosted inference on GPU infrastructure, including containers, inference engines, and compute capacity.
- Optimize latency, throughput, reliability, and cost using techniques such as speculative decoding and serving-engine improvements.
- Partner with frontier labs, inference providers, Applied Research, Models, and Agent Runtime teams.
- Contribute to infrastructure supporting post-training and the broader model lifecycle.
Requirements
- Strong systems thinking and distributed-systems fundamentals.
- Experience designing, building, and operating large-scale production systems.
- Strong judgment regarding tradeoffs between latency, reliability, capacity, and cost.
- Experience owning complex infrastructure from architecture through production operation.
- Interest in applying systems expertise to AI infrastructure and learning rapidly as the technology evolves.
- On-site work in San Francisco, CA is required.
Nice to have
- Experience with ML infrastructure, MLOps, or production inference systems.
- Experience serving LLMs or other large models at scale.
- Experience operating self-hosted inference and GPU infrastructure.
- Familiarity with inference frameworks such as vLLM or SGLang.
- Experience with post-training infrastructure or inference-performance optimization.
Culture & Benefits
- Values include trust, customer obsession, craftsmanship, intensity, and family.
- Unlimited paid time off, medical, dental, and vision benefits for employees and families.
- Life insurance, disability benefits, parental leave, and fertility and family-building support.
- Retirement benefits depend on the country of employment.
- Lunch, snacks, coffee, a discretionary benefit stipend, and free alphorn lessons.
- Eligible full-time employees may participate in equity plans.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Nscale
5 дней назад
Principal AI Product Engineer (AI)
290 000 - 520 000$
3 дня назад
Forward Deployed Engineer (AI)
200 000 - 400 000$
Snowflake
4 дня назад
Senior AI Engineer
156 000 - 224 200$
4 дня назад
Principal Machine Learning Engineer (AI)
163 200 - 264 000$
Baseten
3 часа назад
Distributed Systems Engineer (AI Inference)
180 000 - 360 000$
6 дней назад
Principal Software Engineer (AI)
261 500 - 353 500$