56 минут назад
Distributed Systems Engineer (AI Inference)
180 000 - 360 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Distributed Systems Engineer (AI Inference): Building and operating the distributed runtime and Model APIs that power large-scale LLM inference with an accent on routing, autoscaling, Kubernetes orchestration, and advanced model-serving capabilities. Focus on designing low-latency production systems, debugging GPU and infrastructure reliability issues, and balancing performance, scalability, operational simplicity, and developer experience.
Location: Hybrid in San Francisco, Montreal, New York, Seattle, or Toronto
Salary: $180K–$360K annually, plus equity
Company
Baseten provides infrastructure, applied AI research, and developer tooling for deploying and operating production AI models and large-scale inference.
What you will do
- Build infrastructure and orchestration systems for distributed LLM inference, including routing, autoscaling, scheduling, and runtime management.
- Design, build, and operate Model APIs supporting structured outputs, tool and function calling, and multimodal serving.
- Implement API versioning, validation, usage metering, quotas, and authentication.
- Develop observability through metrics, traces, logs, benchmarks, testing, and release automation.
- Debug and harden production systems across Kubernetes, distributed runtimes, networking, and GPU workloads.
- Own projects end to end and collaborate with inference performance engineers to deliver optimizations to customers.
Requirements
- Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 3+ years of experience building and operating distributed systems, backend infrastructure, or large-scale APIs.
- Experience owning low-latency, reliable backend services with rate limiting, authentication, quotas, metering, and migrations.
- Knowledge of profiling, tracing, capacity planning, SLO management, and production reliability practices.
- Ability to debug performance and reliability issues across application, runtime, and infrastructure layers.
- Strong developer-experience, written communication, and cross-functional collaboration skills.
Nice to have
- Experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or Dynamo.
- Deep Kubernetes experience, including operators and custom resources, plus service meshes or API gateways.
- Experience with distributed scheduling, autoscaling, service orchestration, or production GPU workloads.
- Background in developer-facing infrastructure, APIs, open-source infrastructure, or ML systems.
- Familiarity with observability tooling, CI/CD systems, or release automation.
Culture & Benefits
- Competitive compensation with meaningful equity.
- Flexible PTO and a company-wide Winter Break.
- Paid parental leave and a fertility and family-building stipend.
- U.S.-only: medical, dental, and vision insurance fully covered for employees and dependents.
- U.S.-only: company-facilitated 401(k).
- Exposure to ML startups and opportunities for learning and networking.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Forward Deployed Engineer (AI)
200 000 - 400 000$
Anthropic
6 дней назад
Staff + Sr. Software Engineer, Cloud Inference (AI)
320 000 - 485 000$
Anthropic
6 дней назад
Staff + Sr. Software Engineer, Cloud Inference Launch Engineering (AI)
320 000 - 485 000$
10 часов назад
Software Engineer, AI Systems
115 000 - 159 000$
4 часа назад
API Product Engineer (AI)
6 дней назад
System Software Engineer (AI)
120 500 - 243 000$