3 часа назад
Senior Production Engineer (AI Infrastructure)
170 000 - 205 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Production Engineer (AI Infrastructure): Building and operating reliable managed AI services for large language model workloads with an accent on distributed systems, scalability, observability, and cost-efficient infrastructure. Focus on designing fault-tolerant AI platforms, improving SLI/SLO performance, and solving complex reliability challenges across training and inference clusters.
Location: On-site in San Francisco or Sunnyvale, California, United States
Salary: $170,000–$205,000 per year, plus Restricted Stock Units
Company
builds vertically integrated energy and AI infrastructure, operating systems from power generation through cloud services to support large-scale AI workloads.
What you will do
- Design and operate reliable managed AI services for large language model serving and scaling.
- Build automation and reliability tooling for distributed AI pipelines and inference services.
- Define, measure, and improve SLIs and SLOs across AI workloads.
- Collaborate with AI, platform, and infrastructure teams to optimize large-scale training and inference clusters.
- Automate observability and develop telemetry and performance-tuning strategies for latency-sensitive services.
- Investigate and resolve reliability issues using telemetry, logs, and profiling, while contributing to next-generation AI-focused distributed systems.
Requirements
- Production-grade software engineering experience beyond scripting or Bash.
- Experience designing and implementing distributed systems.
- Hands-on experience with large language models or AI/ML infrastructure.
- SRE experience, including defining SLIs/SLOs, monitoring, observability, reliability improvements, fault tolerance, and automated testing.
- Proficiency in Python, Go, Java, or C++.
- Experience with Kubernetes or container orchestration platforms, plus strong collaboration and communication skills.
Nice to have
- Experience scaling LLM inference or training workloads.
Culture & Benefits
- Health, vision, and dental insurance options for employees and dependents.
- HSA contributions, 401(k) matching up to 4% of salary, and paid parental leave.
- Paid time off, holidays, life insurance, disability coverage, and commuter benefits.
- Tuition reimbursement, cell phone reimbursement, Teladoc, Calm, and MetLife Legal subscriptions.
- Restricted Stock Units in a fast-growing technology company.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 часов назад
Senior Site Reliability Engineer (AI Infrastructure)
215 000 - 275 000$
2 часа назад
Software Engineer, DevOps (Robotics)
115 000 - 170 000$
1 час назад
Developer Experience Engineer (AI/HPC)
150 000 - 275 000$
5 часов назад
Senior Site Reliability Engineer (SRE)
170 000 - 196 000$
3 часа назад
Software Engineer (AI Infrastructure)
1 час назад