8 часов назад
Platform Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Platform Engineer (AI): Building and maintaining observable, secure, and scalable infrastructure for reliable and reproducible ML workloads with an accent on high-performance compute, deployment automation, and release safety. Focus on designing resilient CI/CD and deployment systems, hardening environments through load testing and chaos engineering, and improving reliability through incident response and root-cause analysis.
Location: San Francisco, United States; on-site, with an in-person team.
Requirements
- Experience with high-performance compute environments, including ML clusters, GPU farms, and hyperscaler clouds such as AWS or GCP.
- Infrastructure as code experience with tools such as Terraform or Ansible.
- Experience with containers such as Docker or Apptainer and scheduling systems such as Kubernetes or Slurm.
- Experience with deployment strategies at scale, disaster recovery planning, change management, fault tolerance, and reliable experimental environments.
- Knowledge of compliance and audit standards for deployment and system security.
- Experience with load testing, fault injection, and chaos engineering.
Nice to have
- Experience supporting ML/AI infrastructure, including GPU management and workload optimization.
- Exposure to backend development for ML model serving with tools such as vLLM, Ray, SGLang, or Triton.
- Experience with software release engineering for ML/AI systems.
What you will do
- Build and improve observability systems for monitoring, logging, and alerting.
- Manage infrastructure as a service and CI/CD in partnership with engineering teams.
- Design resilient build and deployment systems across research and production environments.
- Implement secure, auditable release processes with rollback support.
- Collaborate with ML, DevOps, and infrastructure teams to improve reliability and performance.
- Lead incident response, root-cause analysis, and postmortems focused on prevention.
Culture & Benefits
- Research and engineering work follows methodical, step-by-step approaches to ambitious goals.
- New ideas are encouraged, with willingness to make significant bets on promising concepts.
- Medical, dental, vision, and FSA plans.
- Competitive compensation, 401(k) plan, unlimited PTO, and company holidays.
- Relocation and immigration support may be available on a case-by-case basis.
- In-office snacks and meals in a collaborative San Francisco environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
11 часов назад
Infrastructure Engineer (AI)
160 000 - 245 000$
11 часов назад
Infrastructure Engineer (AI)
150 000 - 300 000CAD
7 часов назад
Infrastructure Engineer (AI)
15 часов назад
DevOps Engineer (AI)
85 000 - 180 000$
12 часов назад
Platform Engineer (AI)
6 дней назад
Platform Engineer (AI)
170 000 - 300 000$