6 дней назад
Senior / Staff ML Ops Engineer (AI)
184 000 - 272 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior / Staff ML Ops Engineer (AI/Autonomous Transportation): Building reliable Kubernetes-based training infrastructure, developer tooling, and ML platforms for autonomous trucks and robotaxis with an accent on GPU scheduling, distributed training, data pipelines, and experiment lifecycle management. Focus on reducing iteration time, improving platform observability and reliability, and creating self-service infrastructure that researchers and engineers adopt.
Location: Hybrid; remote work available in the United States and Canada. Office locations include Dallas, Phoenix, Pittsburgh, San Francisco, and Toronto.
Salary: $184,000–$272,000 USD per year for US locations, plus equity incentives, an annual performance bonus, and benefits.
Company
is a technology startup building Physical AI and autonomous transportation systems for commercial autonomous trucks and robotaxis.
What you will do
- Build and evolve Kubernetes-based ML training infrastructure, including GPU scheduling, autoscaling, distributed jobs, capacity planning, operators, and workflow engines.
- Develop developer-facing CLIs, SDKs, job submission flows, templates, and self-service paths for research and engineering teams.
- Improve training iteration speed through local workflows, fast failure detection, platform measurement, and reduced time to first training run.
- Strengthen dataset versioning, sharding, high-throughput multimodal sensor-data loading, experiment tracking, model registries, and lineage.
- Build CI/CD, observability, access controls, data-handling guardrails, and cost governance across the ML platform.
- Drive adoption through prototyping with users, documentation, onboarding, office hours, and reliable production tooling.
Requirements
- 5+ years of software or infrastructure engineering experience with tools or platforms used by other engineers and production ML or data-intensive systems.
- Hands-on Kubernetes expertise, including GPU scheduling, autoscaling, Helm or equivalent, networking, and cluster debugging under load.
- Excellent Python skills and experience designing developer-friendly APIs and CLIs.
- Practical AWS experience with object storage, IAM, GPU compute, networking, cost management, and infrastructure as code such as Terraform or Pulumi.
- Experience with distributed PyTorch training, experiment tracking, model registries, containers, CI/CD, modern build systems, and large monorepos.
- Ability to influence without authority, collaborate across teams, write clearly, operate autonomously, and work on ML, robotics, or autonomous-systems infrastructure.
Nice to have
- Internal developer platform, research platform, or DevEx experience with demonstrated tool adoption.
- Large-scale distributed GPU training, NCCL, high-performance cluster networking, and collective communication tuning.
- High-throughput LiDAR or camera-data pipelines using Parquet or WebDataset.
- Experience with Argo Workflows, Ray, Flyte, Kubeflow, Slurm, Bazel, simulation infrastructure, or batch evaluation pipelines.
- Background in security- and IP-sensitive environments or open-source ML infrastructure and developer tools.
Culture & Benefits
- Hybrid workplace with flexible hours and work-from-home support.
- Medical, dental, and vision coverage for full-time employees.
- Unlimited vacation.
- Competitive compensation, equity awards, and an annual performance bonus.
- Catered meals when working in the office, snacks, drinks, and regular onsite, offsite, and virtual team events.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
8 дней назад
Staff AI Engineer
205 000 - 307 000$
6 дней назад
Software Development Engineer - ML Ops (US Federal)
163 800 - 245 800$
12 дней назад
Senior Site Reliability Engineer (AI)
185 500 - 232 000$
8 дней назад
Sr. ML Engineer (AI)
123 400 - 191 100$
8 дней назад
Senior Machine Learning Systems Engineer (CAD)
154 000 - 193 000CAD
13 дней назад
AI/ML Platform Engineer
147 050 - 230 850$