6 дней назад
Staff ML Systems Engineer, Distributed Systems
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff ML Systems Engineer, Distributed Systems (ML infrastructure/robotics): Architecting and building distributed infrastructure for large-scale machine learning workflows across data processing, model training, evaluation, and post-processing with an accent on scalability, reliability, and performance. Focus on distributed execution, CPU/GPU optimization, resource allocation, fault tolerance, observability, and productionizing research workflows for real-world robotics deployments.
Location: Seattle, WA or Irvine, CA; on-site
Company
builds risk-aware, reliable, field-ready embodied AI systems for real robots, sensors, and robotics deployments.
What you will do
- Design and build scalable distributed machine learning pipelines for data processing, model training, evaluation, and post-processing.
- Architect distributed execution systems covering parallelization, workload scheduling, resource allocation, and fault tolerance.
- Develop reusable abstractions, frameworks, and libraries for distributed pipeline development.
- Optimize CPU and GPU workloads, including data partitioning, memory utilization, serialization, throughput, and compute efficiency.
- Partner with ML engineers, data engineers, and infrastructure teams to productionize research workflows and support large-scale model development.
- Improve engineering standards, observability, debugging, monitoring, and operational tooling for distributed systems.
Requirements
- 5+ years of experience building distributed systems, backend infrastructure, machine learning platforms, or large-scale data processing systems.
- Strong Python skills, including concurrency, performance optimization, and systems development.
- Experience with distributed computing frameworks such as Ray, Spark, Dask, Flink, or similar technologies.
- Experience designing and scaling data pipelines or machine learning workflows.
- Strong system design expertise with a focus on scalability, reliability, and performance optimization.
- Ability to work on-site in Seattle, WA or Irvine, CA.
Nice to have
- Experience building infrastructure for ML training and inference systems.
- Familiarity with PyTorch or TensorFlow.
- Experience with multi-node or multi-GPU training, including DDP, FSDP, or DeepSpeed.
- Experience operating Kubernetes-based infrastructure and large-scale cloud systems.
- Experience with distributed debugging, observability, and workflow orchestration platforms.
Culture & Benefits
- Work on embodied AI systems that are tested on real hardware and improved through field deployments.
- Competitive compensation based on background, geographic location, knowledge, skills, and experience.
- Comprehensive benefits and equity participation.
- Opportunity to contribute to advances in AI and robotics.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Data Infrastructure Engineer (AI)
170 000 - 360 000$
6 дней назад
ML Data Engineer
6 дней назад
Member of Technical Staff – ML Systems & Inference (AI)
250 000 - 350 000$
5 дней назад
(US) Principal ML System Engineer (AI)
195 000 - 217 000$
20 часов назад
Senior Data Scientist (AI/ML)
124 000 - 329 200$
3 дня назад
Senior Staff ML Engineer (Generative AI)
150 000 - 300 000$