Назад
Company hidden
6 дней назад

Staff ML Systems Engineer, Distributed Systems

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff ML Systems Engineer, Distributed Systems (ML infrastructure/robotics): Architecting and building distributed infrastructure for large-scale machine learning workflows across data processing, model training, evaluation, and post-processing with an accent on scalability, reliability, and performance. Focus on distributed execution, CPU/GPU optimization, resource allocation, fault tolerance, observability, and productionizing research workflows for real-world robotics deployments.

Location: Seattle, WA or Irvine, CA; on-site

Company

hirify.global builds risk-aware, reliable, field-ready embodied AI systems for real robots, sensors, and robotics deployments.

What you will do

  • Design and build scalable distributed machine learning pipelines for data processing, model training, evaluation, and post-processing.
  • Architect distributed execution systems covering parallelization, workload scheduling, resource allocation, and fault tolerance.
  • Develop reusable abstractions, frameworks, and libraries for distributed pipeline development.
  • Optimize CPU and GPU workloads, including data partitioning, memory utilization, serialization, throughput, and compute efficiency.
  • Partner with ML engineers, data engineers, and infrastructure teams to productionize research workflows and support large-scale model development.
  • Improve engineering standards, observability, debugging, monitoring, and operational tooling for distributed systems.

Requirements

  • 5+ years of experience building distributed systems, backend infrastructure, machine learning platforms, or large-scale data processing systems.
  • Strong Python skills, including concurrency, performance optimization, and systems development.
  • Experience with distributed computing frameworks such as Ray, Spark, Dask, Flink, or similar technologies.
  • Experience designing and scaling data pipelines or machine learning workflows.
  • Strong system design expertise with a focus on scalability, reliability, and performance optimization.
  • Ability to work on-site in Seattle, WA or Irvine, CA.

Nice to have

  • Experience building infrastructure for ML training and inference systems.
  • Familiarity with PyTorch or TensorFlow.
  • Experience with multi-node or multi-GPU training, including DDP, FSDP, or DeepSpeed.
  • Experience operating Kubernetes-based infrastructure and large-scale cloud systems.
  • Experience with distributed debugging, observability, and workflow orchestration platforms.

Culture & Benefits

  • Work on embodied AI systems that are tested on real hardware and improved through field deployments.
  • Competitive compensation based on background, geographic location, knowledge, skills, and experience.
  • Comprehensive benefits and equity participation.
  • Opportunity to contribute to advances in AI and robotics.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →