Назад
Company hidden
3 часа назад

Senior Software Engineer, ML Infrastructure Platform (AI)

193 930 - 291 150$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Software Engineer, ML Infrastructure Platform (AI): Building and operating distributed GPU training, large-scale data pipelines, and agentic-first ML workflows for autonomous driving with an accent on orchestration, observability, reliability, and cost management. Focus on designing reproducible data-to-training-to-evaluation systems, diagnosing distributed training bottlenecks, and improving operational maturity across critical autonomy pipelines.

Location: Mountain View, California, United States

Salary: $193,930–$291,150 base pay annually, plus performance bonus, equity, and benefits.

Company

hirify.global develops Level 4 autonomous driving technology and a universal autonomy platform for robotaxis, logistics fleets, personal vehicles, and other mobility applications.

What you will do

  • Contribute to training infrastructure across multiple accelerator generations, clusters, scheduling systems, and orchestration layers.
  • Design and operate large-scale batch and streaming data pipelines, including storage layouts and high-throughput data generation.
  • Build agentic-first ML workflows connecting data generation, model training, and evaluation.
  • Develop introspectable, reproducible workflows that autonomy teams can run and extend.
  • Own reliability for critical training and release pipelines through instrumentation, alerting, on-call practices, and incident response.

Requirements

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field, plus 3+ years of relevant experience.
  • Strong proficiency in Python and comfort with C++, Go, or a similar systems language.
  • Hands-on experience operating production infrastructure on Kubernetes.
  • Solid distributed-systems fundamentals, including performance, failure-mode, and reliability analysis.
  • Ownership mindset and experience improving operational maturity through monitoring, alerting, and runbooks.
  • Work location: Mountain View, California, United States.

Nice to have

  • Strong working knowledge of GCP.
  • Experience building large-scale data-generation pipelines and Kubernetes-native orchestration for ML workloads.
  • Knowledge of GPU and distributed-training internals, including NCCL and collective communication.
  • Experience with GPU and training observability tools, infrastructure cost reduction, and reliability improvement.

Culture & Benefits

  • Annual performance bonus and equity eligibility.
  • Competitive benefits package.
  • Inclusive workplace focused on diversity, psychological safety, and equal opportunity.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →