Назад
Company hidden
3 часа назад

Software Engineer, ML Infrastructure Platform (AI)

160 360 - 240 540$
Тип работы
fulltime
Грейд
junior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer, ML Infrastructure Platform (AI): Building and operating infrastructure for distributed GPU training, data pipelines, and agentic-first ML workflows supporting autonomous driving with an accent on orchestration, observability, reliability, and cost management. Focus on designing reproducible data-to-training-to-evaluation systems, improving Kubernetes-based production infrastructure, and diagnosing performance and failure modes across distributed systems.

Location: Mountain View, California, United States

Salary: $160,360–$240,540 base pay per year, plus annual performance bonus, equity, and benefits.

Company

hirify.global is a physical AI company developing Level 4 autonomous driving technology and a universal autonomy platform for vehicles and mobility services.

What you will do

  • Contribute to training infrastructure across multiple generations of accelerators, including multi-cluster scheduling and orchestration.
  • Design and operate large-scale batch and streaming data pipelines, storage layouts, and high-throughput data generation systems.
  • Build agentic-first ML workflows covering data generation, training, and evaluation with reproducible and introspectable pipelines.
  • Own reliability for critical training and release pipelines through instrumentation, alerting, runbooks, on-call practices, and incident response.

Requirements

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a closely related field, plus at least 1 year of relevant experience.
  • Strong proficiency in Python and comfort with C++, Go, or a similar systems language.
  • Hands-on experience running production infrastructure on Kubernetes.
  • Solid distributed-systems fundamentals, including performance, failure-mode, and reliability analysis.
  • Ownership mindset and willingness to improve technical and operational standards through monitoring, alerting, and operational maturity practices.

Nice to have

  • Strong working knowledge of GCP.
  • Experience building large-scale data-generation pipelines and Kubernetes-native orchestration for ML workloads.
  • Depth in GPU and distributed-training internals, including NCCL and collective communication.
  • Experience with GPU and training observability tools, infrastructure cost reduction, and reliability improvements.

Culture & Benefits

  • Annual performance bonus, equity, and a competitive benefits package.
  • Commitment to diversity, inclusion, psychological safety, and equal employment opportunity.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →