Назад
Company hidden
3 дня назад

Senior AI Infrastructure Engineer (Autonomous Vehicles)

180 000 - 240 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior AI Infrastructure Engineer (Autonomous Vehicles) (ML Infrastructure/MLOps): Building and scaling the high-performance AI platform that supports autonomous driving models with an accent on distributed training, GPU orchestration, model deployment, and agentic automation. Focus on optimizing multi-node H100/A100 clusters, developing self-healing infrastructure, and connecting research workflows with reliable production systems.

Location: Onsite five days a week at the Santa Clara, California office

Salary: $180,000–$240,000 per year

Company

hirify.global develops Level 4 autonomous transportation technology for short-haul, B2B middle-mile logistics and operates commercially deployed autonomous trucks.

What you will do

  • Design, build, and scale the AI infrastructure supporting perception, planning, world-model, and VLA research workloads.
  • Enable distributed multi-node and multi-GPU training using PyTorch Distributed, FSDP, DeepSpeed, and Ray Train.
  • Optimize GPU clusters, networking, and inference using NVIDIA hardware, NCCL, InfiniBand or RoCE v2, TensorRT, ONNX Runtime, and Triton.
  • Develop agentic infrastructure automation for cluster monitoring, failure triage, CI/CD, resource optimization, and data curation.
  • Implement MLOps workflows for experiment tracking, model lifecycle management, A/B testing, shadow deployments, and rollbacks.
  • Build cloud-native data pipelines and observability systems with Airflow, Kafka, Spark, Prometheus, Grafana, OpenTelemetry, and ELK.

Requirements

  • 5+ years of experience in ML infrastructure, MLOps, or DevOps for high-scale compute environments.
  • Deep expertise in multi-GPU training, high-performance networking, Kubernetes, Terraform, and Helm.
  • Experience with MLflow, Argo Workflows, Docker, and GPU-native orchestration.
  • Proven experience building or supporting agentic workflows for infrastructure or data automation.
  • Proficiency in Apache Airflow, Kafka, Spark, GitOps automation, Python, and Bash.
  • Ability to work onsite five days per week in Santa Clara, California.

Nice to have

  • Experience with Go or Rust.
  • Familiarity with the Model Context Protocol.
  • Experience managing hybrid cloud and on-premises GPU clusters for physical AI workloads.
  • Experience using LLMs for semantic monitoring and distributed-system log analysis.

Culture & Benefits

  • Collaborative environment focused on safe autonomous transportation and supply-chain resilience.
  • Commitment to diversity, inclusion, respect, agility, and professional growth.
  • Opportunity to work on proprietary software and hardware for commercially deployed Level 4 autonomous trucks.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →