3 дня назад
Senior AI Infrastructure Engineer (Autonomous Vehicles)
180 000 - 240 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior AI Infrastructure Engineer (Autonomous Vehicles) (ML Infrastructure/MLOps): Building and scaling the high-performance AI platform that supports autonomous driving models with an accent on distributed training, GPU orchestration, model deployment, and agentic automation. Focus on optimizing multi-node H100/A100 clusters, developing self-healing infrastructure, and connecting research workflows with reliable production systems.
Location: Onsite five days a week at the Santa Clara, California office
Salary: $180,000–$240,000 per year
Company
develops Level 4 autonomous transportation technology for short-haul, B2B middle-mile logistics and operates commercially deployed autonomous trucks.
What you will do
- Design, build, and scale the AI infrastructure supporting perception, planning, world-model, and VLA research workloads.
- Enable distributed multi-node and multi-GPU training using PyTorch Distributed, FSDP, DeepSpeed, and Ray Train.
- Optimize GPU clusters, networking, and inference using NVIDIA hardware, NCCL, InfiniBand or RoCE v2, TensorRT, ONNX Runtime, and Triton.
- Develop agentic infrastructure automation for cluster monitoring, failure triage, CI/CD, resource optimization, and data curation.
- Implement MLOps workflows for experiment tracking, model lifecycle management, A/B testing, shadow deployments, and rollbacks.
- Build cloud-native data pipelines and observability systems with Airflow, Kafka, Spark, Prometheus, Grafana, OpenTelemetry, and ELK.
Requirements
- 5+ years of experience in ML infrastructure, MLOps, or DevOps for high-scale compute environments.
- Deep expertise in multi-GPU training, high-performance networking, Kubernetes, Terraform, and Helm.
- Experience with MLflow, Argo Workflows, Docker, and GPU-native orchestration.
- Proven experience building or supporting agentic workflows for infrastructure or data automation.
- Proficiency in Apache Airflow, Kafka, Spark, GitOps automation, Python, and Bash.
- Ability to work onsite five days per week in Santa Clara, California.
Nice to have
- Experience with Go or Rust.
- Familiarity with the Model Context Protocol.
- Experience managing hybrid cloud and on-premises GPU clusters for physical AI workloads.
- Experience using LLMs for semantic monitoring and distributed-system log analysis.
Culture & Benefits
- Collaborative environment focused on safe autonomous transportation and supply-chain resilience.
- Commitment to diversity, inclusion, respect, agility, and professional growth.
- Opportunity to work on proprietary software and hardware for commercially deployed Level 4 autonomous trucks.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
AI Infrastructure Engineer (Automotive)
140 000 - 170 000$
3 дня назад
Senior MLOps Engineer (AI)
3 дня назад
ML Training Infrastructure Engineer (AI)
220 000 - 320 000$
3 дня назад
AI Engineer
159 000 - 195 000$
1 день назад
Senior AI Engineer (Generative AI)
165 000 - 225 000$
2 дня назад
AI Engineer (Fintech)
180 000 - 260 000$