13 дней назад
Senior Staff Engineer, ML Ops (R4941)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Staff Engineer, ML Ops (R4941) (Kubernetes/AI infrastructure): Building a Kubernetes-native AI Factory platform for developing, training, evaluating, and deploying autonomy systems with an accent on distributed training, GPU infrastructure, and researcher productivity. Focus on designing scalable MLOps capabilities across cloud, on-premise, sovereign, and air-gapped environments, while optimizing scheduling, model lifecycle management, and platform reliability.
Location: San Mateo, California; on-site
Company
builds autonomy systems for defense applications across air, maritime, and space platforms.
What you will do
- Design and implement the AI Factory Reference Architecture as a Kubernetes-native platform for AI development, distributed training, simulation, evaluation, and deployment.
- Partner with ML researchers and autonomy teams to support foundation model development, reinforcement learning, distributed training, and evolving research workflows.
- Build self-service developer workflows that support movement from local experimentation to large-scale distributed execution.
- Develop shared GPU infrastructure and improve scheduling, storage, networking, observability, resource utilization, and reliability.
- Enable dataset management, experiment tracking, artifact management, model versioning, evaluation, deployment, monitoring, and continuous model improvement.
- Create repeatable infrastructure deployment and lifecycle management solutions for commercial cloud, on-premises, sovereign, and air-gapped environments.
Requirements
- Experience building Kubernetes-native AI or MLOps platforms for distributed machine learning workloads.
- Deep knowledge of PyTorch, Hugging Face Transformers, modern AI training frameworks, and distributed training techniques.
- Experience operating GPU-accelerated infrastructure and distributed training systems.
- Strong understanding of Kubernetes, Linux, networking, security, storage, and distributed systems.
- Experience with Terraform, Helm, Python, Golang, and modern cloud-native technologies.
- Experience translating ML research workflows into scalable platform capabilities.
Nice to have
- Experience with Ray, KAI, Slurm, or other distributed AI orchestration and GPU scheduling technologies.
- Experience with reinforcement learning, simulation-driven training, robotics, autonomy workloads, or edge AI deployment.
- Experience designing infrastructure for classified, sovereign, or air-gapped environments.
- Experience with OpenTelemetry, Prometheus, Grafana, or open-source infrastructure projects.
Culture & Benefits
- Work at the intersection of modern AI, distributed systems, Kubernetes, and defense.
- Build a reference architecture used by internal engineering teams and delivered to customers.
- Contribute to an open-source-oriented AI infrastructure ecosystem.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
12 дней назад
Staff AI Engineer (L4)
3 дня назад
AI Platform Engineer
130 000 - 180 000$
11 дней назад
Staff ML Infrastructure Engineer (Embodied AI)
171 700 - 335 300$
13 дней назад
Software Engineer, Machine Learning Infrastructure - USDS (AI)
136 800 - 359 720$
13 дней назад
Tech Lead Machine Learning Infrastructure Engineer - Recommendations and Search (AI)
187 040 - 438 000$
11 дней назад
Senior ML Infrastructure Engineer (Autonomous Driving)
128 700 - 261 300$