Назад
Company hidden
13 дней назад

Senior Staff Engineer, ML Ops (R4941)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Staff Engineer, ML Ops (R4941) (Kubernetes/AI infrastructure): Building a Kubernetes-native AI Factory platform for developing, training, evaluating, and deploying autonomy systems with an accent on distributed training, GPU infrastructure, and researcher productivity. Focus on designing scalable MLOps capabilities across cloud, on-premise, sovereign, and air-gapped environments, while optimizing scheduling, model lifecycle management, and platform reliability.

Location: San Mateo, California; on-site

Company

hirify.global builds autonomy systems for defense applications across air, maritime, and space platforms.

What you will do

  • Design and implement the AI Factory Reference Architecture as a Kubernetes-native platform for AI development, distributed training, simulation, evaluation, and deployment.
  • Partner with ML researchers and autonomy teams to support foundation model development, reinforcement learning, distributed training, and evolving research workflows.
  • Build self-service developer workflows that support movement from local experimentation to large-scale distributed execution.
  • Develop shared GPU infrastructure and improve scheduling, storage, networking, observability, resource utilization, and reliability.
  • Enable dataset management, experiment tracking, artifact management, model versioning, evaluation, deployment, monitoring, and continuous model improvement.
  • Create repeatable infrastructure deployment and lifecycle management solutions for commercial cloud, on-premises, sovereign, and air-gapped environments.

Requirements

  • Experience building Kubernetes-native AI or MLOps platforms for distributed machine learning workloads.
  • Deep knowledge of PyTorch, Hugging Face Transformers, modern AI training frameworks, and distributed training techniques.
  • Experience operating GPU-accelerated infrastructure and distributed training systems.
  • Strong understanding of Kubernetes, Linux, networking, security, storage, and distributed systems.
  • Experience with Terraform, Helm, Python, Golang, and modern cloud-native technologies.
  • Experience translating ML research workflows into scalable platform capabilities.

Nice to have

  • Experience with Ray, KAI, Slurm, or other distributed AI orchestration and GPU scheduling technologies.
  • Experience with reinforcement learning, simulation-driven training, robotics, autonomy workloads, or edge AI deployment.
  • Experience designing infrastructure for classified, sovereign, or air-gapped environments.
  • Experience with OpenTelemetry, Prometheus, Grafana, or open-source infrastructure projects.

Culture & Benefits

  • Work at the intersection of modern AI, distributed systems, Kubernetes, and defense.
  • Build a reference architecture used by internal engineering teams and delivered to customers.
  • Contribute to an open-source-oriented AI infrastructure ecosystem.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →