Назад
Company hidden
9 часов назад

Senior Cloud Infrastructure Engineer (AI)

180 000 - 240 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Cloud Infrastructure Engineer (AI): Architecting and managing large-scale compute, data infrastructure, and automated pipelines for an autonomous driving AI platform with an accent on Kubernetes GPU clusters, distributed training, and MLOps. Focus on optimizing multi-node GPU communication, building resilient infrastructure-as-code workflows, and deploying models across simulation and production environments.

Location: Onsite five days a week at the Santa Clara, California office

Salary: $180,000–$240,000 per year

Company

hirify.global develops autonomous transportation technology for short-haul, B2B middle-mile logistics, integrating proprietary Level 4 autonomous software and hardware with commercial freight operations.

What you will do

  • Architect and maintain Kubernetes clusters optimized for high-volume GPU and TPU workloads, including GPU scheduling and self-healing infrastructure.
  • Build infrastructure-as-code and GitOps workflows with Terraform, Helm, ArgoCD, and GitLab CI/CD.
  • Develop large-scale autonomy data pipelines with Apache Airflow, Kafka, and Spark.
  • Maintain observability for infrastructure and model-serving performance using Prometheus, Grafana, and OpenTelemetry.
  • Design ML model tracking, feature-store integrations, lifecycle automation, and serving workflows with MLflow, Airflow, Kubernetes, Triton, Ray Serve, and ONNX Runtime.
  • Support distributed training across multi-node GPU clusters and optimize NCCL, InfiniBand, RoCE v2, FSDP, and DeepSpeed workloads.

Requirements

  • 5+ years of experience in cloud infrastructure, DevOps, or MLOps supporting high-scale compute environments.
  • Deep expertise in Kubernetes, Helm, and container orchestration.
  • Strong experience with Apache Airflow, Argo Workflows, MLflow, and Terraform.
  • Practical experience supporting Ray and PyTorch Distributed.
  • Proficiency in Python and Bash scripting, with a solid understanding of IAM and RBAC.
  • Ability to work onsite five days per week at the Santa Clara office.

Nice to have

  • Deep understanding of FSDP and DeepSpeed.
  • Experience building agentic workflows with LangGraph or AutoGen for infrastructure automation or data curation.
  • Familiarity with Model Context Protocol (MCP).

Culture & Benefits

  • Collaborative, respectful, and agile working culture.
  • Focus on diversity, inclusion, and equal opportunities for growth.
  • Opportunity to contribute to autonomous transportation and a more resilient, sustainable supply chain.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →