Назад
Company hidden
8 часов назад

Platform Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Platform Engineer (AI): Building and maintaining observable, secure, and scalable infrastructure for reliable and reproducible ML workloads with an accent on high-performance compute, deployment automation, and release safety. Focus on designing resilient CI/CD and deployment systems, hardening environments through load testing and chaos engineering, and improving reliability through incident response and root-cause analysis.

Location: San Francisco, United States; on-site, with an in-person team.

Requirements

  • Experience with high-performance compute environments, including ML clusters, GPU farms, and hyperscaler clouds such as AWS or GCP.
  • Infrastructure as code experience with tools such as Terraform or Ansible.
  • Experience with containers such as Docker or Apptainer and scheduling systems such as Kubernetes or Slurm.
  • Experience with deployment strategies at scale, disaster recovery planning, change management, fault tolerance, and reliable experimental environments.
  • Knowledge of compliance and audit standards for deployment and system security.
  • Experience with load testing, fault injection, and chaos engineering.

Nice to have

  • Experience supporting ML/AI infrastructure, including GPU management and workload optimization.
  • Exposure to backend development for ML model serving with tools such as vLLM, Ray, SGLang, or Triton.
  • Experience with software release engineering for ML/AI systems.

What you will do

  • Build and improve observability systems for monitoring, logging, and alerting.
  • Manage infrastructure as a service and CI/CD in partnership with engineering teams.
  • Design resilient build and deployment systems across research and production environments.
  • Implement secure, auditable release processes with rollback support.
  • Collaborate with ML, DevOps, and infrastructure teams to improve reliability and performance.
  • Lead incident response, root-cause analysis, and postmortems focused on prevention.

Culture & Benefits

  • Research and engineering work follows methodical, step-by-step approaches to ambitious goals.
  • New ideas are encouraged, with willingness to make significant bets on promising concepts.
  • Medical, dental, vision, and FSA plans.
  • Competitive compensation, 401(k) plan, unlimited PTO, and company holidays.
  • Relocation and immigration support may be available on a case-by-case basis.
  • In-office snacks and meals in a collaborative San Francisco environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →