Назад
Company hidden
5 часов назад

Software Engineer (AI)

Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer (AI): Operating and improving GPU cluster infrastructure for foundation model training and research with an accent on reliability, resource efficiency, and production-quality automation. Focus on debugging distributed compute, storage, networking, and scheduler issues, migrating workloads across GPU providers, and building monitoring and platform abstractions.

Location: San Francisco, United States; hybrid work

Company

Builds general-purpose AI systems designed to run efficiently across data center accelerators and on-device hardware.

What you will do

  • Own the reliability and operation of GPU clusters used for foundation model training and research.
  • Debug issues across compute, storage, networking, schedulers, and distributed workloads.
  • Improve CPU, GPU, and storage utilization through tooling and automation.
  • Onboard and migrate workloads across GPU providers and hardware platforms.
  • Build monitoring, validation, and platform abstractions that reduce operational work for researchers.
  • Contribute to the architecture of training infrastructure and the GPU platform.

Requirements

  • Strong software engineering experience building production-quality infrastructure tooling and automation.
  • Deep knowledge of distributed systems, Linux, networking, and storage.
  • Experience operating a shared compute cluster or distributed training platform.
  • Experience supporting production users and turning recurring failures into durable solutions.
  • Ability to partner effectively with senior research and infrastructure engineers.

Nice to have

  • Experience with SLURM, Kubernetes, Ray, Hadoop, or another distributed compute platform.
  • Experience supporting GPU, HPC, or large-scale AI training infrastructure.
  • Experience with distributed storage, cluster schedulers, cloud providers, or infrastructure control planes.

Culture & Benefits

  • High-impact ownership of infrastructure affecting foundation model training speed and efficiency.
  • Competitive base salary with equity in a unicorn-stage company.
  • 100% coverage of medical, dental, and vision premiums for employees and dependents.
  • 401(k) matching up to 4% of base pay.
  • Unlimited PTO and company-wide Refill Days throughout the year.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →