Назад
Company hidden
4 часа назад

Senior Software Engineer, AI Infrastructure

126 000 - 189 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Software Engineer, AI Infrastructure (Go/Python/HPC): Building and scaling orchestration, scheduling, and execution infrastructure for training AI models across large GPU clusters with an accent on distributed systems, Linux internals, and performance engineering. Focus on automating HPC operations, debugging complex distributed failures, optimizing GPU workloads, and managing cluster compute, networking, and storage.

Location: On-site in Seattle, Washington, United States. On-site requirements vary by position and team.

Salary: $126,000–$189,000 base salary per year, plus bonus plans.

Company

hirify.global is a Seattle-based non-profit AI research institute building open, large-scale AI models, data, robotics, and conservation technologies for public benefit.

What you will do

  • Design and scale orchestration systems that intelligently schedule high-value workloads across large GPU clusters.
  • Build software-defined infrastructure, automation tooling, and cluster health management systems.
  • Optimize distributed workloads and conduct root-cause analysis of complex system failures.
  • Own systems across the Beaker job scheduler and execution runtime.
  • Provide technical input on compute, networking, storage, and large-scale HPC infrastructure deployments.
  • Review code and design documents, mentor engineers, and collaborate with research staff.

Requirements

  • 8+ years of experience developing business-critical software and operating large-scale compute infrastructure.
  • Proficiency in Go and/or Python.
  • Expert-level knowledge of Linux internals and container runtimes such as Docker.
  • Experience designing, debugging, and optimizing high-scale distributed systems and databases.
  • Bachelor’s degree in a related field, or an advanced degree substituting for equivalent technical experience.
  • Authorization to work in the United States is required.

Nice to have

  • Experience with workload schedulers such as Kubernetes or Slurm.
  • Experience with NCCL, InfiniBand, HPC system administration, or SRE.
  • Experience training or fine-tuning frontier AI models.
  • Open-source infrastructure or orchestration contributions.
  • Familiarity with WEKA or Ceph storage systems.

Culture & Benefits

  • Open science environment focused on transparent, open-source AI infrastructure and scientific impact.
  • Medical, dental, vision, employee assistance, HSA, HRA, and flexible spending account plans.
  • 401(k) plan, annual bonuses, and long-term incentive plan.
  • Paid sick leave, personal days, vacation, holidays, and family leave.
  • Monthly commuting or internet support and fitness and wellbeing expenses.
  • Learning opportunities through ongoing education and Ai2 Academy lectures.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →