4 часа назад
Senior Software Engineer, AI Infrastructure
126 000 - 189 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Software Engineer, AI Infrastructure (Go/Python/HPC): Building and scaling orchestration, scheduling, and execution infrastructure for training AI models across large GPU clusters with an accent on distributed systems, Linux internals, and performance engineering. Focus on automating HPC operations, debugging complex distributed failures, optimizing GPU workloads, and managing cluster compute, networking, and storage.
Location: On-site in Seattle, Washington, United States. On-site requirements vary by position and team.
Salary: $126,000–$189,000 base salary per year, plus bonus plans.
Company
is a Seattle-based non-profit AI research institute building open, large-scale AI models, data, robotics, and conservation technologies for public benefit.
What you will do
- Design and scale orchestration systems that intelligently schedule high-value workloads across large GPU clusters.
- Build software-defined infrastructure, automation tooling, and cluster health management systems.
- Optimize distributed workloads and conduct root-cause analysis of complex system failures.
- Own systems across the Beaker job scheduler and execution runtime.
- Provide technical input on compute, networking, storage, and large-scale HPC infrastructure deployments.
- Review code and design documents, mentor engineers, and collaborate with research staff.
Requirements
- 8+ years of experience developing business-critical software and operating large-scale compute infrastructure.
- Proficiency in Go and/or Python.
- Expert-level knowledge of Linux internals and container runtimes such as Docker.
- Experience designing, debugging, and optimizing high-scale distributed systems and databases.
- Bachelor’s degree in a related field, or an advanced degree substituting for equivalent technical experience.
- Authorization to work in the United States is required.
Nice to have
- Experience with workload schedulers such as Kubernetes or Slurm.
- Experience with NCCL, InfiniBand, HPC system administration, or SRE.
- Experience training or fine-tuning frontier AI models.
- Open-source infrastructure or orchestration contributions.
- Familiarity with WEKA or Ceph storage systems.
Culture & Benefits
- Open science environment focused on transparent, open-source AI infrastructure and scientific impact.
- Medical, dental, vision, employee assistance, HSA, HRA, and flexible spending account plans.
- 401(k) plan, annual bonuses, and long-term incentive plan.
- Paid sick leave, personal days, vacation, holidays, and family leave.
- Monthly commuting or internet support and fitness and wellbeing expenses.
- Learning opportunities through ongoing education and Ai2 Academy lectures.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Lead AI Software Engineer (AI Platform & Architecture)
180 000 - 225 000$
5 дней назад
Staff AI Cloud Engineer (AI)
180 000 - 225 000$
2 часа назад
Senior/Staff Software Engineer (AI)
235 030 - 352 290$
5 дней назад
AI Infrastructure Engineer (GPU)
170 500 - 315 490$
5 часов назад
Software Engineer (AI/Data Processing)
215 000 - 230 000$
5 часов назад
Systems Engineer (AI/HPC)
160 000 - 320 000$