5 часов назад
Software Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineer (AI): Operating and improving GPU cluster infrastructure for foundation model training and research with an accent on reliability, resource efficiency, and production-quality automation. Focus on debugging distributed compute, storage, networking, and scheduler issues, migrating workloads across GPU providers, and building monitoring and platform abstractions.
Location: San Francisco, United States; hybrid work
Company
Builds general-purpose AI systems designed to run efficiently across data center accelerators and on-device hardware.
What you will do
- Own the reliability and operation of GPU clusters used for foundation model training and research.
- Debug issues across compute, storage, networking, schedulers, and distributed workloads.
- Improve CPU, GPU, and storage utilization through tooling and automation.
- Onboard and migrate workloads across GPU providers and hardware platforms.
- Build monitoring, validation, and platform abstractions that reduce operational work for researchers.
- Contribute to the architecture of training infrastructure and the GPU platform.
Requirements
- Strong software engineering experience building production-quality infrastructure tooling and automation.
- Deep knowledge of distributed systems, Linux, networking, and storage.
- Experience operating a shared compute cluster or distributed training platform.
- Experience supporting production users and turning recurring failures into durable solutions.
- Ability to partner effectively with senior research and infrastructure engineers.
Nice to have
- Experience with SLURM, Kubernetes, Ray, Hadoop, or another distributed compute platform.
- Experience supporting GPU, HPC, or large-scale AI training infrastructure.
- Experience with distributed storage, cluster schedulers, cloud providers, or infrastructure control planes.
Culture & Benefits
- High-impact ownership of infrastructure affecting foundation model training speed and efficiency.
- Competitive base salary with equity in a unicorn-stage company.
- 100% coverage of medical, dental, and vision premiums for employees and dependents.
- 401(k) matching up to 4% of base pay.
- Unlimited PTO and company-wide Refill Days throughout the year.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Cognition
6 дней назад
Software Engineer, Infrastructure (AI)
4 часа назад
Software Engineer (AI Infrastructure)
4 часа назад
Production Engineer (AI Infrastructure)
172 000 - 209 000$
3 часа назад
Cluster Software Engineer (AI)
6 часов назад
Forward Deployed Engineers (AI Infrastructure)
180 000 - 240 000$
3 часа назад
Software Engineer: Infrastructure and Tools
125 000 - 170 000$