4 дня назад
Software Engineer, Infrastructure (AI)
300 000 - 400 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineer, Infrastructure (AI): Designing and operating distributed systems for large-scale model training and inference across thousands of accelerators with an accent on orchestration, scheduling, storage, resource management, and observability. Focus on debugging complex distributed failures, building Python and Go libraries and APIs, and improving infrastructure reliability and performance.
Location: On-site in San Francisco, CA; New York is also listed as a location.
Annual salary: $300,000–$400,000 USD, depending on background, skills, and experience.
Company
Thinking Machines Lab builds AI systems designed to extend human will and judgment, including frontier model training, customizable model infrastructure, and human-AI interfaces.
What you will do
- Design, build, and operate distributed systems for large-scale model training and inference across thousands of accelerators.
- Build and maintain orchestration, scheduling, storage, and resource management infrastructure.
- Improve the reliability, performance, and observability of infrastructure used by research and product teams.
- Partner with researchers and platform engineers to turn infrastructure needs into robust, well-abstracted systems.
- Debug complex distributed failures across networking, storage, compute, and scheduling layers.
- Develop and maintain internal libraries and APIs primarily in Python and Go.
Requirements
- Demonstrated expertise designing and developing large-scale distributed systems.
- Strong proficiency in Python and Go.
- Experience building, deploying, and operating production infrastructure at scale.
- Strong understanding of consensus, consistency, fault tolerance, and networking.
- Ability to work on-site in San Francisco, CA.
- Visa sponsorship is available for qualified candidates.
Nice to have
- Experience with ML infrastructure, including training orchestration, job schedulers, or distributed storage and data systems.
- Experience operating large-scale GPU or TPU clusters.
- Experience with Kubernetes and infrastructure-as-code.
- Contributions to open-source infrastructure projects.
- Comfort working with high autonomy in a fast-changing, early-stage environment.
Culture & Benefits
- Foundational infrastructure ownership in a fast-moving startup environment.
- Work directly with research and product teams on systems operating at large scale.
- Health, dental, and vision benefits.
- Unlimited paid time off and paid parental leave.
- Relocation support as needed.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior AI Engineer (MLOps)
164 800 - 247 000$
Microsoft AI
9 дней назад
Software Engineer (AI Infra and Model Foundry)
142 800 - 274 800$
Baseten
8 дней назад
Software Engineer (AI Inference)
165 000 - 330 000$
5 дней назад
Senior AI/ML Platform Engineer
148 500 - 221 000$
4 дня назад
ML Platform Engineer (AI)
100 000 - 160 000$
Reddit
4 дня назад
Staff Machine Learning Engineer (ML Efficiency)
230 000 - 322 000$