4 дня назад
Software Engineer, Distributed Systems (AI)
300 000 - 350 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineer, Distributed Systems (AI): Building core distributed systems for AI training and serving platforms across thousands of GPU and TPU machines with an accent on orchestration, scheduling, storage, networking, and fault tolerance. Focus on designing consensus and replication mechanisms, solving performance and resource-allocation challenges, and maintaining reliable operation under hardware failures and network partitions.
Location: On-site in San Francisco, CA
Annual salary: $300,000–$350,000 USD, depending on background, skills, and experience.
Company
Thinking Machines Lab builds AI systems that extend human will and judgment, including frontier model training and model-serving platforms.
What you will do
- Design and build distributed systems for compute orchestration, scheduling, storage, and networking across large GPU and TPU clusters.
- Develop fault-tolerant systems that remain correct during hardware failures, network partitions, and workload growth.
- Build distributed storage and data-orchestration layers for large volumes of training and model data.
- Improve collective communication, scheduling, and resource allocation across thousands of machines.
- Partner with research and infrastructure teams to identify bottlenecks and design solutions from first principles.
- Write production-quality code and help shape company-wide systems architecture.
Requirements
- 5+ years of experience building large-scale distributed systems.
- Proficiency in Python and Go, C++, or another systems-level language.
- Strong understanding of consensus, consistency, replication, and fault tolerance.
- Experience with network programming, load balancing, or distributed storage systems.
- Fluent English and experience operating with high autonomy in a fast-changing environment.
Nice to have
- Experience building distributed compute or orchestration systems for AI or ML workloads.
- Experience with containerization, orchestration, and distributed compute frameworks.
- Experience integrating GPUs or TPUs into distributed training or serving systems.
- Background in AI research labs, high-performance computing centers, or similarly demanding environments.
- Published work or open-source contributions related to distributed systems or performance engineering.
Culture & Benefits
- Visa sponsorship is available, with support through the visa process for suitable candidates.
- Relocation support is available as needed.
- Health, dental, and vision benefits.
- Unlimited paid time off and paid parental leave.
- Work on frontier AI training and serving infrastructure in an early-stage environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
8 дней назад
Software Engineer (Robotics/Cloud Infrastructure)
200 000 - 300 000$
9 дней назад
Staff Software Engineer, Agentic AI
151 000 - 297 000$
Baseten
6 дней назад
Partner Engineer (AI)
265 000 - 330 000$
Deepgram
8 дней назад
Software Engineer (AI)
197 000 - 307 000$
Anthropic
9 дней назад
Software Engineer, Infrastructure, Interpretability (AI)
405 000 - 485 000$
TypeSafe AI
8 дней назад
Member of Technical Staff Backend Platform (AI)
150 000 - 250 000$