Назад
4 дня назад

Software Engineer, Infrastructure (AI)

300 000 - 400 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer, Infrastructure (AI): Designing and operating distributed systems for large-scale model training and inference across thousands of accelerators with an accent on orchestration, scheduling, storage, resource management, and observability. Focus on debugging complex distributed failures, building Python and Go libraries and APIs, and improving infrastructure reliability and performance.

Location: On-site in San Francisco, CA; New York is also listed as a location.

Annual salary: $300,000–$400,000 USD, depending on background, skills, and experience.

Company

Thinking Machines Lab builds AI systems designed to extend human will and judgment, including frontier model training, customizable model infrastructure, and human-AI interfaces.

What you will do

  • Design, build, and operate distributed systems for large-scale model training and inference across thousands of accelerators.
  • Build and maintain orchestration, scheduling, storage, and resource management infrastructure.
  • Improve the reliability, performance, and observability of infrastructure used by research and product teams.
  • Partner with researchers and platform engineers to turn infrastructure needs into robust, well-abstracted systems.
  • Debug complex distributed failures across networking, storage, compute, and scheduling layers.
  • Develop and maintain internal libraries and APIs primarily in Python and Go.

Requirements

  • Demonstrated expertise designing and developing large-scale distributed systems.
  • Strong proficiency in Python and Go.
  • Experience building, deploying, and operating production infrastructure at scale.
  • Strong understanding of consensus, consistency, fault tolerance, and networking.
  • Ability to work on-site in San Francisco, CA.
  • Visa sponsorship is available for qualified candidates.

Nice to have

  • Experience with ML infrastructure, including training orchestration, job schedulers, or distributed storage and data systems.
  • Experience operating large-scale GPU or TPU clusters.
  • Experience with Kubernetes and infrastructure-as-code.
  • Contributions to open-source infrastructure projects.
  • Comfort working with high autonomy in a fast-changing, early-stage environment.

Culture & Benefits

  • Foundational infrastructure ownership in a fast-moving startup environment.
  • Work directly with research and product teams on systems operating at large scale.
  • Health, dental, and vision benefits.
  • Unlimited paid time off and paid parental leave.
  • Relocation support as needed.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →