Назад
Company hidden
обновлено 6 дней назад

Machine Learning Infrastructure (AI)

150 000 - 350 000$
Формат работы
remote (только USA)/onsite
Тип работы
fulltime
Английский
b2
Страна
UK/US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning Infrastructure (AI) (Distributed Systems/ML Infrastructure): Building the orchestration layer that schedules, routes, and coordinates production AI workloads across thousands of nodes and heterogeneous hardware with an accent on distributed control planes, resource management, and large-scale ML serving. Focus on designing reliable systems for concurrency, coordination, fault tolerance, failure recovery, and production-scale operation.

Location: San Francisco, California; London Area/London (on-site); Houston, Texas; Mountain View, California; Santa Clara, California; remote work available in the United States.

Salary: $150,000–$350,000 annually.

Company

hirify.global represents a well-funded Series A AI infrastructure company building systems for the next generation of AI compute.

What you will do

  • Design and operate the orchestration layer for production AI compute.
  • Build systems that schedule, route, and coordinate AI workloads across thousands of nodes and heterogeneous hardware.
  • Develop production schedulers, orchestration platforms, distributed control planes, and resource-management or queueing systems.
  • Build Kubernetes-adjacent infrastructure rather than only operating Kubernetes clusters.
  • Develop large-scale ML serving and distributed-compute infrastructure.
  • Shape foundational infrastructure and contribute to engineering culture.

Requirements

  • Hands-on experience building and owning production schedulers or orchestration platforms.
  • Experience with distributed control planes, resource management, queueing systems, or large-scale ML infrastructure.
  • Ability to reason about concurrency, coordination, fault tolerance, failure recovery, and production reliability.
  • Experience with Go, C++, Python, RPC, or asynchronous systems.
  • Ability to clearly explain systems personally designed, shipped, scaled, and operated.

Culture & Benefits

  • Opportunity to influence foundational infrastructure for next-generation AI compute.
  • Opportunity to help shape the engineering culture of a Series A company.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →