Назад
Company hidden
8 дней назад

Technical Lead (AI Infrastructure)

270 000 - 330 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior/lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Technical Lead (AI Infrastructure): Building and operating the machines layer of a high-performance serverless AI platform, with an accent on bare-metal fleets, cloud hosts, GPUs, networking, and distributed control-plane services. Focus on automating hardware provisioning, acceptance testing, machine recovery, kernel and image management, and fleet reliability across many datacenters.

Location: New York, United States; work is on-site.

Salary: $270,000–$330,000 per year

Company

hirify.global is building an AI infrastructure layer and a high-performance serverless platform for running functions, sandboxes, and training jobs.

What you will do

  • Lead 3–8 engineers building and operating the machines layer of the serverless platform.
  • Own the full machine lifecycle, including hardware onboarding, network bring-up, kernel and image management, GPU and disk health tracking, and automated remediation.
  • Design control-plane services for provisioning, monitoring, repairing, and managing bare-metal and cloud hosts.
  • Automate integration, acceptance testing, and benchmarking of CPU, GPU, storage, interconnect, and network hardware.
  • Standardize network configuration and monitor reliability across multiple datacenters.
  • Set technical direction and remain hands-on while participating in the on-call rotation and responding to production incidents.

Requirements

  • 7+ years of experience writing high-quality production code.
  • 3+ years of direct people-management experience, ideally leading engineering teams through planning, growth, and performance conversations.
  • Experience operating large-scale physical hardware fleets or building control planes that manage them, including bare-metal provisioning, BMC/IPMI, PXE, network boot, and firmware.
  • Strong cloud skills and knowledge of low-level operating-system foundations, including the Linux kernel, drivers, networking, file systems, and containers.
  • Experience working with hardware and colocation providers, including hardware acceptance testing and benchmarking.
  • Experience with GPUs and the NVIDIA software stack in production, plus prior experience with Go.

Nice to have

  • Experience with GPU health monitoring, RDMA, and NVLink.
  • Experience designing automatic remediation for power cycling, reimaging, and GPU recovery.
  • Experience managing heterogeneous CPU, GPU, storage, and network hardware across many datacenters.

Culture & Benefits

  • Hands-on engineering role with ownership across hardware, operating systems, networking, and distributed services.
  • Participation in an on-call rotation supporting production reliability.
  • Opportunity to shape the long-term infrastructure direction of an AI platform.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →