8 дней назад
Technical Lead (AI Infrastructure)
270 000 - 330 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Technical Lead (AI Infrastructure): Building and operating the machines layer of a high-performance serverless AI platform, with an accent on bare-metal fleets, cloud hosts, GPUs, networking, and distributed control-plane services. Focus on automating hardware provisioning, acceptance testing, machine recovery, kernel and image management, and fleet reliability across many datacenters.
Location: New York, United States; work is on-site.
Salary: $270,000–$330,000 per year
Company
is building an AI infrastructure layer and a high-performance serverless platform for running functions, sandboxes, and training jobs.
What you will do
- Lead 3–8 engineers building and operating the machines layer of the serverless platform.
- Own the full machine lifecycle, including hardware onboarding, network bring-up, kernel and image management, GPU and disk health tracking, and automated remediation.
- Design control-plane services for provisioning, monitoring, repairing, and managing bare-metal and cloud hosts.
- Automate integration, acceptance testing, and benchmarking of CPU, GPU, storage, interconnect, and network hardware.
- Standardize network configuration and monitor reliability across multiple datacenters.
- Set technical direction and remain hands-on while participating in the on-call rotation and responding to production incidents.
Requirements
- 7+ years of experience writing high-quality production code.
- 3+ years of direct people-management experience, ideally leading engineering teams through planning, growth, and performance conversations.
- Experience operating large-scale physical hardware fleets or building control planes that manage them, including bare-metal provisioning, BMC/IPMI, PXE, network boot, and firmware.
- Strong cloud skills and knowledge of low-level operating-system foundations, including the Linux kernel, drivers, networking, file systems, and containers.
- Experience working with hardware and colocation providers, including hardware acceptance testing and benchmarking.
- Experience with GPUs and the NVIDIA software stack in production, plus prior experience with Go.
Nice to have
- Experience with GPU health monitoring, RDMA, and NVLink.
- Experience designing automatic remediation for power cycling, reimaging, and GPU recovery.
- Experience managing heterogeneous CPU, GPU, storage, and network hardware across many datacenters.
Culture & Benefits
- Hands-on engineering role with ownership across hardware, operating systems, networking, and distributed services.
- Participation in an on-call rotation supporting production reliability.
- Opportunity to shape the long-term infrastructure direction of an AI platform.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
Platform Engineer (AI)
250 000 - 315 000$
11 дней назад
Senior Platform Engineer (AI)
190 000 - 265 000$
vCluster
10 дней назад
Staff Software Engineer (Kubernetes)
75 000 - 145 000€
12 дней назад
Agentic Harness Engineer (AI)
180 000 - 235 000$
13 дней назад
Senior Software Engineer, Infrastructure (AWS)
135 150 - 187 000$
10 дней назад
Staff Software Engineer (Kubernetes)
215 000 - 265 000$