4 дня назад
Infrastructure Engineer (AI)
250 000 - 300 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Infrastructure Engineer (AI) (AI infrastructure): Building and operating the machines layer and control plane for a high-performance serverless AI platform with an accent on bare-metal fleet automation, GPUs, networking, storage, and hardware lifecycle management. Focus on designing automatic remediation, integrating heterogeneous CPU and GPU capacity, debugging across hardware and software layers, and maintaining reliable infrastructure across many datacenters.
Location: On-site in San Francisco or New York, United States
Salary: $250,000–$300,000 per year
Company
is building an AI infrastructure layer and a serverless platform for Functions, Sandboxes, and training workloads.
What you will do
- Design, build, and maintain the machines layer and control plane for bare-metal and cloud hosts.
- Automate the integration, acceptance testing, benchmarking, imaging, and production deployment of CPU, GPU, storage, and network servers.
- Configure and operate GPUs, RDMA, networking, storage, firmware, kernels, machine images, and custom network boot infrastructure.
- Build automatic remediation for unhealthy machines, including power cycling, reimaging, and GPU recovery.
- Monitor network and hardware health across datacenters and improve fleet reliability.
- Participate in the on-call rotation and respond to production incidents across hardware and software layers.
Requirements
- 5+ years of experience writing high-quality production code.
- Experience operating physical hardware fleets or building control planes for bare-metal infrastructure.
- Strong cloud skills and knowledge of Linux kernels, drivers, networking, filesystems, and containers.
- Experience debugging across infrastructure layers, from BGP and Linux networking to GPU firmware and control-plane services.
- Production experience with GPUs and the NVIDIA software stack, including drivers, health monitoring, XIDs, RDMA, or NVLink.
- Willingness to participate in an on-call rotation and respond to production incidents.
Nice to have
- Prior experience with Go.
Culture & Benefits
- Work on high-performance AI infrastructure and a rapidly growing serverless platform.
- Collaborate with engineers who have created open-source projects, conducted academic research, and led engineering and product organizations.
- Operate infrastructure spanning multiple hardware providers and datacenters.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Staff Software Engineer (Kubernetes)
215 000 - 265 000$
10 дней назад
Member of Technical Staff, Infrastructure (AI)
150 000 - 390 000$
7 дней назад
Senior Software Engineer - Platform (AI)
145 000 - 198 000$
9 дней назад
Technology Enablement Engineer (AI Infrastructure)
129 600 - 190 067$
3 дня назад
GPU Infrastructure Engineer
150 000 - 300 000$
10 дней назад
Platform Engineer - AI Engineering (AI)
130 000 - 190 000$