Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Engineering Manager (AI Infrastructure): Leading teams responsible for deploying and operating production GPU fleet infrastructure with an accent on reliability, bare-metal provisioning, orchestration, and datacenter systems. Focus on managing large-scale compute environments, driving cross-functional deployments, improving automation and efficiency, and maintaining reliable systems with real SLAs.
Location: Hybrid, with presence in the San Francisco, San Jose, or Bellevue office 4 days per week; Tuesday is the designated work-from-home day.
Annual salary: $330,000–$440,000 for San Francisco or San Jose; $297,000–$396,000 for Bellevue.
Company
AI cloud infrastructure company building large-scale GPU compute systems for researchers, enterprises, and hyperscalers.
What you will do
- Lead and grow distributed engineering teams responsible for production systems infrastructure.
- Coordinate cross-functional projects and deployments, aligning stakeholders and delivering against deadlines.
- Improve efficiency through tooling, process optimization, and automation.
- Provide visibility into project progress, risks, outcomes, staffing, priorities, and deliverables.
- Participate in new technology qualification, incident management, and incident review programs.
- Conduct 1:1s, provide feedback, and support team members’ career development.
Requirements
- 3+ years of experience leading or managing engineers in AI/ML infrastructure or another large-scale compute environment.
- Experience owning production systems with real SLAs and balancing operational reliability with long-term technical improvements.
- Confidence working in Linux and debugging across operating system, hardware, and networking layers.
- Ability to lead technical design for medium-to-large initiatives, drive alignment, and deliver solutions.
- Experience building high-performing teams through hiring, upskilling, skills planning, performance management, and clear expectations.
- Strong problem-solving and troubleshooting skills, with interest in the intersection of hardware, software, and physical datacenter infrastructure.
Nice to have
- Linux systems administration, TCP/IP networking, automation, and scripting.
- Bare-metal provisioning and lifecycle management with PXE, Redfish, IPMI, BMC, DHCP, or DNS.
- Strong coding ability, APIs, distributed systems, and automation pipelines.
- GPU acceleration, virtualization, cloud computing, InfiniBand, racks, switches, and power domains.
- NetBox or similar source-of-truth and DCIM tooling, Linux distribution building, OS customization, or imaging.
Culture & Benefits
- Generous cash and equity compensation.
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles.
- 401(k) plan with a 2% company match for US employees.
- Flexible paid time off plan.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior Software Engineering Manager (AI Infrastructure)
5 дней назад
Engineering Manager (AI)
172 000 - 233 000$
4 дня назад
Senior Engineering Manager (AI)
215 000 - 250 000$
5 дней назад
Engineering Manager (AI Infrastructure)
250 000 - 340 000$
7 дней назад
Engineering Manager, Factory Core
220 000 - 280 000$
7 часов назад
Software Engineering Manager, Public Sector (AI)
162 400 - 270 000$