Назад
4 дня назад

Manager, Technical Support Engineering (Bare Metal)

157 000 - 210 000$
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Manager, Technical Support Engineering (Bare Metal) (GPU infrastructure and data center operations): Leading support operations for physical infrastructure that powers AI workloads, with an accent on hardware reliability, incident response, and customer escalations. Focus on building a scalable bare metal support team, diagnosing rack-scale CPU and GPU systems, and improving operational performance across 24/7 support workflows.

Location: San Francisco, CA; Seattle, WA; or Sunnyvale, CA, United States. Travel of up to 30% annually is required. The position requires access to export-controlled information and may require U.S. person status or eligibility for applicable export authorization.

Salary: $157,000–$210,000 base salary, plus discretionary bonus, equity awards, and benefits.

Company

AI cloud infrastructure company providing high-performance compute, tools, and technical expertise for AI labs, startups, and enterprises.

What you will do

  • Lead infrastructure support operations across multiple client environments and physical compute locations.
  • Build and manage a dedicated bare metal infrastructure support team.
  • Oversee infrastructure incidents, escalations, hardware operations, and customer communications.
  • Improve support processes, reliability, operational efficiency, and downtime performance.
  • Partner with product, infrastructure, engineering, and other internal teams to deliver infrastructure resources.
  • Mentor engineers and develop team capabilities through coaching, training, and performance management.

Requirements

  • 5+ years of experience leading teams in infrastructure support, data center operations, or physical compute environments.
  • Hands-on Linux system administration and command-line experience.
  • Experience diagnosing, troubleshooting, replacing, and managing server, power, cabling, CPU, and GPU hardware.
  • Knowledge of rack-scale GPU infrastructure, including NVIDIA A100/H100 systems, PCIe, NVLink, and liquid cooling, or the ability to learn HPC environments quickly.
  • Experience owning production-impacting incidents, client escalations, ticket workflows, and operational metrics such as MTTR, SLOs, backlog, and ticket trends.
  • Experience managing scheduling, shift coverage, and team logistics in 24/7 or hybrid support environments.

Nice to have

  • Experience scaling infrastructure support teams in high-growth environments.
  • Server and GPU hardware lifecycle management, including deployment, maintenance, thermal and power management, RMA coordination, and decommissioning.
  • Familiarity with AI/ML workloads, cluster utilization, and GPU-heavy customer infrastructure.

Culture & Benefits

  • Entrepreneurial, collaborative, fast-paced environment focused on innovative infrastructure solutions.
  • Medical, dental, and vision insurance fully paid by the employer, plus life and disability insurance.
  • 401(k) with employer match, flexible spending and health savings accounts, and employee stock purchase program.
  • Flexible PTO, paid parental leave, tuition reimbursement, mental wellness benefits, and family-forming support.
  • Flexible childcare support, catered meals at office and data center locations, and a casual work environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →