Manager, Technical Support Engineering (Bare Metal)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Location: San Francisco, CA; Seattle, WA; or Sunnyvale, CA, United States. Travel of up to 30% annually is required. The position requires access to export-controlled information and may require U.S. person status or eligibility for applicable export authorization.
Salary: $157,000–$210,000 base salary, plus discretionary bonus, equity awards, and benefits.
Company
AI cloud infrastructure company providing high-performance compute, tools, and technical expertise for AI labs, startups, and enterprises.
What you will do
- Lead infrastructure support operations across multiple client environments and physical compute locations.
- Build and manage a dedicated bare metal infrastructure support team.
- Oversee infrastructure incidents, escalations, hardware operations, and customer communications.
- Improve support processes, reliability, operational efficiency, and downtime performance.
- Partner with product, infrastructure, engineering, and other internal teams to deliver infrastructure resources.
- Mentor engineers and develop team capabilities through coaching, training, and performance management.
Requirements
- 5+ years of experience leading teams in infrastructure support, data center operations, or physical compute environments.
- Hands-on Linux system administration and command-line experience.
- Experience diagnosing, troubleshooting, replacing, and managing server, power, cabling, CPU, and GPU hardware.
- Knowledge of rack-scale GPU infrastructure, including NVIDIA A100/H100 systems, PCIe, NVLink, and liquid cooling, or the ability to learn HPC environments quickly.
- Experience owning production-impacting incidents, client escalations, ticket workflows, and operational metrics such as MTTR, SLOs, backlog, and ticket trends.
- Experience managing scheduling, shift coverage, and team logistics in 24/7 or hybrid support environments.
Nice to have
- Experience scaling infrastructure support teams in high-growth environments.
- Server and GPU hardware lifecycle management, including deployment, maintenance, thermal and power management, RMA coordination, and decommissioning.
- Familiarity with AI/ML workloads, cluster utilization, and GPU-heavy customer infrastructure.
Culture & Benefits
- Entrepreneurial, collaborative, fast-paced environment focused on innovative infrastructure solutions.
- Medical, dental, and vision insurance fully paid by the employer, plus life and disability insurance.
- 401(k) with employer match, flexible spending and health savings accounts, and employee stock purchase program.
- Flexible PTO, paid parental leave, tuition reimbursement, mental wellness benefits, and family-forming support.
- Flexible childcare support, catered meals at office and data center locations, and a casual work environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →