Назад
3 дня назад

Operations Manager (AI Infrastructure)

109 000 - 145 000$
Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Operations Manager (AI Infrastructure): Manage and improve quality and operational processes for AI cloud infrastructure hardware and data center systems with an accent on root cause analysis, quality management systems, and cross-team collaboration. Focus on analyzing production metrics, driving corrective actions, and ensuring operational reliability in a hybrid work environment.

Location: San Jose, CA office with hybrid work format; presence required 4 days per week

Salary: $109,000–$145,000 per year

Company

Lambda is a leading AI cloud infrastructure company focused on delivering superintelligence compute power to researchers and enterprises worldwide.

What you will do

  • Track and manage quality issues in data center deployment and production environments
  • Perform root cause analysis for hardware, software, and process failures
  • Analyze system metrics and quality data to identify trends and weak points
  • Drive corrective and preventive actions and improve RMA turnaround time
  • Collaborate cross-functionally with operations, engineering, supply chain, and vendors
  • Maintain and update quality management systems and report quality KPIs to leadership

Requirements

  • Must be located in or near San Jose, CA with hybrid work presence 4 days per week
  • Experience with hardware, data center, or infrastructure systems
  • Strong skills in data analysis, statistics, and root cause analysis methods
  • Experience with quality management systems and tools
  • Effective cross-team communication and stakeholder management skills
  • Fluent English communication skills (written and verbal)

Nice to have

  • Experience in machine learning, AI infrastructure, GPU, HPC, or computer hardware industries
  • Familiarity with data center standards and certifications (e.g., ISO, Uptime Institute)
  • Knowledge of firmware, embedded systems, and reliability engineering
  • Skills in scripting or automation (Python, SQL) for data processing
  • Exposure to cloud or hyperscaler infrastructure operations

Culture & Benefits

  • Generous cash and equity compensation
  • Health, dental, and vision coverage for employees and dependents
  • Wellness and commuter stipends for select roles
  • 401k plan with 2% company match (for US employees)
  • Flexible paid time off plan actively used by employees

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →