Назад
4 дня назад

Senior HPC Support Engineer (AI)

122 000 - 162 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Senior HPC Support Engineer (AI): Troubleshooting and optimizing high-performance computing infrastructure for an AI cloud platform with an accent on hardware-level debugging and cluster orchestration. Focus on solving complex kernel and driver issues, optimizing GPU infrastructure, and improving operational tooling.

Location: Remote (USA)

Salary: $122k – $162k

Company

Lambda is a leader in AI cloud infrastructure providing GPU compute and superintelligence tools to researchers and enterprises.

What you will do

  • Serve as the senior technical escalation point for the hardest infrastructure and platform issues, debugging down to the hardware, driver, or kernel level.
  • Identify and resolve hardware failures, driver issues, and customer workload misconfigurations.
  • Build scripts and internal automations using AI tools to close operational gaps.
  • Perform root-cause analysis across distributed systems, clusters, and GPU infrastructure.
  • Collaborate with engineering teams to transform recurring customer pain points into permanent fixes.
  • Mentor junior support engineers and lead the resolution of major incidents during on-call rotations.

Requirements

  • 3+ years of hands-on HPC experience in an administration, support, or engineering role.
  • Advanced Linux system administration skills.
  • Proven expertise in Linux cluster administration, specifically with Kubernetes and/or Slurm.
  • Proficiency with monitoring and logging tools such as Prometheus, Grafana, and Datadog.
  • Deep knowledge of CUDA, NCCL, NVLink, GPUDirect RDMA, and high-throughput networking (IB/RoCE).
  • Must be based in the USA.

Nice to have

  • Experience with Docker, Kubernetes, or neocloud/GPU cloud providers.
  • Familiarity with infrastructure-as-code tools like Terraform and Ansible.
  • Experience with Nvidia GPUs and Infiniband.
  • Experience with high-performance storage systems.

Culture & Benefits

  • Generous cash and equity compensation packages.
  • Comprehensive health, dental, and vision coverage for employees and dependents.
  • 401k plan with a 2% company match for USA employees.
  • Flexible paid time off (PTO) plan.
  • Wellness and commuter stipends for select roles.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →