Назад
Company hidden
4 часа назад

Software Engineer (AI Infrastructure)

170 000 - 205 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer (AI Infrastructure): Developing diagnostic, observability, automation, and repair tooling for high-performance GPU compute clusters with an accent on distributed systems, reliability, and hardware operations. Focus on building AI agents for component-level diagnosis and remediation, validating GPU systems with PyTorch and NVIDIA NCCL, and maintaining fleet availability across data center environments.

Location: San Francisco, CA, United States; on-site

Salary: $170,000–$205,000 per year plus bonus

Company

hirify.global builds vertically integrated energy and AI infrastructure, operating systems from power generation through cloud computing to support large-scale AI workloads.

What you will do

  • Develop deep-level diagnostics and troubleshooting tools for hardware faults in GPU racks and high-density compute systems.
  • Build automation and troubleshooting tooling for NVIDIA A100, H200, GB200, B200, and AMD 350X/355X GPU platforms.
  • Develop AI agents for component-level diagnosis, remediation, critical-environment management, and hardware repair workflows.
  • Create post-repair validation and testing tools using burn-in testing, PyTorch, and NVIDIA NCCL.
  • Own deployment, monitoring, and operational support for tooling that improves GPU fleet availability and performance.
  • Develop automation for facility power management and direct liquid-cooling hardware systems.

Requirements

  • 4–6 years of software engineering experience.
  • Strong programming skills in at least one of Go, Python, Java, or Rust.
  • Expertise in distributed systems, reliability, and cloud platforms such as Kubernetes, infrastructure as code, and GCP.
  • Ability to identify problems, develop scalable solutions, set technical direction, and deliver independently.
  • Strong analytical, problem-solving, communication, and collaboration skills.
  • Ability to work on-site in San Francisco, United States.

Nice to have

  • Experience with Temporal and Kubernetes.
  • Experience working directly with hardware vendors.
  • Background in large-scale GPU fleet operations or hyperscale data center environments.

Culture & Benefits

  • Industry-competitive compensation with restricted stock units and bonus eligibility.
  • Health, vision, dental, HSA, life insurance, and disability coverage.
  • 401(k) with a 100% employer match up to 4% of salary.
  • Paid parental leave, generous paid time off, and holiday schedule.
  • Tuition reimbursement, cell phone reimbursement, Calm subscription, and legal services.
  • Company-paid commuter benefit of $50 per pay period.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →