Назад
15 часов назад

Senior Software Engineer (DCIE) (GPU Infrastructure)

170 000 - 205 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Страна
US/Canada
vacancy_detail.hirify_telegram_tooltipВакансия из Telegram канала -

Мэтч & Сопровод

Покажет вашу совместимость и напишет письмо

Описание вакансии

TL;DR
Senior Software Engineer (DCIE) (GPU Infrastructure): Develop software for GPU servers and data centers with an accent on diagnostics, observability, automation, and reliability. Focus on building AI agents and validation tooling for NVIDIA and AMD GPU clusters, troubleshooting hardware faults, and supporting liquid-cooling and facility power-management systems.

Senior Software Engineer (DCIE)

Company

Crusoe

Conditions

6 days agoSeniorSalary: 170K - 205KSan Francisco, CA - US Onsite Full Time Engineering Jobs by Crusoe

Skills

Ai Agent Amd Automation Diagnostics Distributed Systems Gcp Go Gpu Infrastructure-As-Code Java Kubernetes Liquid Cooling Nccl Nvidia Observability Python Pytorch Reliability Rust Software Engineering Temporal Troubleshooting

Candidate Availability

Onsite · Canada · Required Onsite · San Francisco · Required Onsite · United States · Required Onsite · Required Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You develop software that manages GPU servers and data centers. You build diagnostics, observability, automation, repair, validation, and operational tooling for high-performance GPU clusters, deploy and monitor those tools, and support facilities management and liquid cooling systems.

Requirements

  • 4-6 years of software engineering experience
  • Distributed systems expertise
  • Reliability expertise
  • Cloud platform expertise
  • Experience with Kubernetes, infrastructure as code, and GCP
  • Proficiency in Go, Python, Java, or Rust
  • Analytical skills
  • Problem-solving skills
  • Communication skills
  • Collaboration skills
  • Experience with Temporal and Kubernetes
  • Experience working with hardware vendors
  • GPU fleet operations or hyperscale data center experience

Responsibilities

  • Develop deep-level diagnostics and troubleshoot hardware faults within GPU racks and high-density compute systems
  • Develop troubleshooting and automation tooling for NVIDIA and AMD GPU platforms
  • Develop automation and AI agents for component-level diagnosis and hardware remediation
  • Develop tooling and AI agents for managing critical data center environments
  • Develop post-repair validation and testing tools using burn-in, PyTorch, and NVIDIA NCCL
  • Deploy, monitor, and provide operational support for developed tooling
  • Develop automation and operational tooling for facility power management and direct liquid cooling systems

Benefits

  • Restricted Stock Units
  • Health insurance including HDHP and PPO options
  • Vision insurance
  • Dental insurance
  • Employer contributions to HSA accounts
  • Paid parental leave
  • Paid life insurance
  • Short-term disability insurance
  • Long-term disability insurance
  • Teladoc
  • 401(k) with a 100% match up to 4% of salary
  • Paid time off
  • Paid holidays
  • Cell phone reimbursement
  • Tuition reimbursement
  • Calm app subscription
  • MetLife Legal
  • Company-paid commuter benefit

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →

Текст вакансии взят без изменений

Источник -