Назад
5 дней назад

Principal Software Engineer (Fleet Management)

240 000 - 400 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Software Engineer (Fleet Management/Python): Building and leading production-grade workflow automation for provisioning, testing, monitoring, and remediation of GPU nodes and network switches at scale with an accent on distributed systems, infrastructure reliability, and hardware lifecycle automation. Focus on designing orchestration systems, integrating bare-metal and cloud infrastructure tooling, and mentoring senior engineers while driving architecture and operational excellence.

Location: AMER

Base salary: $240,000–$400,000 USD per year, excluding potential bonus, equity, and commission.

Company

Nscale is building a vertically integrated GenAI cloud spanning sustainable data centers, AI infrastructure, and enterprise applications.

What you will do

  • Lead the architecture, technical roadmap, and delivery of Fleet Manager workflow automation systems.
  • Build device provisioning, enrolment, validation, burn-in testing, GPU health monitoring, and remediation workflows at scale.
  • Design workflow orchestration for GPU nodes and network switches across their full lifecycle.
  • Establish standards for reliability, observability, monitoring, alerting, SLOs, incident response, and operational excellence.
  • Integrate DCIMs, NetBox, OpenStack, MAAS, Ironic, IPMI, and bare-metal APIs.
  • Mentor senior engineers and partner with Infrastructure, Platform, and SRE teams on scalable automation.

Requirements

  • 12–15+ years of software engineering experience building and operating production systems.
  • Strong Python fundamentals and experience leading complex, multi-service distributed systems.
  • Proven ownership of technical roadmaps and delivery of large-scale automation systems from ambiguous requirements to production.
  • Deep understanding of infrastructure reliability, scalability, security, monitoring, alerting, incident response, and continuous improvement.
  • Strong mentorship, communication, and technical leadership skills.
  • Regular use of AI development tools such as Claude or Cursor.

Nice to have

  • Experience with Temporal, Airflow, Prefect, or similar workflow orchestration tools.
  • Experience with DCIMs, NetBox, OpenStack, ERP systems, MAAS, Ironic, IPMI, PXE boot, or network automation.
  • GPU infrastructure, HPC, datacenter topology, InfiniBand, RoCE, or hardware lifecycle automation experience.
  • Knowledge of Kubernetes, Terraform, Pulumi, AWS, and GCP.
  • Open-source contributions in infrastructure automation or cloud-native tooling.

Culture & Benefits

  • Culture centered on open collaboration, ownership, and engineering excellence.
  • Medical, dental, and vision coverage.
  • Flexible paid time off and parental leave.
  • Retirement plan participation.
  • Potential bonus, equity, and/or commission programs.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →