5 дней назад
Principal Software Engineer (Fleet Management)
240 000 - 400 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Software Engineer (Fleet Management/Python): Building and leading production-grade workflow automation for provisioning, testing, monitoring, and remediation of GPU nodes and network switches at scale with an accent on distributed systems, infrastructure reliability, and hardware lifecycle automation. Focus on designing orchestration systems, integrating bare-metal and cloud infrastructure tooling, and mentoring senior engineers while driving architecture and operational excellence.
Location: AMER
Base salary: $240,000–$400,000 USD per year, excluding potential bonus, equity, and commission.
Company
Nscale is building a vertically integrated GenAI cloud spanning sustainable data centers, AI infrastructure, and enterprise applications.
What you will do
- Lead the architecture, technical roadmap, and delivery of Fleet Manager workflow automation systems.
- Build device provisioning, enrolment, validation, burn-in testing, GPU health monitoring, and remediation workflows at scale.
- Design workflow orchestration for GPU nodes and network switches across their full lifecycle.
- Establish standards for reliability, observability, monitoring, alerting, SLOs, incident response, and operational excellence.
- Integrate DCIMs, NetBox, OpenStack, MAAS, Ironic, IPMI, and bare-metal APIs.
- Mentor senior engineers and partner with Infrastructure, Platform, and SRE teams on scalable automation.
Requirements
- 12–15+ years of software engineering experience building and operating production systems.
- Strong Python fundamentals and experience leading complex, multi-service distributed systems.
- Proven ownership of technical roadmaps and delivery of large-scale automation systems from ambiguous requirements to production.
- Deep understanding of infrastructure reliability, scalability, security, monitoring, alerting, incident response, and continuous improvement.
- Strong mentorship, communication, and technical leadership skills.
- Regular use of AI development tools such as Claude or Cursor.
Nice to have
- Experience with Temporal, Airflow, Prefect, or similar workflow orchestration tools.
- Experience with DCIMs, NetBox, OpenStack, ERP systems, MAAS, Ironic, IPMI, PXE boot, or network automation.
- GPU infrastructure, HPC, datacenter topology, InfiniBand, RoCE, or hardware lifecycle automation experience.
- Knowledge of Kubernetes, Terraform, Pulumi, AWS, and GCP.
- Open-source contributions in infrastructure automation or cloud-native tooling.
Culture & Benefits
- Culture centered on open collaboration, ownership, and engineering excellence.
- Medical, dental, and vision coverage.
- Flexible paid time off and parental leave.
- Retirement plan participation.
- Potential bonus, equity, and/or commission programs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Scale AI
10 дней назад
Senior Software Engineer, Orchestration Platform
216 000 - 270 000$
12 дней назад
Software Engineer – Scalable Systems (AI)
99 000 - 206 000$
Anthropic
12 дней назад
Staff + Sr. Software Engineer, Cloud Inference (AI)
320 000 - 485 000$
Anthropic
12 дней назад
Staff + Sr. Software Engineer, Cloud Inference Launch Engineering (AI)
320 000 - 485 000$
WRITER
12 дней назад
Software Engineer Agents (AI)
132 000 - 240 000$
Nebius
6 дней назад
Staff Software Engineer (Hardware Infrastructure)
175 000 - 225 000$