Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Staff Software Engineer (Fleet Management) (Python/GPU infrastructure): Build and operate Fleet Manager, a workflow automation platform that provisions, validates, monitors, and remediates GPU nodes and network switches at scale with an accent on distributed systems, infrastructure automation, and hardware lifecycle management. Focus on designing durable event-driven workflows, ensuring idempotency and resumability, integrating datacenter systems, and maintaining reliable production operations.
Location: US
Salary: $220,000–$320,000 USD per year, plus potential bonus, equity, and/or commission.
Company
Nscale provides cost-effective, high-performance GPU cloud infrastructure for AI startups and enterprise customers.
What you will do
- Own domain-level architecture for provisioning, validation, remediation, or another major Fleet Manager area.
- Build Python automation for GPU node and network switch provisioning, burn-in testing, health validation, and remediation.
- Design durable workflows with checkpoints, retries, replay, failure handling, human approval gates, and support for thousands of concurrent executions.
- Integrate Fleet Manager with DCIM, NetBox, bare-metal provisioning, credential, monitoring, and cloud orchestration systems.
- Establish shared engineering patterns, libraries, conventions, and operational runbooks.
- Operate production systems through observability, alerting, incident response, mentoring, and technical influence across teams.
Requirements
- Extensive experience designing, building, and operating production distributed systems.
- Strong proficiency in Python.
- Strong understanding of event-driven and workflow architecture, including reliable delivery, idempotency, retries, replay, and failure handling.
- Experience delivering automation from ambiguous requirements through production and handling monitoring, incidents, and performance optimization.
- Ability to lead complex technical work across team boundaries through influence rather than formal authority.
- Excellent communication skills and experience using AI tools such as Claude or Cursor in development workflows.
Nice to have
- Experience with Temporal, Airflow, Prefect, or similar workflow orchestration tools.
- Experience with DCIM, NetBox, OpenStack, ERP, MAAS, Ironic, IPMI, PXE boot, or network automation.
- Experience automating hardware lifecycles, GPU infrastructure, burn-in testing, or cluster management.
- Knowledge of HPC networking, InfiniBand, RoCE, Kubernetes, Terraform, Pulumi, AWS, or GCP.
- Open-source contributions in infrastructure automation or cloud-native tooling.
Culture & Benefits
- Culture centered on innovation, ownership, accountability, openness, and transparency.
- Medical, dental, and vision benefits may be available.
- Flexible paid time off and parental leave may be available.
- Retirement plan participation may be available.
- Potential bonus, equity, and/or commission programs may be available.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →