2 месяца назад
Infrastructure Software Engineer, Fleet & Automation (AI/HPC)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Infrastructure Software Engineer, Fleet & Automation (AI/HPC) (Python/infrastructure automation): Building control-plane, orchestration, and automation systems for GPU and network hardware lifecycle management with an accent on scalability, reliability, and observability. Focus on designing provisioning and remediation workflows, operating distributed infrastructure services, and improving AI/HPC fleet performance and operational efficiency.
Location: Houston, New York, San Francisco, or Seattle, United States
Company
Nscale provides a GPU cloud infrastructure platform for AI startups and enterprise customers, focused on high-performance, cost-effective AI computing.
What you will do
- Design the architecture, roadmap, and implementation of workflow automation systems for fleet, network, and observability operations.
- Own device provisioning, validation, testing, and remediation workflows for GPU nodes and network switches at scale.
- Build workflow orchestration and production-grade Python systems for hardware lifecycle management.
- Partner with Infrastructure, Platform, SRE, product, design, and operations teams to translate operational needs into scalable automation.
- Establish reliability, observability, and operational excellence standards across infrastructure services.
- Use AI tools to improve delivery, workflows, and operational efficiency.
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience.
- 5+ years of relevant experience building large-scale infrastructure applications or similar systems.
- Experience with C, C++, Java, and Python, including API design and unit testing.
- Deep knowledge of Linux, TCP/IP, BGP, and configuration management tools such as Ansible and Terraform.
- Experience building, operating, and debugging distributed infrastructure, stateful and stateless services, compute, storage, or hardware systems.
- Experience integrating with DCIMs, NetBox, OpenStack, and bare-metal APIs such as MAAS, Ironic, and IPMI.
Nice to have
- Master’s degree or PhD in Engineering, Computer Science, or a related field.
- Experience with AI/HPC infrastructure, NVIDIA GPUs, InfiniBand, high-speed Ethernet, NCCL, or SLURM.
- Experience with Prometheus, Grafana, OpenTelemetry, Kubernetes, Docker, and infrastructure-as-code principles.
- Experience improving system efficiency, scalability, performance, observability, and operator experience.
Culture & Benefits
- Collaborative, supportive, and innovation-focused environment.
- Competitive base salary and equity package.
- Compensation reviews every 12 months.
- Progression plan with opportunities to lead, challenge existing approaches, and own impact.
- Inclusive workplace with accommodations available for individual needs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Software Engineer (AI)
210 000 - 265 000$
6 дней назад
Software Engineer (Network Platforms)
68 000 - 91 000€
Anthropic
5 дней назад
Software Engineer, Infrastructure, Interpretability (AI)
405 000 - 485 000$
Lambda
7 дней назад
Senior Software Engineer (AI)
255 000 - 346 000$
Baseten
2 дня назад
Partner Engineer (AI)
265 000 - 330 000$
2 дня назад
Software Engineer – Scalable Systems (AI)
99 000 - 206 000$