Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Infrastructure Software Engineer, Fleet & Automation (AI/HPC) (Python/infrastructure automation): Building control-plane, orchestration, and automation systems for GPU and network hardware lifecycle management with an accent on scalability, reliability, and observability. Focus on designing provisioning and remediation workflows, operating distributed infrastructure services, and improving AI/HPC fleet performance and operational efficiency.
Location: Houston, New York, San Francisco, or Seattle, United States
Company
Nscale provides a GPU cloud infrastructure platform for AI startups and enterprise customers, focused on high-performance, cost-effective AI computing.
What you will do
- Design the architecture, roadmap, and implementation of workflow automation systems for fleet, network, and observability operations.
- Own device provisioning, validation, testing, and remediation workflows for GPU nodes and network switches at scale.
- Build workflow orchestration and production-grade Python systems for hardware lifecycle management.
- Partner with Infrastructure, Platform, SRE, product, design, and operations teams to translate operational needs into scalable automation.
- Establish reliability, observability, and operational excellence standards across infrastructure services.
- Use AI tools to improve delivery, workflows, and operational efficiency.
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience.
- 5+ years of relevant experience building large-scale infrastructure applications or similar systems.
- Experience with C, C++, Java, and Python, including API design and unit testing.
- Deep knowledge of Linux, TCP/IP, BGP, and configuration management tools such as Ansible and Terraform.
- Experience building, operating, and debugging distributed infrastructure, stateful and stateless services, compute, storage, or hardware systems.
- Experience integrating with DCIMs, NetBox, OpenStack, and bare-metal APIs such as MAAS, Ironic, and IPMI.
Nice to have
- Master’s degree or PhD in Engineering, Computer Science, or a related field.
- Experience with AI/HPC infrastructure, NVIDIA GPUs, InfiniBand, high-speed Ethernet, NCCL, or SLURM.
- Experience with Prometheus, Grafana, OpenTelemetry, Kubernetes, Docker, and infrastructure-as-code principles.
- Experience improving system efficiency, scalability, performance, observability, and operator experience.
Culture & Benefits
- Collaborative, supportive, and innovation-focused environment.
- Competitive base salary and equity package.
- Compensation reviews every 12 months.
- Progression plan with opportunities to lead, challenge existing approaches, and own impact.
- Inclusive workplace with accommodations available for individual needs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Ground Systems Software Engineer (Aerospace)
131 000 - 184 000$
5 дней назад
Software Engineer (Kubernetes)
215 000 - 275 000$
3 дня назад
Senior Backend Engineer (AI)
180 000 - 250 000$
Lambda
3 дня назад
Staff Software Engineer (AI Cloud)
314 000 - 419 000$
3 дня назад
Senior Software Engineer, Training & Experimentation (AI)
180 000 - 250 000$
5 дней назад