Назад
3 дня назад

Infrastructure Software Engineer, Fleet & Automation (AI/HPC)

Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Infrastructure Software Engineer, Fleet & Automation (AI/HPC) (Python/infrastructure automation): Building control-plane, orchestration, and automation systems for GPU and network hardware lifecycle management with an accent on scalability, reliability, and observability. Focus on designing provisioning and remediation workflows, operating distributed infrastructure services, and improving AI/HPC fleet performance and operational efficiency.

Location: Houston, New York, San Francisco, or Seattle, United States

Company

Nscale provides a GPU cloud infrastructure platform for AI startups and enterprise customers, focused on high-performance, cost-effective AI computing.

What you will do

  • Design the architecture, roadmap, and implementation of workflow automation systems for fleet, network, and observability operations.
  • Own device provisioning, validation, testing, and remediation workflows for GPU nodes and network switches at scale.
  • Build workflow orchestration and production-grade Python systems for hardware lifecycle management.
  • Partner with Infrastructure, Platform, SRE, product, design, and operations teams to translate operational needs into scalable automation.
  • Establish reliability, observability, and operational excellence standards across infrastructure services.
  • Use AI tools to improve delivery, workflows, and operational efficiency.

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience.
  • 5+ years of relevant experience building large-scale infrastructure applications or similar systems.
  • Experience with C, C++, Java, and Python, including API design and unit testing.
  • Deep knowledge of Linux, TCP/IP, BGP, and configuration management tools such as Ansible and Terraform.
  • Experience building, operating, and debugging distributed infrastructure, stateful and stateless services, compute, storage, or hardware systems.
  • Experience integrating with DCIMs, NetBox, OpenStack, and bare-metal APIs such as MAAS, Ironic, and IPMI.

Nice to have

  • Master’s degree or PhD in Engineering, Computer Science, or a related field.
  • Experience with AI/HPC infrastructure, NVIDIA GPUs, InfiniBand, high-speed Ethernet, NCCL, or SLURM.
  • Experience with Prometheus, Grafana, OpenTelemetry, Kubernetes, Docker, and infrastructure-as-code principles.
  • Experience improving system efficiency, scalability, performance, observability, and operator experience.

Culture & Benefits

  • Collaborative, supportive, and innovation-focused environment.
  • Competitive base salary and equity package.
  • Compensation reviews every 12 months.
  • Progression plan with opportunities to lead, challenge existing approaches, and own impact.
  • Inclusive workplace with accommodations available for individual needs.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →