Назад
2 дня назад

Software Engineering Manager (AI Infrastructure)

25 000 - 29 167$
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Software Engineering Manager (AI Infrastructure): Leading the team building Fleet Manager, a Python-based workflow automation platform that provisions, tests, monitors, and remediates GPU nodes and network switches at scale with an accent on distributed systems, infrastructure reliability, and operational excellence. Focus on designing reliable automation workflows, scaling engineering delivery, managing incidents and SLOs, and developing a high-performing software engineering team.

Location: US

Salary: $300,000–$350,000 USD per year, plus potential bonus and equity.

Company

Nscale provides GPU cloud infrastructure for AI startups and enterprise customers, reducing the complexity of AI development and deployment.

What you will do

  • Hire, onboard, coach, and develop software engineers while managing performance and career growth.
  • Turn the Fleet Manager roadmap into executable plans, manage dependencies and risks, and deliver committed projects.
  • Own planning, prioritization, stakeholder communication, escalation processes, and team operational load.
  • Maintain engineering standards across code review, testing, CI/CD, incident response, postmortems, SLOs, observability, and alerting.
  • Partner with Principal and Staff engineers on architecture, reliability, maintainability, and automation complexity.
  • Collaborate with Product, Infrastructure, Platform, SRE, and UI/UX teams while remaining hands-on with design reviews, incidents, and code.

Requirements

  • 8+ years of software engineering experience building and operating production systems, including 2+ years managing software engineers.
  • Strong Python and distributed systems background, with the ability to contribute to design and code reviews.
  • Experience delivering complex projects from ambiguous requirements through production and handling monitoring, incidents, and performance optimization.
  • Proven experience hiring, coaching, managing performance, and developing engineers toward senior and staff levels.
  • Strong focus on infrastructure reliability, scalability, security, and continuous improvement.
  • Excellent communication skills for building consensus with internal and external stakeholders.

Nice to have

  • Experience with workflow orchestration systems such as Temporal, Airflow, or Prefect.
  • Infrastructure tooling experience with DCIMs, NetBox, OpenStack, ERP systems, MAAS, Ironic, IPMI, PXE boot, or network automation.
  • GPU infrastructure, hardware lifecycle automation, HPC, datacenter networking, InfiniBand, or RoCE experience.
  • Working knowledge of Kubernetes, Terraform, Pulumi, AWS, and GCP.
  • Experience scaling engineering teams through rapid growth.

Culture & Benefits

  • Collaborative, supportive, and innovation-focused engineering environment.
  • Flexible workplace with autonomy over how the workday is structured.
  • Performance reviews every 12 months and a progression plan aligned with career ambitions.
  • Potential bonus, equity, medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.
  • Opportunity to shape global AI capacity planning and deployment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →