Назад
2 дня назад

Software Engineer - Compute Infra / HPC

119 800 - 304 200$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer - Compute Infra / HPC (AI compute infrastructure): Building distributed services, control planes, and automation for bringing large-scale GPU clusters into production and maintaining their health with an accent on Kubernetes, fleet lifecycle management, and hardware-aware operations. Focus on designing resilient workflows, diagnosing failures across software and hardware boundaries, and automating certification, maintenance, recovery, and safe rollouts at hyperscale.

Location: Mountain View, United States. Employees living within 50 miles of a U.S. Microsoft office are expected to work from a designated office at least four days per week.

Base pay: USD $119,800–$234,700 per year for Software Engineering IC4, or USD $142,800–$274,800 per year for IC5. San Francisco Bay Area and New York City ranges are USD $160,200–$261,000 for IC4 and USD $188,000–$304,200 for IC5.

Company

Microsoft AI develops AI systems intended to advance science, education, productivity, and global well-being.

What you will do

  • Design and build distributed services, control planes, APIs, and workflow engines for the full lifecycle of AI compute clusters.
  • Automate cluster bootstrap, Kubernetes control planes, networking, identity, image distribution, secrets, and infrastructure-as-code across Azure and partner clouds.
  • Develop rack and node qualification, topology validation, scale testing, and capacity transition workflows for large GPU fleets.
  • Build health, diagnostics, telemetry, lifecycle, and remediation systems for hardware, hosts, networks, and storage.
  • Implement policy-driven automation for certification, maintenance, safe rollouts, rollbacks, and recovery from failures.
  • Lead architecture decisions, design reviews, mentoring, and complex initiatives with research, hardware, networking, storage, security, and datacenter teams.

Requirements

  • Bachelor’s degree in computer science, computer engineering, electrical engineering, or a related field and 4+ years of software engineering experience, or equivalent experience.
  • 4+ years designing scalable software for cloud, datacenter, cluster, or fleet infrastructure using languages such as Go, Rust, C++, C#, Java, or Python.
  • Experience with Kubernetes or comparable orchestration systems, Linux, public-cloud infrastructure, and infrastructure-as-code or declarative configuration.
  • Ability to debug complex behavior across services, control planes, operating systems, networking, and hardware.
  • Experience leading projects across multiple teams and communicating technical tradeoffs clearly.

Nice to have

  • Hyperscale compute infrastructure experience with thousands of nodes, multiple clusters, heterogeneous accelerators, or multiple cloud and datacenter providers.
  • Depth in Kubernetes internals, cluster and node lifecycle systems, cloud foundations, durable workflows, state machines, telemetry pipelines, and automated remediation.
  • Low-level systems experience with Linux kernels, device drivers, firmware, BMC or Redfish, secure boot, attestation, virtualization, or hardware-health agents.
  • Experience with GPU systems, InfiniBand, RoCE, RDMA, NVLink, NCCL, high-performance storage, or data movement.
  • Experience owning production reliability, incident response, mentoring, and multi-quarter infrastructure initiatives.

Culture & Benefits

  • Hands-on work at the intersection of distributed systems and AI supercomputing.
  • Collaboration with research, hardware health, networking, storage, security, and datacenter specialists.
  • Potential eligibility for benefits and additional compensation.
  • Equal employment opportunity and reasonable accommodation support during the application process.

Hiring process

  • Applications are accepted on an ongoing basis until the position is filled, with the posting open for at least five days.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →