Software Engineer - Compute Infra / HPC
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Location: Mountain View, United States. Employees living within 50 miles of a U.S. Microsoft office are expected to work from a designated office at least four days per week.
Base pay: USD $119,800–$234,700 per year for Software Engineering IC4, or USD $142,800–$274,800 per year for IC5. San Francisco Bay Area and New York City ranges are USD $160,200–$261,000 for IC4 and USD $188,000–$304,200 for IC5.
Company
Microsoft AI develops AI systems intended to advance science, education, productivity, and global well-being.
What you will do
- Design and build distributed services, control planes, APIs, and workflow engines for the full lifecycle of AI compute clusters.
- Automate cluster bootstrap, Kubernetes control planes, networking, identity, image distribution, secrets, and infrastructure-as-code across Azure and partner clouds.
- Develop rack and node qualification, topology validation, scale testing, and capacity transition workflows for large GPU fleets.
- Build health, diagnostics, telemetry, lifecycle, and remediation systems for hardware, hosts, networks, and storage.
- Implement policy-driven automation for certification, maintenance, safe rollouts, rollbacks, and recovery from failures.
- Lead architecture decisions, design reviews, mentoring, and complex initiatives with research, hardware, networking, storage, security, and datacenter teams.
Requirements
- Bachelor’s degree in computer science, computer engineering, electrical engineering, or a related field and 4+ years of software engineering experience, or equivalent experience.
- 4+ years designing scalable software for cloud, datacenter, cluster, or fleet infrastructure using languages such as Go, Rust, C++, C#, Java, or Python.
- Experience with Kubernetes or comparable orchestration systems, Linux, public-cloud infrastructure, and infrastructure-as-code or declarative configuration.
- Ability to debug complex behavior across services, control planes, operating systems, networking, and hardware.
- Experience leading projects across multiple teams and communicating technical tradeoffs clearly.
Nice to have
- Hyperscale compute infrastructure experience with thousands of nodes, multiple clusters, heterogeneous accelerators, or multiple cloud and datacenter providers.
- Depth in Kubernetes internals, cluster and node lifecycle systems, cloud foundations, durable workflows, state machines, telemetry pipelines, and automated remediation.
- Low-level systems experience with Linux kernels, device drivers, firmware, BMC or Redfish, secure boot, attestation, virtualization, or hardware-health agents.
- Experience with GPU systems, InfiniBand, RoCE, RDMA, NVLink, NCCL, high-performance storage, or data movement.
- Experience owning production reliability, incident response, mentoring, and multi-quarter infrastructure initiatives.
Culture & Benefits
- Hands-on work at the intersection of distributed systems and AI supercomputing.
- Collaboration with research, hardware health, networking, storage, security, and datacenter specialists.
- Potential eligibility for benefits and additional compensation.
- Equal employment opportunity and reasonable accommodation support during the application process.
Hiring process
- Applications are accepted on an ongoing basis until the position is filled, with the posting open for at least five days.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →