Назад
Company hidden
4 дня назад

System Software Engineer — Node & Cluster Management (AI)

200 000 - 350 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
System Software Engineer — Node & Cluster Management (AI) (Linux/C/Go/Rust): Building the management and observability layer for high-performance AI accelerator systems across nodes, racks, and clusters with an accent on infrastructure APIs, hardware telemetry, Linux daemons, and BMC interfaces. Focus on designing failover and recovery mechanisms, debugging across kernel and firmware boundaries, and developing provisioning, fleet management, and device recovery tooling.

Location: Bay Area, California, United States

Salary: $200,000–$350,000 per year

Company

Early-stage AI hardware company building high-performance accelerator systems for large-scale AI workloads.

What you will do

  • Design and build node-level management planes for AI accelerator systems.
  • Expose health, inventory, telemetry, diagnostics, and control through HTTP/REST APIs.
  • Develop cluster management, failover, recovery, availability, and fleet-wide health aggregation.
  • Build Linux management and telemetry daemons, CLI tools, provisioning workflows, and test automation.
  • Integrate host-side software with BMC and out-of-band interfaces, device software, and firmware.
  • Debug issues across APIs, userspace daemons, kernel drivers, firmware, and hardware.

Requirements

  • 8+ years of experience in systems, infrastructure, platform, or hardware-management software.
  • Strong Linux systems programming experience and experience developing low-level userspace software and Linux daemons.
  • Strong C skills plus experience with Go, Rust, C++, and/or Python.
  • Experience building REST APIs and CLI tooling for hardware or infrastructure systems.
  • Understanding of device drivers, PCIe devices, hardware telemetry, firmware interfaces, and BMC-managed subsystems.
  • Ability to collaborate with firmware and hardware engineers in an early-stage environment with evolving specifications.

Nice to have

  • Experience with Redfish, OpenBMC, IPMI, or gNMI.
  • GPU, accelerator, HPC, datacenter fleet, or large-scale Linux infrastructure experience.
  • Experience with secure boot, device attestation, firmware recovery, hardware bring-up, or lab infrastructure.

Culture & Benefits

  • Hands-on work in an early-stage AI hardware environment.
  • Direct collaboration across software, firmware, and hardware engineering.
  • Opportunity to define management interfaces for a new AI platform from individual servers through full-scale clusters.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →