Назад
Company hidden
2 дня назад

Principal Engineer (AI Infrastructure)

285 000 - 335 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Engineer (AI Infrastructure): Building Conductor, a control plane for managing the full lifecycle and operation of a global fleet of 100,000+ GPUs with an accent on distributed systems, topology modeling, reconciliation, and policy-gated automation. Focus on defining the platform architecture, orchestrating bare-metal and firmware workflows, correlating high-cardinality observability data, and enabling autonomous site operations.

Location: San Francisco, CA, US; on-site

Salary: $285,000–$335,000 annually plus bonus; restricted stock units included.

Company

hirify.global builds vertically integrated AI infrastructure spanning energy, data centers, hardware, and cloud services.

What you will do

  • Define the end-to-end architecture and operating standards for Conductor, hirify.global’s AI infrastructure operating platform.
  • Design twin sources of truth for runtime observability and infrastructure inventory, including a traversable topology graph.
  • Build reconciliation, policy enforcement, workflow orchestration, and lifecycle automation for more than 100,000 GPUs.
  • Orchestrate provisioning, imaging, firmware upgrades, validation, repair, RMA, re-admission, and rollback workflows.
  • Develop unified observability across GPU, networking, storage, orchestration, and workload signals.
  • Lead architecture across hardware, networking, compute, validation, data center, and product teams while mentoring senior engineers.

Requirements

  • 10+ years building infrastructure-layer systems at scale, including fleet management, distributed control planes, provisioning, inventory, or hardware lifecycle automation.
  • Deep experience with distributed systems, state reconciliation, event-driven orchestration, workflow engines, and policy enforcement.
  • Hands-on experience with bare-metal provisioning using PXE, Redfish, or IPMI; firmware and BIOS management; GPU telemetry; and InfiniBand or RoCE fabrics.
  • Experience designing high-cardinality observability or telemetry platforms across compute, network, and storage layers.
  • Strong software engineering fundamentals in Go, Rust, C++, or a similar systems language.
  • Demonstrated technical leadership across multiple teams, including architecture reviews, senior-engineer mentorship, and multi-quarter roadmap delivery.

Nice to have

  • Experience with graph data models for infrastructure topology.
  • Experience with security attestation or SBOM tooling.

Culture & Benefits

  • Competitive compensation, bonus, equity, and restricted stock units.
  • Paid time off, holidays, parental leave, and leave of absence programs.
  • Health, dental, and vision insurance, including employer HSA contributions.
  • 401(k) plan with company matching up to 4% of salary.
  • Professional development, tuition reimbursement, mental health support, commuter benefits, and daily meal allowance.
  • Global travel insurance, emergency assistance, volunteer time off, and location-specific programs.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →