2 дня назад
Principal Engineer (AI Infrastructure)
285 000 - 335 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Engineer (AI Infrastructure): Building Conductor, a control plane for managing the full lifecycle and operation of a global fleet of 100,000+ GPUs with an accent on distributed systems, topology modeling, reconciliation, and policy-gated automation. Focus on defining the platform architecture, orchestrating bare-metal and firmware workflows, correlating high-cardinality observability data, and enabling autonomous site operations.
Location: San Francisco, CA, US; on-site
Salary: $285,000–$335,000 annually plus bonus; restricted stock units included.
Company
builds vertically integrated AI infrastructure spanning energy, data centers, hardware, and cloud services.
What you will do
- Define the end-to-end architecture and operating standards for Conductor, ’s AI infrastructure operating platform.
- Design twin sources of truth for runtime observability and infrastructure inventory, including a traversable topology graph.
- Build reconciliation, policy enforcement, workflow orchestration, and lifecycle automation for more than 100,000 GPUs.
- Orchestrate provisioning, imaging, firmware upgrades, validation, repair, RMA, re-admission, and rollback workflows.
- Develop unified observability across GPU, networking, storage, orchestration, and workload signals.
- Lead architecture across hardware, networking, compute, validation, data center, and product teams while mentoring senior engineers.
Requirements
- 10+ years building infrastructure-layer systems at scale, including fleet management, distributed control planes, provisioning, inventory, or hardware lifecycle automation.
- Deep experience with distributed systems, state reconciliation, event-driven orchestration, workflow engines, and policy enforcement.
- Hands-on experience with bare-metal provisioning using PXE, Redfish, or IPMI; firmware and BIOS management; GPU telemetry; and InfiniBand or RoCE fabrics.
- Experience designing high-cardinality observability or telemetry platforms across compute, network, and storage layers.
- Strong software engineering fundamentals in Go, Rust, C++, or a similar systems language.
- Demonstrated technical leadership across multiple teams, including architecture reviews, senior-engineer mentorship, and multi-quarter roadmap delivery.
Nice to have
- Experience with graph data models for infrastructure topology.
- Experience with security attestation or SBOM tooling.
Culture & Benefits
- Competitive compensation, bonus, equity, and restricted stock units.
- Paid time off, holidays, parental leave, and leave of absence programs.
- Health, dental, and vision insurance, including employer HSA contributions.
- 401(k) plan with company matching up to 4% of salary.
- Professional development, tuition reimbursement, mental health support, commuter benefits, and daily meal allowance.
- Global travel insurance, emergency assistance, volunteer time off, and location-specific programs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Principal Engineer (Kubernetes)
174 051 - 227 233$
CrowdStrike
18 часов назад
Principal Software Engineer (Sensor, Telemetry & Observability)
195 000 - 290 000$
3 дня назад
Principal Architect, Simulation Platform (AI)
200 000 - 240 000$
CoreWeave
15 часов назад
Senior Software Engineer - AI Infrastructure Performance Insights & Observability
182 000 - 242 000$
2 часа назад
Principal Engineer - Software Systems
150 000 - 200 000$
2 дня назад
Enterprise Product Engineer (AI)
180 000 - 500 000$