7 ΡΠ°ΡΠΎΠ² Π½Π°Π·Π°Π΄
Infrastructure Engineer (AI)
208Β 000 - 269Β 000$
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Π’Π΅ΠΊΡΡ:
TL;DR
Infrastructure Engineer (AI): Building and operating hyperscale compute infrastructure and control planes for AI with an accent on observability, API design, and distributed systems. Focus on automating fleet management, ensuring system reliability, and integrating new hardware generations into a unified production environment.
Location: On-site in San Francisco, CA; Austin, TX; New York, NY; or Seattle, WA
Salary: $208,000 β $269,000 + Equity
Company
is building civilization-scale infrastructure for AI, focusing on rapid deployment of compute power through hardware and software innovation.
What you will do
- Own and build the observability platform to make tens of thousands of GPUs legible in real-time.
- Design and build the API surface for infrastructure to manage and operate the hyperscale fleet.
- Develop the production control plane, including unified machine management and distributed command execution.
- Maintain fleet state as a source of truth across provisioning, operations, and customer-facing platforms.
- Integrate new hardware generations into the platform using ZTP, DHCP, and DNS.
Requirements
- Must be based in or able to work on-site in San Francisco, Austin, New York, or Seattle.
- Experience shipping production services that other teams depend on at scale.
- Proficiency in designing APIs that age well and avoiding leaky abstractions.
- Ability to move toward ambiguity and build maps in unfamiliar domains.
- Fluency with AI tooling (LLM APIs, MCP servers, agentic frameworks) and AI-assisted coding.
- Comfortable carrying a pager and managing incidents end-to-end.
Nice to have
- Experience with distributed systems and data pipeline engineering.
- Knowledge of time-series observability stacks like Prometheus, Thanos, or VictoriaMetrics.
- Familiarity with workflow and orchestration engines such as Temporal or Cadence.
- Experience with BMC/Redfish or hardware telemetry.
- Proficiency in Go, Python, and Postgres.
Culture & Benefits
- Competitive total compensation package including base salary and equity.
- Comprehensive health, dental, and vision insurance.
- Retirement or pension plan.
- Generous PTO policy.
- High-intensity, ownership-driven work environment focused on solving frontier AI infrastructure challenges.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β
ΠΠΎΡ ΠΎΠΆΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Nscale
3 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Infrastructure Software Engineer (AI)
150Β 000 - 215Β 000$
Anthropic
5 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Software Engineer, Infrastructure, Interpretability (AI)
320Β 000 - 485Β 000$
5 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Senior DevOps Engineer (AI)
170Β 000 - 185Β 000$
5 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Developer Experience Engineer (AI/HPC)
150Β 000 - 275Β 000$
5 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Production Engineer (AI Infrastructure)
172Β 000 - 209Β 000$
5 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Senior Site Reliability Engineer (AI Infrastructure)
215Β 000 - 275Β 000$