2 ΡΠ°ΡΠ° Π½Π°Π·Π°Π΄
Production Engineer (AI Infrastructure)
172Β 000 - 209Β 000$
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Π’Π΅ΠΊΡΡ:
TL;DR
Production Engineer (AI Infrastructure) (GPU cloud/SRE): Building and operating reliable, scalable infrastructure for AI and HPC workloads with an accent on production reliability, observability, and automation. Focus on designing self-healing systems, resolving complex distributed-system incidents, and improving resilience across compute, networking, storage, and platform services.
Location: On-site in San Francisco or Sunnyvale, California, US
Salary: $172,000β$209,000 annually plus bonus and Restricted Stock Units
Company
builds an energy-efficient, AI-optimized cloud platform for demanding AI and high-performance computing workloads.
What you will do
- Define, measure, and improve availability metrics, service-level indicators, and service-level objectives for the cloud platform.
- Respond to production incidents, diagnose service disruptions, and contribute to post-incident reviews and root-cause analysis.
- Build and improve infrastructure observability with Prometheus, Grafana, Alertmanager, and OpenTelemetry.
- Identify reliability risks, performance bottlenecks, and early indicators of production issues across distributed systems.
- Develop automation, self-healing infrastructure, and operational tooling that reduce toil and improve recovery times.
- Partner with compute, networking, storage, and platform teams to strengthen resilience, disaster recovery, and reliability practices.
Requirements
- 5+ years of experience in Production Engineering, SRE, or large-scale infrastructure operations.
- Experience supporting GPU workloads, HPC environments, or latency- and throughput-sensitive distributed systems.
- Strong Linux/Unix debugging skills across kernel and user space.
- Knowledge of Kubernetes, distributed systems, virtualization, and cloud platforms such as AWS or GCP.
- Experience with monitoring and observability, infrastructure-as-code or configuration management, and scripting or programming in Go, Python, C, or C++.
- Bachelorβs degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience; strong cross-team communication and incident troubleshooting skills.
Nice to have
- Experience operating Kubernetes or container orchestration platforms at scale.
- Experience with automated remediation, event-driven tooling, self-healing systems, change management, or operational readiness reviews.
- Interest in scaling AI or HPC infrastructure and mentoring others in Production Engineering.
Culture & Benefits
- Competitive compensation, bonus, and Restricted Stock Units.
- Health, vision, dental, HSA contributions, life insurance, disability coverage, and Teladoc.
- 401(k) with a 100% employer match up to 4% of salary.
- Paid parental leave, generous paid time off, and holidays.
- Tuition reimbursement, cell phone reimbursement, a $300 monthly commuter benefit, Calm subscription, and MetLife Legal.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β
ΠΠΎΡ ΠΎΠΆΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
1 ΡΠ°Ρ Π½Π°Π·Π°Π΄
Cluster Software Engineer (AI)
4 ΡΠ°ΡΠ° Π½Π°Π·Π°Π΄
Forward Deployed Engineers (AI Infrastructure)
180Β 000 - 240Β 000$
2 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Senior AI DevOps Developer (AI)
150Β 000 - 206Β 000$
4 ΡΠ°ΡΠ° Π½Π°Π·Π°Π΄
Senior DevOps Engineer (AI)
170Β 000 - 185Β 000$
2 ΡΠ°ΡΠ° Π½Π°Π·Π°Π΄
Software Engineer (AI Infrastructure)
4 ΡΠ°ΡΠ° Π½Π°Π·Π°Π΄
Senior Site Reliability Engineer (AI Infrastructure)
215Β 000 - 275Β 000$