Назад
Company hidden
2 часа Π½Π°Π·Π°Π΄

Production Engineer (AI Infrastructure)

172Β 000 - 209Β 000$
Π€ΠΎΡ€ΠΌΠ°Ρ‚ Ρ€Π°Π±ΠΎΡ‚Ρ‹
onsite
Π’ΠΈΠΏ Ρ€Π°Π±ΠΎΡ‚Ρ‹
fulltime
Π“Ρ€Π΅ΠΉΠ΄
senior
Английский
b2
Π‘Ρ‚Ρ€Π°Π½Π°
US
Вакансия ΠΈΠ· списка Hirify.GlobalВакансия ΠΈΠ· Hirify Global, списка ΠΌΠ΅ΠΆΠ΄ΡƒΠ½Π°Ρ€ΠΎΠ΄Π½Ρ‹Ρ… tech-ΠΊΠΎΠΌΠΏΠ°Π½ΠΈΠΉ
Для мэтча ΠΈ ΠΎΡ‚ΠΊΠ»ΠΈΠΊΠ° Π½ΡƒΠΆΠ΅Π½ Plus

ΠœΡΡ‚Ρ‡ & Π‘ΠΎΠΏΡ€ΠΎΠ²ΠΎΠ΄

Для мэтча с этой вакансиСй Π½ΡƒΠΆΠ΅Π½ Plus

ОписаниС вакансии

ВСкст:
/
TL;DR
Production Engineer (AI Infrastructure) (GPU cloud/SRE): Building and operating reliable, scalable infrastructure for AI and HPC workloads with an accent on production reliability, observability, and automation. Focus on designing self-healing systems, resolving complex distributed-system incidents, and improving resilience across compute, networking, storage, and platform services.

Location: On-site in San Francisco or Sunnyvale, California, US

Salary: $172,000–$209,000 annually plus bonus and Restricted Stock Units

Company

hirify.global builds an energy-efficient, AI-optimized cloud platform for demanding AI and high-performance computing workloads.

What you will do

  • Define, measure, and improve availability metrics, service-level indicators, and service-level objectives for the cloud platform.
  • Respond to production incidents, diagnose service disruptions, and contribute to post-incident reviews and root-cause analysis.
  • Build and improve infrastructure observability with Prometheus, Grafana, Alertmanager, and OpenTelemetry.
  • Identify reliability risks, performance bottlenecks, and early indicators of production issues across distributed systems.
  • Develop automation, self-healing infrastructure, and operational tooling that reduce toil and improve recovery times.
  • Partner with compute, networking, storage, and platform teams to strengthen resilience, disaster recovery, and reliability practices.

Requirements

  • 5+ years of experience in Production Engineering, SRE, or large-scale infrastructure operations.
  • Experience supporting GPU workloads, HPC environments, or latency- and throughput-sensitive distributed systems.
  • Strong Linux/Unix debugging skills across kernel and user space.
  • Knowledge of Kubernetes, distributed systems, virtualization, and cloud platforms such as AWS or GCP.
  • Experience with monitoring and observability, infrastructure-as-code or configuration management, and scripting or programming in Go, Python, C, or C++.
  • Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience; strong cross-team communication and incident troubleshooting skills.

Nice to have

  • Experience operating Kubernetes or container orchestration platforms at scale.
  • Experience with automated remediation, event-driven tooling, self-healing systems, change management, or operational readiness reviews.
  • Interest in scaling AI or HPC infrastructure and mentoring others in Production Engineering.

Culture & Benefits

  • Competitive compensation, bonus, and Restricted Stock Units.
  • Health, vision, dental, HSA contributions, life insurance, disability coverage, and Teladoc.
  • 401(k) with a 100% employer match up to 4% of salary.
  • Paid parental leave, generous paid time off, and holidays.
  • Tuition reimbursement, cell phone reimbursement, a $300 monthly commuter benefit, Calm subscription, and MetLife Legal.

Π‘ΡƒΠ΄ΡŒΡ‚Π΅ остороТны: Ссли Ρ€Π°Π±ΠΎΡ‚ΠΎΠ΄Π°Ρ‚Π΅Π»ΡŒ просит Π²ΠΎΠΉΡ‚ΠΈ Π² ΠΈΡ… систСму, ΠΈΡΠΏΠΎΠ»ΡŒΠ·ΡƒΡ iCloud/Google, ΠΏΡ€ΠΈΡΠ»Π°Ρ‚ΡŒ ΠΊΠΎΠ΄/ΠΏΠ°Ρ€ΠΎΠ»ΡŒ, Π·Π°ΠΏΡƒΡΡ‚ΠΈΡ‚ΡŒ ΠΊΠΎΠ΄/ПО, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡ‚Π΅ этого - это мошСнники. ΠžΠ±ΡΠ·Π°Ρ‚Π΅Π»ΡŒΠ½ΠΎ ΠΆΠΌΠΈΡ‚Π΅ "ΠŸΠΎΠΆΠ°Π»ΠΎΠ²Π°Ρ‚ΡŒΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡˆΠΈΡ‚Π΅ Π² ΠΏΠΎΠ΄Π΄Π΅Ρ€ΠΆΠΊΡƒ. ΠŸΠΎΠ΄Ρ€ΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β†’