Назад
ΠΎΠ±Π½ΠΎΠ²Π»Π΅Π½ΠΎ 11 часов Π½Π°Π·Π°Π΄

Site Reliability Engineer (AI)

Π€ΠΎΡ€ΠΌΠ°Ρ‚ Ρ€Π°Π±ΠΎΡ‚Ρ‹
onsite
Π’ΠΈΠΏ Ρ€Π°Π±ΠΎΡ‚Ρ‹
fulltime
Английский
b2
Π‘Ρ‚Ρ€Π°Π½Π°
US
Вакансия ΠΈΠ· списка Hirify.GlobalВакансия ΠΈΠ· Hirify Global, списка ΠΌΠ΅ΠΆΠ΄ΡƒΠ½Π°Ρ€ΠΎΠ΄Π½Ρ‹Ρ… tech-ΠΊΠΎΠΌΠΏΠ°Π½ΠΈΠΉ
Для мэтча ΠΈ ΠΎΡ‚ΠΊΠ»ΠΈΠΊΠ° Π½ΡƒΠΆΠ΅Π½ Plus

ΠœΡΡ‚Ρ‡ & Π‘ΠΎΠΏΡ€ΠΎΠ²ΠΎΠ΄

Для мэтча с этой вакансиСй Π½ΡƒΠΆΠ΅Π½ Plus

ОписаниС вакансии

ВСкст:
/

TL;DR

Site Reliability Engineer (AI): Own production reliability and platform engineering for user-facing AI products Devin and Windsurf with an accent on SLOs, incident response, and CI/CD pipelines. Focus on building monitoring and observability systems, automating toil reduction, and ensuring infrastructure scales with hundreds of thousands of daily users.

Location: On-site in San Francisco Bay Area

Company

Applied AI lab building end-to-end software agents like Devin, the first AI software engineer, and Windsurf, an AI-native IDE.

What you will do

  • Define and own SLOs, SLIs, error budgets, monitoring, alerting, and observability for Devin and Windsurf.
  • Lead incident response, run blameless postmortems, and build runbooks and tooling for sustainable on-call.
  • Own deployment pipelines, release infrastructure, CI/CD, and internal developer tooling to enable fast shipping.
  • Manage cloud infrastructure as code with reproducible, version-controlled environments.
  • Perform capacity planning, performance profiling, and growth modeling.
  • Integrate security into reliability practices and foster reliability culture across product and engineering teams.

Requirements

  • Deep experience running production systems at scale: SLOs, error budgets, on-call rotations, incident command.
  • Strong software engineering fundamentals; write real code.
  • Proficiency with cloud infrastructure (AWS, GCP, or Azure), Kubernetes, Terraform or equivalent IaC.
  • Experience building and owning CI/CD pipelines and deployment infrastructure.
  • Strong observability skills: instrumentation, dashboards, effective alerting.
  • Track record of systematic toil reduction through automation.
  • Comfort owning incidents end-to-end and product empathy for user-facing reliability.

Nice to have

  • Experience with developer-facing products or platforms.

Culture & Benefits

  • Small, talent-dense team of competitive programmers, founders, and AI researchers from Scale AI, Palantir, Cursor, Google DeepMind.
  • High ownership and trust: set your own reliability standards.
  • Proactive, systematic environment treating reliability as a craft.
  • Ship products used by hundreds of thousands of developers daily.

Π‘ΡƒΠ΄ΡŒΡ‚Π΅ остороТны: Ссли Ρ€Π°Π±ΠΎΡ‚ΠΎΠ΄Π°Ρ‚Π΅Π»ΡŒ просит Π²ΠΎΠΉΡ‚ΠΈ Π² ΠΈΡ… систСму, ΠΈΡΠΏΠΎΠ»ΡŒΠ·ΡƒΡ iCloud/Google, ΠΏΡ€ΠΈΡΠ»Π°Ρ‚ΡŒ ΠΊΠΎΠ΄/ΠΏΠ°Ρ€ΠΎΠ»ΡŒ, Π·Π°ΠΏΡƒΡΡ‚ΠΈΡ‚ΡŒ ΠΊΠΎΠ΄/ПО, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡ‚Π΅ этого - это мошСнники. ΠžΠ±ΡΠ·Π°Ρ‚Π΅Π»ΡŒΠ½ΠΎ ΠΆΠΌΠΈΡ‚Π΅ "ΠŸΠΎΠΆΠ°Π»ΠΎΠ²Π°Ρ‚ΡŒΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡˆΠΈΡ‚Π΅ Π² ΠΏΠΎΠ΄Π΄Π΅Ρ€ΠΆΠΊΡƒ. ΠŸΠΎΠ΄Ρ€ΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β†’