Назад
Company hidden
5 часов Π½Π°Π·Π°Π΄

Lead Site Reliability Engineer (Kubernetes)

Π€ΠΎΡ€ΠΌΠ°Ρ‚ Ρ€Π°Π±ΠΎΡ‚Ρ‹
remote (Ρ‚ΠΎΠ»ΡŒΠΊΠΎ USA)
Π’ΠΈΠΏ Ρ€Π°Π±ΠΎΡ‚Ρ‹
fulltime
Π“Ρ€Π΅ΠΉΠ΄
lead
Английский
b2
Π‘Ρ‚Ρ€Π°Π½Π°
US
Вакансия ΠΈΠ· списка Hirify.GlobalВакансия ΠΈΠ· Hirify Global, списка ΠΌΠ΅ΠΆΠ΄ΡƒΠ½Π°Ρ€ΠΎΠ΄Π½Ρ‹Ρ… tech-ΠΊΠΎΠΌΠΏΠ°Π½ΠΈΠΉ
Для мэтча ΠΈ ΠΎΡ‚ΠΊΠ»ΠΈΠΊΠ° Π½ΡƒΠΆΠ΅Π½ Plus

ΠœΡΡ‚Ρ‡ & Π‘ΠΎΠΏΡ€ΠΎΠ²ΠΎΠ΄

Для мэтча с этой вакансиСй Π½ΡƒΠΆΠ΅Π½ Plus

ОписаниС вакансии

ВСкст:
/
TL;DR
Lead Site Reliability Engineer (Kubernetes/Cloud Infrastructure): Modernizing and operating Intellum’s highly available SaaS platform across multiple cloud providers with an accent on container orchestration, infrastructure as code, deployment systems, and observability. Focus on designing portable infrastructure, solving complex distributed-system failures, improving SLI/SLO practices, and leading incident response and technical mentorship.

Location: Remote, United States; collaboration across US and European time zones and participation in an on-call rotation are required.

Company

hirify.global provides corporate education technology for customer, partner, and employee learning programs.

What you will do

  • Lead infrastructure modernization from legacy compute environments to portable, container-orchestrated infrastructure.
  • Design and maintain infrastructure as code across multiple cloud providers, with a focus on portability, maintainability, and scalability.
  • Improve CI/CD systems, deployment tooling, observability, monitoring, alerting, and load-testing capabilities.
  • Establish and evolve SLI and SLO practices and lead incident troubleshooting, root cause analysis, and corrective actions.
  • Provide architecture guidance, mentorship, and technical direction across the Systems Engineering function.
  • Partner with Security and Engineering on access controls, infrastructure hardening, compliance, cost management, and developer experience.

Requirements

  • 8+ years of hands-on experience in infrastructure, DevOps, platform engineering, site reliability engineering, or a related discipline.
  • Experience designing, operating, and troubleshooting highly available production infrastructure across more than one major cloud provider, with depth in AWS or Google Cloud and working fluency in the other.
  • Significant production experience with Kubernetes and container orchestration, including cluster operations, workload configuration, reliability, and troubleshooting.
  • Experience modernizing VM-based or legacy infrastructure toward containerized or cloud-native architectures, plus strong Terraform or comparable infrastructure-as-code experience.
  • Experience operating CI/CD and deployment infrastructure, responding to incidents, diagnosing distributed-system failures, and conducting post-incident reviews.
  • Strong Linux administration, scripting or programming in Ruby, Python, or a comparable language, communication skills, and experience mentoring engineers.

Nice to have

  • Experience operating AWS and Google Cloud simultaneously.
  • People leadership, technical leadership, or player-coach experience.
  • Cloud cost management or FinOps experience at scale.
  • Experience with Spinnaker, Jenkins, Ruby on Rails, SOC 2, AI-assisted development tooling, or learning management systems.

Culture & Benefits

  • Remote-first work environment with distributed team members.
  • Medical, dental, and vision coverage with 100% of employee premiums covered for selected individual plans.
  • 401(k) with matching for US-based employees.
  • Flexible PTO, Calm subscription, and LinkedIn Learning.
  • Personal development budgets and an annual company retreat.

Π‘ΡƒΠ΄ΡŒΡ‚Π΅ остороТны: Ссли Ρ€Π°Π±ΠΎΡ‚ΠΎΠ΄Π°Ρ‚Π΅Π»ΡŒ просит Π²ΠΎΠΉΡ‚ΠΈ Π² ΠΈΡ… систСму, ΠΈΡΠΏΠΎΠ»ΡŒΠ·ΡƒΡ iCloud/Google, ΠΏΡ€ΠΈΡΠ»Π°Ρ‚ΡŒ ΠΊΠΎΠ΄/ΠΏΠ°Ρ€ΠΎΠ»ΡŒ, Π·Π°ΠΏΡƒΡΡ‚ΠΈΡ‚ΡŒ ΠΊΠΎΠ΄/ПО, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡ‚Π΅ этого - это мошСнники. ΠžΠ±ΡΠ·Π°Ρ‚Π΅Π»ΡŒΠ½ΠΎ ΠΆΠΌΠΈΡ‚Π΅ "ΠŸΠΎΠΆΠ°Π»ΠΎΠ²Π°Ρ‚ΡŒΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡˆΠΈΡ‚Π΅ Π² ΠΏΠΎΠ΄Π΄Π΅Ρ€ΠΆΠΊΡƒ. ΠŸΠΎΠ΄Ρ€ΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β†’