5 ΡΠ°ΡΠΎΠ² Π½Π°Π·Π°Π΄
Lead Site Reliability Engineer (Kubernetes)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Π’Π΅ΠΊΡΡ:
TL;DR
Lead Site Reliability Engineer (Kubernetes/Cloud Infrastructure): Modernizing and operating Intellumβs highly available SaaS platform across multiple cloud providers with an accent on container orchestration, infrastructure as code, deployment systems, and observability. Focus on designing portable infrastructure, solving complex distributed-system failures, improving SLI/SLO practices, and leading incident response and technical mentorship.
Location: Remote, United States; collaboration across US and European time zones and participation in an on-call rotation are required.
Company
provides corporate education technology for customer, partner, and employee learning programs.
What you will do
- Lead infrastructure modernization from legacy compute environments to portable, container-orchestrated infrastructure.
- Design and maintain infrastructure as code across multiple cloud providers, with a focus on portability, maintainability, and scalability.
- Improve CI/CD systems, deployment tooling, observability, monitoring, alerting, and load-testing capabilities.
- Establish and evolve SLI and SLO practices and lead incident troubleshooting, root cause analysis, and corrective actions.
- Provide architecture guidance, mentorship, and technical direction across the Systems Engineering function.
- Partner with Security and Engineering on access controls, infrastructure hardening, compliance, cost management, and developer experience.
Requirements
- 8+ years of hands-on experience in infrastructure, DevOps, platform engineering, site reliability engineering, or a related discipline.
- Experience designing, operating, and troubleshooting highly available production infrastructure across more than one major cloud provider, with depth in AWS or Google Cloud and working fluency in the other.
- Significant production experience with Kubernetes and container orchestration, including cluster operations, workload configuration, reliability, and troubleshooting.
- Experience modernizing VM-based or legacy infrastructure toward containerized or cloud-native architectures, plus strong Terraform or comparable infrastructure-as-code experience.
- Experience operating CI/CD and deployment infrastructure, responding to incidents, diagnosing distributed-system failures, and conducting post-incident reviews.
- Strong Linux administration, scripting or programming in Ruby, Python, or a comparable language, communication skills, and experience mentoring engineers.
Nice to have
- Experience operating AWS and Google Cloud simultaneously.
- People leadership, technical leadership, or player-coach experience.
- Cloud cost management or FinOps experience at scale.
- Experience with Spinnaker, Jenkins, Ruby on Rails, SOC 2, AI-assisted development tooling, or learning management systems.
Culture & Benefits
- Remote-first work environment with distributed team members.
- Medical, dental, and vision coverage with 100% of employee premiums covered for selected individual plans.
- 401(k) with matching for US-based employees.
- Flexible PTO, Calm subscription, and LinkedIn Learning.
- Personal development budgets and an annual company retreat.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β