Senior Site Reliability Engineer (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Location: On-site in El Segundo, California, United States. Occasional travel to customer sites and other locations is required, with extended hours and weekend work as needed. U.S. work authorization and U.S. person status are required under ITAR restrictions.
Salary: $153,000–$185,000 per year, plus potential stock options and/or long-term cash awards.
Company
develops commercial space infrastructure, including in-orbit pharmaceutical processing payloads, reentry capsules, heat shields, and satellite buses.
What you will do
- Deploy, maintain, and operate mission-critical applications and infrastructure supporting spacecraft and company-wide systems.
- Build Infrastructure as Code frameworks with Terraform and design infrastructure for cloud services and embedded spacecraft software.
- Implement observability systems, including metrics, logging, tracing, and actionable alerting.
- Build and maintain CI/CD pipelines and GitOps-based automation for safe, repeatable deployments.
- Identify bottlenecks and reliability risks, tune system performance, and improve scalability and resilience.
- Respond to production incidents, perform root-cause analysis, lead blameless postmortems, and participate in on-call rotations.
Requirements
- Bachelor’s degree in computer science, engineering, or a related STEM field with 5+ years of SRE experience, or 7+ years of progressive DevOps, SRE, or systems engineering experience in lieu of a degree.
- Production experience operating Kubernetes or similar container orchestration platforms.
- Experience with Infrastructure as Code, including Terraform, server provisioning, and configuration management.
- Experience with Prometheus, Grafana, InfluxDB, or similar observability technologies.
- Knowledge of software-defined networking and scripting with Python, Bash, PowerShell, or similar tools.
- Must be authorized to work in the United States and qualify as a U.S. person for access to export-controlled technology.
Nice to have
- Experience managing scalable Azure infrastructure and implementing CI/CD, GitOps, Ansible, Salt, or ArgoCD solutions.
- Strong Linux and container-runtime expertise, including Docker or containerd.
- Experience with GPU workloads, high-throughput computing, or HPC environments using schedulers such as Slurm.
- Experience with hybrid cloud, on-premises, or edge environments and debugging distributed systems at scale.
- Experience with databases and data modeling.
Culture & Benefits
- Full-time employees receive flexible PTO and 12 paid holidays.
- Company-paid medical, dental, and vision insurance, with FSA and employer-matched HSA options.
- Parental leave, family-building benefits, wellness reimbursement, and sponsored One Medical memberships.
- 401(k) plan with a 6% immediately vested employer match and equity incentives.
- Relocation support may be available for new hires, along with daily lunch, twice-weekly dinner, team events, and EV charging.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →