Site Reliability Engineer II (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Site Reliability Engineer II (Kubernetes): Building and operating reliable cloud gaming infrastructure with an accent on production ownership, API gateways, service mesh technologies, and scalable deployments. Focus on troubleshooting distributed systems, tuning traffic management and performance, and leading incident response for high-scale gaming services.
Location: Aliso Viejo, California, United States; hybrid working policy applies
Salary: $145,700–$218,500 USD base pay annually, with potential bonus eligibility
Company
A global interactive entertainment technology organization delivering hardware, network services, and cloud gaming experiences to more than 100 million people.
What you will do
- Own production services, operational readiness, reliability, and stability throughout the software development lifecycle.
- Operate and troubleshoot Kong API Gateway or similar enterprise gateways, including routing, plugins, TLS certificates, authentication, authorization, and latency performance.
- Manage service mesh infrastructure and resolve connectivity, latency, traffic-management, and policy issues.
- Implement safe upgrades, configuration changes, rollouts, rollbacks, monitoring, and infrastructure-as-code or GitOps practices.
- Participate in on-call rotations and lead or support incident response and post-incident reviews.
- Collaborate with engineers and stakeholders while providing operational feedback and improving production code quality.
Requirements
- At least 5 years of experience in software development and/or Linux systems administration in production environments.
- Bachelor’s degree in Computer Science or equivalent professional experience.
- Hands-on experience with Kong API Gateway or comparable enterprise API gateway technology.
- Experience with service mesh technologies such as Istio, Linkerd, or Kuma, including mTLS and zero-trust networking.
- Strong understanding of routing, retries, timeouts, circuit breaking, canary releases, and blue/green traffic shifting.
- Programming experience with Python, Bash, Go, Java, C++, or Rust; willingness to participate in an on-call rotation is required.
Nice to have
- Experience with Ceph, Rook, MongoDB, Redis, Cassandra, Elasticsearch, Kafka, PostgreSQL, or MySQL at scale.
- Experience with Prometheus, Grafana, Kubernetes, Rancher, release engineering, package distribution, performance analysis, or load testing.
- QA or SDET experience.
Culture & Benefits
- Work on cloud gaming and large-scale interactive entertainment services.
- Medical, dental, and vision coverage.
- 401(k) matching, paid time off, wellness program, and employee discounts.
- Potential eligibility for a bonus package.
- Inclusive workplace with equal-opportunity and fair-chance employment practices.
Hiring process
- Background checks are conducted at the offer stage and may include criminal background checks for some roles.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →