обновлено 10 дней назад
Staff Site Reliability Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Site Reliability Engineer (Kubernetes/AWS): Operating and evolving the reliability, performance, and availability of Develocity SaaS instances and supporting production services with an accent on incident response, observability, automation, and disaster recovery. Focus on defining SRE operating standards, leading complex incidents, improving cloud platform scalability, and mentoring a new SRE team.
Location: Remote, must be located in the GMT timezone (UK, Ireland, Portugal, or the Canary Islands)
Company
builds Develocity, an AI-powered toolchain observability and intelligence platform for software organizations, supporting build and test acceleration across multiple development ecosystems.
What you will do
- Operate and maintain Develocity instances, artifact registries, and supporting production services.
- Define SRE standards and operating models covering on-call, incident response, postmortems, SLOs, and error budgets.
- Lead technical escalation and blameless retrospectives for complex or high-severity incidents.
- Drive automation for deployments, upgrades, monitoring, self-healing, recovery, and operational workflows.
- Build observability across managed services, including logging, metrics, tracing, and alerting.
- Own disaster recovery, backups, business continuity, performance, cost optimization, and reliability improvements.
Requirements
- 7+ years of experience in SRE, DevOps, or an equivalent production operations role.
- Experience leading reliability initiatives across multiple teams or services and influencing technical direction without direct authority.
- Strong production Kubernetes experience and cloud infrastructure expertise, preferably with AWS, EKS, RDS, S3, and EC2.
- Proficiency with Prometheus, Grafana, Terraform, Python, and Bash.
- Experience with SLOs, error budgets, incident management, and 24/7 on-call operations.
- Strong written and verbal English communication skills.
Nice to have
- Experience establishing SRE practices in a growing SaaS organization.
- Familiarity with Develocity or JVM languages such as Java and Kotlin.
- Experience with customer-facing and executive-level incident communications.
Culture & Benefits
- Remote-first environment with asynchronous communication and written documentation.
- Hands-on ownership of production systems used by major software organizations and open-source projects.
- Culture focused on automation rather than heroics.
- In-person annual offsites and team meetings.
- Competitive salary and equity grants.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
18 минут назад
DevOps/SRE Engineer
1 день назад
Senior Site Reliability Engineer (Satellite Operations)
142 800 - 178 500$
10 дней назад
Staff Site Reliability Engineer (AI)
252 000 - 308 000$
12 дней назад
Site Reliability Engineer (AWS/Kubernetes)
10 дней назад
Engineering Manager, SRE (AI)
11 дней назад