4 дня назад
Engineering Manager, Edge SRE
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Engineering Manager, Edge SRE (SRE/Distributed Systems): Leading the London Edge SRE team responsible for Cloudflare edge production reliability and tooling with an accent on scalability, observability, and incident response. Focus on building systems and APIs for production debugging, measuring availability and performance, coordinating reliability initiatives, and developing engineers.
Location: Hybrid in London, United Kingdom. Candidates progressing to the offer stage may be asked to attend an in-person interview at a office or hub.
Company
operates a global network that protects and accelerates Internet applications, serving customers from individual bloggers to large enterprises.
What you will do
- Lead and develop the Edge SRE team responsible for keeping edge production reliable and scalable.
- Build tools and practices that help engineering teams debug production systems, measure availability and performance, and track operational thresholds.
- Prioritize the SRE roadmap and drive alignment across engineering, infrastructure, and product teams.
- Partner with Engineering Managers to improve reliability outcomes for their services.
- Participate in technical design discussions and ensure systems meet quality and reliability standards.
- Mentor engineers, support personal development plans, and empower the team to make decisions.
Requirements
- 5+ years of software engineering, reliability, or operations experience in a customer-focused environment.
- 2+ years managing a team of 5 or more engineers working with distributed systems, tooling, Linux, internetworking, infrastructure security, or infrastructure management.
- Experience building services and systems from inception through production.
- Experience leading cross-team projects, setting technical direction, and communicating with upper management.
- Knowledge of incident management, root cause analysis, and production follow-up practices.
- Experience with distributed systems, proxies, DNS, databases, Internet infrastructure, security, tools, or APIs.
Nice to have
- Experience with observability tools such as Jaeger, OpenTracing, ELK, Prometheus, Thanos, Grafana, or ClickHouse.
- Experience leading and hiring teams that build and operate tools and platforms.
- Experience planning execution against deadlines and short release cycles.
Culture & Benefits
- Work within an infrastructure organization providing production reliability across edge and core environments.
- Collaborate with SRE teams across Asia, Europe, and the United States to support follow-the-sun coverage.
- Contribute to projects supporting a free and open Internet, including protection for journalism, civil society, and election-related services.
- Equal opportunity employment and reasonable accommodations are available.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Anthropic
8 дней назад
Staff Software Engineer (AI Reliability Engineering)
325 000 - 390 000GBP
Anthropic
8 дней назад
Incident Response Manager (AI)
290 000 - 365 000$
6 дней назад
Engineering Manager (AI)
8 дней назад
Engineering Manager, Services (Python/Go/Ruby)
150 000 - 175 000CAD
11 дней назад
Senior Site Reliability Engineer (Kubernetes)
Anthropic
8 дней назад
Engineering Manager, Connectivity (AI)
255 000 - 325 000GBP