2 дня назад
Senior Staff Site Reliability Engineer
181 000 - 263 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Staff Site Reliability Engineer (Kubernetes/Terraform): Defining organization-wide SRE strategy and developing reliable infrastructure for globally distributed products with an accent on observability, automation, cloud security, and FinOps. Focus on designing distributed systems, improving multi-region architecture, leading high-impact incident response, and setting production readiness standards across engineering teams.
Location: Hybrid in San Francisco, New York, or Seattle, United States
Annual base compensation: $181,000–$263,000
Company
provides a data collaboration network for brands, retailers, financial services providers, and healthcare organizations, supporting marketing AI and customer insights.
What you will do
- Define the organization-wide SRE strategy, including SLOs, SLAs, error budgets, and operational excellence frameworks.
- Develop and operate complex infrastructure across multiple products, services, environments, and regions.
- Lead distributed-systems architecture reviews, API design, automation, performance optimization, and production readiness standards.
- Act as the final escalation point for high-impact global incidents and lead postmortems with organization-wide follow-up actions.
- Drive FinOps strategy across Kubernetes, cloud resources, and database infrastructure.
- Mentor Staff Engineers and influence technical decisions across engineering teams without direct authority.
Requirements
- 10+ years of experience in SRE, production engineering, or platform engineering, including 3+ years at senior or staff level.
- B.S./M.S. in Computer Science, Software Engineering, or equivalent experience.
- Expertise with Terraform, Kubernetes, highly available globally distributed systems, and real-time or NoSQL databases.
- Strong proficiency in Python and/or Go for production-grade internal tooling.
- Advanced experience with observability engineering, FinOps, CI/CD platforms, and cloud security across GCP and/or AWS.
- Exceptional communication skills and the ability to lead technical initiatives across engineering organizations.
Nice to have
- Experience with multi-region active-active architectures and chaos engineering frameworks.
- Contributions to open-source observability or infrastructure tooling.
- Staff-plus individual contributor or global SRE technical leadership experience.
- Familiarity with LLMs, AI-assisted development, and agentic infrastructure automation workflows.
Culture & Benefits
- Organization-wide scope across global products, infrastructure, and engineering teams.
- Opportunity to influence architecture, reliability standards, and technical strategy.
- Work across multiple products, services, and regions.
- Compensation is determined based on experience, skills, geography, and internal equity.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
8 дней назад
Head of Site Reliability Engineering (AI)
195 000 - 285 000$
5 дней назад
Sr Software Development Engineer, SRE (US Federal)
163 800 - 245 800$
3 дня назад
Site Reliability Engineer
123 000 - 150 000$
3 дня назад
Staff Site Reliability Engineer (Remote, AWS)
170 000 - 210 000$
4 дня назад
Senior Site Reliability Engineer (Fintech)
160 000 - 200 000$
8 дней назад
Staff Site Reliability Engineer (AI)
220 000 - 260 000$