Director of Site Reliability Engineering (Blockchain)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Director of Site Reliability Engineering (Blockchain): Leading a high-leverage SRE function and building the infrastructure, reliability frameworks, and operational practices that support Stellar engineering teams with an accent on cloud platforms, Kubernetes, observability, and service ownership. Focus on designing scalable self-service infrastructure, improving incident response and deployment safety, and evaluating AI-assisted workflows to reduce operational toil.
Location: Hybrid role based in New York, United States
Salary: $205,000–$305,000 base salary, plus lumen-denominated grants and benefits.
Company
is a nonprofit organization supporting the development of the Stellar blockchain network and equitable access to the global financial system.
What you will do
- Lead, coach, and develop a distributed SRE team while defining its vision, charter, operating model, roadmap, and success measures.
- Establish a service ownership and maturity framework across engineering based on service criticality.
- Own and improve cloud foundations, Kubernetes and compute patterns, CI/CD, observability, secrets management, GitHub workflows, and infrastructure automation.
- Enable engineering teams to improve operational readiness through standards, dashboards, runbooks, alerting, escalation paths, and deployment practices.
- Improve resilience, self-healing, disaster recovery, incident response, postmortems, and on-call health.
- Partner with Security, Compliance, Legal, Finance, Procurement, and Corporate IT on infrastructure, access, cloud operations, vendors, and controls.
Requirements
- 10+ years of experience in SRE, infrastructure, platform engineering, cloud infrastructure, production operations, or related engineering roles.
- 5+ years of experience leading, managing, or formally developing infrastructure, SRE, platform, or reliability engineers.
- Deep technical judgment across cloud infrastructure, production operations, distributed systems, reliability, automation, and operational risk.
- 3+ years of experience with AWS, GCP, or similar cloud environments, plus Kubernetes, infrastructure as code, declarative systems, CI/CD, and deployment safety.
- Strong experience with observability, monitoring, alerting, logging, dashboards, SLOs/SLIs, incident response, postmortems, and on-call practices.
- Clear executive communication skills and the ability to partner directly with a CTO and senior engineering leaders.
Nice to have
- Experience supporting globally distributed teams or 24/7 operational coverage.
- Experience with infrastructure security, secrets management, access controls, cloud security, or compliance-related controls.
- Experience in financial services, regulated environments, blockchain, crypto, Web3, or other high-reliability ecosystems.
- Practical experience applying AI-assisted or agentic workflows to infrastructure, reliability, observability, or developer productivity.
Culture & Benefits
- Mission-driven nonprofit organization focused on equitable access to the global financial system.
- Flexible time off, 15 company holidays, and a company-wide holiday break.
- Health, dental, and vision coverage, plus parental leave and disability benefits.
- 401(k) with a 4% match, FSA and HSA contributions, and commuter benefits.
- $1,500 annual learning and development budget, gym reimbursement, wellbeing benefits, office meals, and company retreats.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →