Principal Site Reliability Engineer (Infrastructure Observability)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Location: Owings Mills, Maryland, United States. Hybrid work is available, with up to three days per week from home. Applicants must have US work authorization that does not now or in the future require visa sponsorship.
Base salary: $159,000–$272,000 annually for Maryland, Colorado, Washington, and remote workers. Other listed ranges are $175,000–$299,000 for Washington, D.C. and $199,000–$339,000 for New York and California.
Company
Global asset management organization providing investment solutions across equity, fixed income, and multi-asset capabilities.
What you will do
- Design technology solutions and automations that prevent or minimize service disruptions.
- Develop observability, sustainability, scalability, measurability, and recoverability across cloud and on-premises environments.
- Analyze incidents and reliability trends across a complex, distributed technology portfolio.
- Drive adoption of SRE methodologies, including blameless post-mortems, error budgets, SLOs, and SLIs.
- Consolidate information from disconnected systems into cohesive views for identifying trends, redundancies, and risks.
- Contribute to target-state architecture and lead initiatives across multifunctional teams.
Requirements
- Bachelor’s degree or equivalent education and experience, plus 10+ years designing and operating cloud infrastructure with senior-level impact.
- 5+ years building and supporting solutions in Amazon AWS and 5+ years building and running DevOps or SRE functions.
- Experience with chaos engineering at scale, strategic program implementation, automation, incident remediation, and 24x7 monitoring and support.
- Fluency in multiple programming languages, such as Python, Java, Go, Node.js, or .NET Core, plus database development experience.
- Experience defining and managing SLOs, SLIs, availability metrics, error budgets, recovery plans, observability dashboards, APM, infrastructure monitoring, and application logging.
- Experience with observability tools such as New Relic, SolarWinds DPA, Elastic Stack, Prometheus, Grafana, Splunk, and cloud-native tools; ability to work on-call or during off-hours.
Nice to have
- Cloud or SRE-related certifications.
- Working knowledge of Azure.
Culture & Benefits
- Collaborative and inclusive work environment focused on diversity, learning, and meaningful impact.
- Competitive compensation and annual discretionary bonus eligibility.
- Retirement plan and health and wellness benefits, including online therapy.
- Paid time off for vacation, illness, medical appointments, and volunteering.
- Family care resources, including fertility and adoption benefits.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →