Назад
Company hidden
7 дней назад

Principal Site Reliability Engineer (Infrastructure Observability)

159 000 - 272 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
c1
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Site Reliability Engineer (Infrastructure Observability) (AWS/SRE): Designing and operating observability, reliability, and recovery solutions for complex distributed cloud and on-premises infrastructure with an accent on automation, scalability, incident prevention, and service-level objectives. Focus on building SRE practices, implementing chaos engineering at scale, standardizing monitoring and dashboards, and reducing technology failures across the enterprise.

Location: Owings Mills, Maryland, United States. Hybrid work is available, with up to three days per week from home. Applicants must have US work authorization that does not now or in the future require visa sponsorship.

Base salary: $159,000–$272,000 annually for Maryland, Colorado, Washington, and remote workers. Other listed ranges are $175,000–$299,000 for Washington, D.C. and $199,000–$339,000 for New York and California.

Company

Global asset management organization providing investment solutions across equity, fixed income, and multi-asset capabilities.

What you will do

  • Design technology solutions and automations that prevent or minimize service disruptions.
  • Develop observability, sustainability, scalability, measurability, and recoverability across cloud and on-premises environments.
  • Analyze incidents and reliability trends across a complex, distributed technology portfolio.
  • Drive adoption of SRE methodologies, including blameless post-mortems, error budgets, SLOs, and SLIs.
  • Consolidate information from disconnected systems into cohesive views for identifying trends, redundancies, and risks.
  • Contribute to target-state architecture and lead initiatives across multifunctional teams.

Requirements

  • Bachelor’s degree or equivalent education and experience, plus 10+ years designing and operating cloud infrastructure with senior-level impact.
  • 5+ years building and supporting solutions in Amazon AWS and 5+ years building and running DevOps or SRE functions.
  • Experience with chaos engineering at scale, strategic program implementation, automation, incident remediation, and 24x7 monitoring and support.
  • Fluency in multiple programming languages, such as Python, Java, Go, Node.js, or .NET Core, plus database development experience.
  • Experience defining and managing SLOs, SLIs, availability metrics, error budgets, recovery plans, observability dashboards, APM, infrastructure monitoring, and application logging.
  • Experience with observability tools such as New Relic, SolarWinds DPA, Elastic Stack, Prometheus, Grafana, Splunk, and cloud-native tools; ability to work on-call or during off-hours.

Nice to have

  • Cloud or SRE-related certifications.
  • Working knowledge of Azure.

Culture & Benefits

  • Collaborative and inclusive work environment focused on diversity, learning, and meaningful impact.
  • Competitive compensation and annual discretionary bonus eligibility.
  • Retirement plan and health and wellness benefits, including online therapy.
  • Paid time off for vacation, illness, medical appointments, and volunteering.
  • Family care resources, including fertility and adoption benefits.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →