Назад
Company hidden
обновлено 6 дней назад

Senior Site Reliability Engineer (Observability)

160 000 - 200 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (Observability): Building observability, infrastructure-as-code, and incident management capabilities for payment infrastructure across Azure and AWS with an accent on New Relic, Terraform, Windows-based systems, and production reliability. Focus on designing SLOs and alerting, reducing alert noise, establishing incident response practices, and troubleshooting high-volume enterprise treasury systems.

Location: United States; positions based in Illinois require in-office collaboration for 10+ days per month.

Annual base salary for positions based in Illinois: $160,000–$200,000 USD, excluding equity, bonuses, and commissions.

Company

hirify.global builds crypto solutions for financial institutions, businesses, governments, and developers, including hirify.global Treasury technology for managing liquidity across traditional and digital assets.

What you will do

  • Design and implement monitoring, alerting, dashboards, NRQL queries, SLOs, SLIs, and error budgets in New Relic across Azure and AWS.
  • Build and govern Terraform-based observability infrastructure, monitoring resources, alert configurations, and Azure DevOps pipelines.
  • Administer Incident.IO workflows, alert routing, notifications, integrations, runbooks, on-call rotations, escalation policies, and incident severity models.
  • Respond to production incidents, facilitate post-incident reviews, and track MTTR, MTTD, incident frequency, availability, and error budgets.
  • Coach engineering teams on structured logging, metrics, distributed tracing, dashboard design, and operational maturity.
  • Partner with the Subsystems Platform Team to create self-service observability and incident management capabilities.

Requirements

  • 7+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering with a strong focus on observability and production operations.
  • Expert hands-on experience with New Relic and NRQL, including APM, infrastructure monitoring, logs, synthetics, and alerts.
  • Strong Terraform, Azure, Azure DevOps, PowerShell, and incident management experience; working knowledge of AWS is required.
  • Experience designing SLOs, SLIs, error budgets, alerting strategies, incident response workflows, on-call rotations, and escalation policies.
  • Experience with Azure App Services, Virtual Machines, Azure SQL, networking, monitoring, Octopus Deploy, Slack, and both Windows and Linux environments.
  • Ability to deliver hands-on engineering work while coaching and mentoring cross-functional Agile/Scrum teams.

Nice to have

  • Experience with alert noise reduction, observability cost optimization, chaos engineering, game days, or failure injection.
  • Knowledge of VM-hosted SQL Server monitoring, FinTech compliance requirements, SOC 2, ISO 27001, and audit evidence collection.
  • Python or Bash scripting experience and familiarity with Jira for incident tracking and workflow automation.

Culture & Benefits

  • Fast-paced start-up environment with experienced industry leaders and a professional development budget.
  • Flexible team and manager decisions about in-office collaboration, with 10+ office days per month for moments that matter.
  • Competitive salary, bonuses, equity, healthcare, retirement benefits, family support, and a mobile phone stipend.
  • Vacation policy, R&R days, wellness reimbursement, parental leave, and family planning benefits.
  • Catered lunches, stocked kitchens, team offsites, bonding activities, and company events.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →