11 дней назад
Site Reliability Engineer, Observability
160 000 - 200 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer, Observability (New Relic/Terraform): Building observability infrastructure, reliability practices, and incident management foundations for high-volume enterprise treasury systems across Azure and AWS with an accent on monitoring, alert quality, infrastructure as code, and production operations. Focus on designing SLOs and error budgets, reducing alert noise, establishing incident response workflows, and troubleshooting predominantly Windows-based distributed systems.
Location: Chicago, Illinois, United States; hybrid collaboration with managers and teams deciding which 10+ days per month employees come into the office.
Annual base salary: $160,000–$200,000 USD for positions based in Illinois. Equity, bonuses, and commissions are excluded.
Company
builds crypto and treasury technology for financial institutions, businesses, governments, developers, and enterprise finance customers.
What you will do
- Design and implement New Relic monitoring, alerting, dashboards, instrumentation, distributed tracing, and NRQL-based troubleshooting across Azure and AWS.
- Define SLOs, SLIs, and error budgets while improving observability maturity, signal quality, and monitoring cost efficiency.
- Develop and govern Terraform infrastructure as code for monitoring resources and observability infrastructure.
- Author and troubleshoot Azure DevOps pipelines and support deployment visibility, change tracking, and release hygiene.
- Configure Incident.IO workflows, alert routing, integrations, on-call rotations, escalation policies, severity classification, and post-incident reviews.
- Coach engineering teams through workshops, documentation, training, and hands-on guidance on reliability and incident management practices.
Requirements
- 7+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering with a focus on observability and production operations.
- Expert hands-on experience with New Relic and NRQL, including APM, infrastructure monitoring, logs, synthetics, and alerts.
- Strong experience defining SLOs, SLIs, error budgets, incident response workflows, on-call rotations, and escalation policies.
- Strong Terraform, Azure, Azure DevOps, PowerShell, and observability infrastructure experience; working knowledge of AWS.
- Experience with Incident.IO, PagerDuty, OpsGenie, or similar incident management platforms, plus Octopus Deploy.
- Comfort working across Windows and Linux server environments and collaborating in Agile/Scrum teams.
Nice to have
- Experience reducing alert noise, optimizing observability costs, and managing log ingestion, pipeline rules, and cardinality.
- Experience with chaos engineering, game days, failure injection, or VM-hosted SQL Server monitoring.
- Knowledge of FinTech compliance requirements such as SOC 2 and ISO 27001.
- Python or Bash scripting, Jira, and organizational reliability metrics such as MTTR, MTTD, availability, and error budgets.
Culture & Benefits
- Fast-paced startup environment with experienced industry leaders and opportunities to work with current technologies.
- Professional development budget, competitive benefits, retirement support, healthcare, family-forming benefits, and family support.
- Flexible manager- and team-led hybrid scheduling with in-office collaboration for important moments.
- Vacation policy, R&R days, wellness reimbursement, parental leave, family planning benefits, and weekly onsite and virtual programming.
- Team offsites, bonding activities, catered lunches, stocked kitchens, employee giving match, and a mobile phone stipend.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Replit
13 дней назад
Site Reliability Engineer
210 000 - 275 000$
Replit
13 дней назад
Staff Site Reliability Engineer (Kubernetes/GCP)
250 000 - 325 000$
Okta
8 часов назад
Staff Site Reliability Engineer (Splunk)
194 000 - 267 000$
12 дней назад
Site Reliability Engineering Manager (AWS/Kubernetes)
205 000 - 255 000$
13 дней назад
Staff Site Reliability Engineer (AI Ops)
152 000 - 228 000$
13 дней назад
Site Reliability Engineering Manager (Fintech)
250 000 - 280 000$