5 дней назад
Senior Site Reliability Engineer (Linux Systems & Application Observability)
180 000 - 200 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Linux Systems & Application Observability): Building fault-tolerant infrastructure, internal tooling, and observability for a brokerage platform with an accent on Linux systems, distributed services, and production reliability. Focus on scaling a HashiCorp Nomad service fabric, instrumenting Ruby, Java, and Elixir services, and designing telemetry, SLOs, and error-budget practices for critical trading flows.
Location: Chicago, Illinois; hybrid with 3 days per week in the office
Base salary: $180,000–$200,000 per year. Discretionary performance bonus: 15–20% of base salary.
Company
is a retail brokerage within IG North America and IG Group, building trading platforms for options, futures, equities, forex, and digital assets.
What you will do
- Build self-healing, fault-tolerant infrastructure and internal tooling that reduces operational toil.
- Identify and close telemetry, logging, and alerting gaps across the observability stack.
- Own capacity planning, load testing, and scalability improvements for the HashiCorp Nomad service fabric.
- Extend Prometheus, Honeycomb, and OpenTelemetry instrumentation to expose production failure modes.
- Define SLOs, error budgets, and multi-window burn-rate alerting for critical brokerage flows.
- Mentor engineers and help establish a site reliability practice across infrastructure and application teams.
Requirements
- Hands-on experience designing and shipping fault-tolerant, self-healing distributed systems.
- Deep knowledge of distributed systems, Linux systems, cloud-native architectures, or containerization.
- Experience scaling production systems through capacity planning and architectural bottleneck analysis.
- Hands-on experience with OpenTelemetry, Prometheus, and Grafana, including direct service instrumentation.
- Strong Linux internals and networking fundamentals, including TCP/IP, UDP/multicast, packet capture, and flow analysis.
- Production on-call experience, blameless incident reviews, SLOs, error budgets, and programming in Python, Ruby, Java, or a similar language.
Nice to have
- Experience with HashiCorp Nomad, Consul, or Vault.
Culture & Benefits
- Performance bonuses and stock purchase options.
- Medical, vision, dental, 401(k), paid vacation, and paid sick leave.
- Gym reimbursement, in-building gym, commuter benefits, and shuttle service to and from Metra.
- Pet insurance, wellness and mental health programs, charitable donation matching, and paid volunteer days.
- Daily catered lunch, snacks, and beverages when working in the office.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
11 дней назад
Senior Site Reliability Engineer (Kubernetes)
147 600 - 221 400$
11 дней назад
Site Reliability Engineer - Enterprise Technology
200 000 - 250 000$
Reddit
12 дней назад
Staff Site Reliability Engineer, Ads
217 000 - 303 900$
9 дней назад
Senior Site Reliability Engineer (Cloud-Native Infrastructure)
142 800 - 178 500$
Okta
5 дней назад
Staff Site Reliability Engineer (Splunk)
194 000 - 267 000$
Okta
5 дней назад
Staff Site Reliability Engineer (Splunk)
174 000 - 239 000$