Назад
обновлено 9 часов назад

Staff Site Reliability Engineer (Splunk)

194 000 - 267 000$
Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Site Reliability Engineer (Splunk) (Observability/Splunk): Building and operating a scalable observability platform for distributed systems with an accent on Splunk Cloud, infrastructure as code, and high-availability operations. Focus on automating agent deployment with Terraform and Go or SPL, optimizing log collection and storage, and leading incident reviews for systemic reliability improvements.

Location: Hybrid role based in Bellevue, Washington; Chicago, Illinois; New York, New York; San Francisco, California; or Washington, DC. In-person onboarding is required in the San Francisco or Chicago office during the first week.

Salary: $194,000–$267,000 USD annual base salary for specified U.S. locations; $174,000–$238,000 USD annual base salary for additional listed locations. Equity, bonus, and benefits may also apply.

Company

Okta builds identity and security infrastructure that helps organizations securely adopt digital and AI technologies.

What you will do

  • Own and evolve the Splunk ecosystem as part of a comprehensive, scalable observability platform.
  • Design, build, and maintain observability infrastructure using Terraform and infrastructure-as-code practices.
  • Optimize Splunk log collection, processing, storage, Workload Management, and HTTP Event Collector performance.
  • Automate deployment and scaling of observability agents and collectors across distributed systems.
  • Participate in on-call rotations, lead post-incident reviews, and drive observability-focused reliability improvements.

Requirements

  • U.S. Person status is required to access federal environments or protected federal data.
  • At least 5 years of experience scaling and managing Splunk Cloud at scale, including environments with 1,000+ services, Workload Management, and HEC optimization.
  • At least 5 years of experience in SRE, DevOps, or systems engineering focused on high-availability systems.
  • Strong programming skills in SPL and Go, with experience building internal tools and automating workflows.
  • Deep knowledge of Linux internals, networking including TCP/IP, DNS, and load balancing, and Kubernetes/EKS.
  • Strong troubleshooting skills for complex cross-service performance bottlenecks and experience creating actionable Splunk dashboards.

Nice to have

  • Experience with OpenTelemetry, Vector, or similar telemetry frameworks.
  • Experience implementing Splunk charge-back applications for usage reporting.
  • Experience managing observability-native tools in AWS or GCP.

Culture & Benefits

  • Health, dental, and vision insurance.
  • 401(k), flexible spending account, paid time off, and parental leave.
  • Equity and bonus opportunities where applicable.
  • Connection and collaboration across a global community spanning more than 20 offices.
  • In-person onboarding designed to connect new hires with the organization and role.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →