Назад
обновлено 1 день назад

Staff Site Reliability Engineer (Splunk)

194 000 - 267 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Site Reliability Engineer (Splunk) (Observability/SRE): Building and scaling a comprehensive observability platform and Splunk ecosystem for distributed systems with an accent on infrastructure as code, log management, and high-availability operations. Focus on automating agent deployment with Terraform and Go, optimizing Splunk Cloud performance, and leading incident response and post-incident improvements.

Location: Hybrid role associated with offices in Bellevue, Chicago, New York, San Francisco, and Washington, DC; in-person onboarding is required in San Francisco or Chicago during the first week. Candidates must be able to provide documentation establishing U.S. Person status to access federal environments or protected federal data.

Salary: $194,000–$267,000 USD annual base salary for candidates in California excluding the San Francisco Bay Area, Colorado, Illinois, New York, and Washington; a separate range of $174,000–$239,000 USD is also listed.

Company

Okta builds identity and security infrastructure for organizations and supports secure adoption of AI.

What you will do

  • Own and evolve the Splunk ecosystem as part of a scalable observability platform.
  • Design, build, and maintain observability infrastructure using Terraform and infrastructure-as-code practices.
  • Optimize Splunk log collection, processing, storage, Workload Management, and HEC performance.
  • Automate deployment and scaling of observability agents and collectors across distributed systems.
  • Participate in on-call rotations, lead post-incident reviews, and drive systemic reliability improvements.

Requirements

  • At least 5 years of experience scaling and managing Splunk Cloud, including environments with 1,000+ services, Workload Management, and HEC optimization.
  • At least 5 years in SRE, DevOps, or systems engineering roles focused on high-availability systems.
  • Strong proficiency with SPL and Go for internal tools and workflow automation.
  • Deep knowledge of Linux internals, TCP/IP, DNS, load balancing, and Kubernetes/EKS.
  • Ability to create actionable Splunk dashboards and debug complex cross-service performance bottlenecks.
  • Ability to establish U.S. Person status and attend in-person onboarding in San Francisco or Chicago.

Nice to have

  • Experience with OpenTelemetry, Vector, or similar telemetry frameworks.
  • Experience implementing Splunk charge-back applications for usage reporting.
  • Experience managing observability-native tools in AWS or GCP.

Culture & Benefits

  • Health, dental, and vision insurance.
  • 401(k) and flexible spending account.
  • Paid leave, including PTO and parental leave.
  • Equity and bonus opportunities where applicable.
  • In-person onboarding designed to connect employees with the mission and team.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →