Back
updated 4 days ago

Staff Site Reliability Engineer (Splunk)

194 000 - 267 000$
Work format
hybrid
Work type
fulltime
Grade
senior
English
b2
Country
US
This vacancy is from Hirify.Global listVacancy from Hirify Global, list of international tech companies
Plus is required to make matches and apply

Match & Cover letter

Plus required for matching with this vacancy

Job description

Text:
/
TL;DR
Staff Site Reliability Engineer (Splunk) (Observability/SRE): Building and scaling a comprehensive observability platform and Splunk ecosystem for distributed systems with an accent on infrastructure as code, log management, and high-availability operations. Focus on automating agent deployment with Terraform and Go, optimizing Splunk Cloud performance, and leading incident response and post-incident improvements.

Location: Hybrid role associated with offices in Bellevue, Chicago, New York, San Francisco, and Washington, DC; in-person onboarding is required in San Francisco or Chicago during the first week. Candidates must be able to provide documentation establishing U.S. Person status to access federal environments or protected federal data.

Salary: $194,000–$267,000 USD annual base salary for candidates in California excluding the San Francisco Bay Area, Colorado, Illinois, New York, and Washington; a separate range of $174,000–$239,000 USD is also listed.

Company

Okta builds identity and security infrastructure for organizations and supports secure adoption of AI.

What you will do

  • Own and evolve the Splunk ecosystem as part of a scalable observability platform.
  • Design, build, and maintain observability infrastructure using Terraform and infrastructure-as-code practices.
  • Optimize Splunk log collection, processing, storage, Workload Management, and HEC performance.
  • Automate deployment and scaling of observability agents and collectors across distributed systems.
  • Participate in on-call rotations, lead post-incident reviews, and drive systemic reliability improvements.

Requirements

  • At least 5 years of experience scaling and managing Splunk Cloud, including environments with 1,000+ services, Workload Management, and HEC optimization.
  • At least 5 years in SRE, DevOps, or systems engineering roles focused on high-availability systems.
  • Strong proficiency with SPL and Go for internal tools and workflow automation.
  • Deep knowledge of Linux internals, TCP/IP, DNS, load balancing, and Kubernetes/EKS.
  • Ability to create actionable Splunk dashboards and debug complex cross-service performance bottlenecks.
  • Ability to establish U.S. Person status and attend in-person onboarding in San Francisco or Chicago.

Nice to have

  • Experience with OpenTelemetry, Vector, or similar telemetry frameworks.
  • Experience implementing Splunk charge-back applications for usage reporting.
  • Experience managing observability-native tools in AWS or GCP.

Culture & Benefits

  • Health, dental, and vision insurance.
  • 401(k) and flexible spending account.
  • Paid leave, including PTO and parental leave.
  • Equity and bonus opportunities where applicable.
  • In-person onboarding designed to connect employees with the mission and team.

Be careful: if the employer asks you to log into their system using iCloud/Google, send codes/passwords, or run code/software, don't do it - these are scammers. Always click "Report" or contact support. More in guide →