Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Site Reliability Engineer (Splunk) (Splunk/Observability): Building and operating a scalable observability platform centered on Splunk Cloud, Terraform, and automated agents across distributed systems with an accent on log management, infrastructure as code, and high-availability engineering. Focus on optimizing Splunk collection and storage, automating deployments, investigating cross-service performance bottlenecks, and leading incident reviews.
Location: Hybrid role in Bellevue, Chicago, New York, San Francisco, or Washington, DC; in-person onboarding is required in the San Francisco or Chicago office during the first week.
Salary: $174,000–$239,000 USD annual base salary; location-specific ranges may reach $267,000.
Company
Okta builds identity and security infrastructure for organizations and AI systems.
What you will do
- Design, build, and maintain scalable observability infrastructure with Terraform.
- Own and evolve the Splunk ecosystem, optimizing log collection, processing, storage, reliability, and latency.
- Automate deployment and scaling of observability agents and collectors across distributed systems.
- Participate in on-call rotations and lead post-incident reviews to drive systemic improvements.
- Build internal tools and workflows using SPL and Go.
Requirements
- U.S. Person status is required to access federal environments or protected federal data.
- At least 5 years of experience scaling and managing Splunk Cloud, including environments with 1,000+ services, Workload Management, and HEC optimization.
- At least 5 years of experience in SRE, DevOps, or systems engineering focused on highly available systems.
- Strong proficiency in SPL and Go, plus knowledge of Linux internals, networking, TCP/IP, DNS, load balancing, and Kubernetes/EKS.
- Ability to create actionable Splunk dashboards and debug complex cross-service performance bottlenecks.
Nice to have
- Experience with OpenTelemetry, Vector, or similar telemetry frameworks.
- Experience implementing Splunk charge-back applications for usage reporting.
- Experience managing native observability tools in AWS or GCP.
Culture & Benefits
- Health, dental, and vision insurance.
- 401(k) and flexible spending account.
- Paid leave, including PTO and parental leave.
- Equity and bonus opportunities where applicable.
- In-person onboarding designed to connect new hires with the mission and team.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Site Reliability Engineer (AWS)
180 000 - 220 000$
7 дней назад
Lead Site Reliability Engineer (Kubernetes)
7 дней назад
Site Reliability Engineer - Enterprise Technology
200 000 - 250 000$
8 дней назад
Sr Site Reliability Engineer (AWS)
95 000 - 135 000$
2 дня назад
Senior Site Reliability Engineer (AI)
191 000 - 226 000$
Reddit
8 дней назад
Staff Site Reliability Engineer, Ads
217 000 - 303 900$