6 дней назад
Senior Site Reliability Engineer (Databricks)
108 000 - 216 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Databricks/AWS): Building and operating reliable data platforms and customer service infrastructure with an accent on observability, automation, and production incident response. Focus on designing resilient Databricks, AWS, GCP, Kubernetes, and Kafka environments, defining SLOs, optimizing performance, and improving disaster recovery and operational reliability.
Location: 39 Tesla, Irvine, CA 92618, United States
Salary: $108,000–$216,000 per year, plus annual or quarterly performance incentives.
Company
provides customer service platforms and supporting technology infrastructure within Walmart.
What you will do
- Design and evolve monitoring, observability, alerting, and telemetry frameworks for proactive detection and root cause analysis.
- Build automation for CI/CD, testing, infrastructure management, and repetitive operational workflows.
- Participate in on-call rotations, diagnose production incidents, troubleshoot performance and availability bottlenecks, and lead blameless post-incident reviews.
- Define and manage SLIs, SLOs, and SLAs while monitoring reliability and performance across multiple applications and systems.
- Architect resilient Databricks, AWS, GCP, Kubernetes, and Kafka environments, including disaster recovery and minimum operating environment planning.
- Collaborate with engineering, data engineering, data science, offshore, Central Ops, and QA teams to improve reliability, observability, deployment safety, and operational efficiency.
Requirements
- Onsite work in Irvine, California, United States.
- Master’s degree in a relevant field and 1 year of related experience, or bachelor’s degree in a relevant field and 3 years of related experience; any amount of experience with the required skills will be accepted.
- Experience building, supporting, and maintaining Databricks platforms and infrastructure, including Spark job troubleshooting and optimization.
- Experience managing AWS services such as S3, IAM, VPC, subnets, and VPC endpoints, as well as supporting applications on AWS and GCP EKS and Kafka clusters.
- Experience with Terraform infrastructure as code, Airflow orchestration, PagerDuty, Monte Carlo, and CloudWatch monitoring.
- Experience with AWS-to-GCP migrations, infrastructure setup, data movement from S3 to Google Cloud Storage, and production support.
Culture & Benefits
- Full-time employment with health, vision, and dental coverage.
- 401(k), stock purchase opportunities, company-paid life insurance, and performance-based incentives.
- Paid time off, including sick leave, parental leave, family care leave, bereavement leave, jury duty, and voting leave.
- Short- and long-term disability coverage, education assistance including company-paid college degrees, company discounts, military service pay, and adoption expense reimbursement.
- Equal opportunity employment and a drug-free workplace.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
12 дней назад
Senior Data Ops Engineer, Data Activation & Products (SRE)
102 800 - 190 204$
13 дней назад
Principal Site Reliability Engineer (AI)
165 000 - 185 000$
11 дней назад
Senior/Lead Site Reliability Engineer (Federal)
159 000 - 230 000$
Replit
13 дней назад
Staff Site Reliability Engineer (Kubernetes/GCP)
250 000 - 325 000$
9 дней назад
Cloud Site Reliability Engineer (AWS)
120 000 - 130 000$
8 дней назад
Senior Manager, Site Reliability Engineering (AI Ops)
222 000 - 300 500$