12 дней назад
Senior Data Ops Engineer, Data Activation & Products (SRE)
102 800 - 190 204$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Data Ops Engineer, Data Activation & Products (SRE) (Kubernetes/Databricks/Spark): Improving the reliability, observability, and operational maturity of data platforms, Kubernetes deployment systems, internal applications, and cloud environments with an accent on incident response, streaming workloads, and production data services. Focus on building dashboards and alerting, troubleshooting Databricks and Spark environments, automating operational work, and strengthening deployment and rollback safety.
Location: Santa Monica, California, United States; onsite work is available at the Los Angeles (Pen Factory) office Monday through Thursday, or the role may be remote.
Salary: $102,800–$190,204 annual base pay in the U.S.; incentive compensation may also be available.
Company
develops, publishes, and distributes interactive gaming and entertainment products, including major video game franchises.
What you will do
- Monitor service health, respond to alerts, and participate in incident response across cloud, Kubernetes, application, and data platform environments.
- Investigate reliability issues across Kubernetes, networking, DNS, Databricks, Spark, orchestration systems, event systems, and dependent services.
- Build dashboards, alerting, runbooks, operational documentation, and Grafana-based monitoring for platform and streaming signals.
- Improve reliability for Databricks workflows, Spark streaming, Airflow/Astronomer DAGs, dbt jobs, event systems, APIs, workers, and internal applications.
- Automate repetitive operational work, improve deployment and rollback safety, and support platform modernization and migration efforts.
- Partner with data engineers, analytics engineers, and software engineers on reliability standards and post-incident corrective actions.
Requirements
- 5+ years of experience in SRE, DevOps, cloud infrastructure, platform engineering, software engineering, data platform operations, or related production-support roles.
- Hands-on experience with Kubernetes-based workloads, deployment systems, cloud infrastructure, or production application environments.
- Familiarity with Linux, HTTP, DNS, containers, Kubernetes, Git-based workflows, and scripting in Bash, Python, or similar languages.
- Experience with monitoring, logs, metrics, dashboards, alerting, and incident management practices.
- Experience with event systems such as Kafka, Google Pub/Sub, Kinesis, or similar technologies, including consumers, retries, lag, and dead-letter queues.
- Strong troubleshooting and communication skills, with an interest in data platforms, orchestration, streaming workloads, and production data services.
Nice to have
- Experience with Helm, ArgoCD, GitOps, CI/CD automation, or Terraform.
- Familiarity with Databricks, Spark Structured Streaming, Airflow, Astronomer, dbt, Kafka, Pub/Sub, object storage, or lakehouse architectures.
- Experience with Grafana, Prometheus, Cloud Monitoring, Datadog, Splunk, or similar observability platforms.
- Exposure to data reliability concepts, SLOs, SLAs, error budgets, postmortems, or formal reliability practices.
Culture & Benefits
- Onsite or remote work arrangements, subject to business needs.
- Medical, dental, vision, healthcare spending, life, disability, and related insurance programs.
- 401(k) with company match, tuition reimbursement, and charitable donation matching.
- Paid holidays, vacation, sick time, parental leave, and other leave programs.
- Mental health, wellbeing, fitness, gaming, and additional voluntary benefit programs.
- Relocation assistance may be available when the company requires a geographic move for the role.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
12 дней назад
Senior Data Ops Engineer, Data Activation & Products - Activision
102 800 - 190 204$
Replit
13 дней назад
Site Reliability Engineer
210 000 - 275 000$
13 дней назад
Principal Site Reliability Engineer (AI)
165 000 - 185 000$
12 дней назад
Site Reliability Engineering Manager (AWS/Kubernetes)
205 000 - 255 000$
Replit
13 дней назад
Staff Site Reliability Engineer (Kubernetes/GCP)
250 000 - 325 000$
12 дней назад
Senior Site Reliability Engineer (AI)
185 500 - 232 000$