5 дней назад
Senior Site Reliability Engineer (Data & ML Platform)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Data & ML Platform): Building and operating resilient, observable, and scalable infrastructure for mission-critical data and ML workloads with an accent on Databricks and Snowflake platform operations, cloud reliability, and data-platform automation. Focus on designing highly available systems, advancing observability, automating CI/CD and infrastructure with Terraform, and enabling cross-cloud data movement.
Location: United States
Company
is a healthcare data collaboration platform that provides secure, accessible, and actionable data solutions for providers, health plans, researchers, and life sciences organizations.
What you will do
- Operate and improve Databricks and Snowflake platforms, including lifecycle automation, workspace governance, job orchestration, and cost optimization.
- Architect resilient, scalable, secure infrastructure across cloud environments, including failover, autoscaling, chaos testing, and capacity planning.
- Build platform-wide monitoring, alerting, and logging with Datadog and open tooling, and define SLOs and SLAs for critical services.
- Automate deployments of data pipelines, ML workflows, and infrastructure components using GitHub Actions, Terraform, and related infrastructure-as-code tools.
- Develop patterns for inter- and intra-cloud data movement across Snowflake, S3, Delta Lake, and Kafka.
- Partner with data, ML, analytics, and application engineering teams and contribute to data platform architecture and ML enablement strategy.
Requirements
- 6+ years of experience in SRE, platform engineering, or DevOps roles supporting data-intensive or ML-powered applications.
- Hands-on experience with Databricks, including workspace setup, cluster and job management, CI/CD integration, and data orchestration; experience with Snowflake.
- Strong knowledge of AWS or similar cloud-native infrastructure, including VPCs, IAM, event-driven patterns, and serverless computing.
- Expertise with observability solutions, especially Datadog, and platform-wide logging and monitoring.
- Strong command of CI/CD, GitHub Actions, Terraform, deployment automation, shell scripting, and Python.
- Experience building highly available, fault-tolerant systems and collaborating effectively across teams.
Nice to have
- DevSecOps experience with IaC security, CI/CD security, secret management, and audit logging.
- Experience with MLflow, feature stores, GPU workload orchestration, or large-scale Databricks and Snowflake lakehouse environments.
- Knowledge of Iceberg, Glue, Azure, multi-cloud or hybrid-cloud data environments, compliance-aware architecture, or open-source infrastructure projects.
Culture & Benefits
- Collaborative, high-performance environment focused on transforming healthcare through data logistics products and services.
- Equal employment opportunity workplace committed to inclusion and reasonable accommodations.
- Total rewards program for employees.
- Some client assignments may require post-offer health screenings and vaccination documentation.
- This job is not eligible for employment sponsorship.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Senior Platform Engineer (Data)
175 500 - 195 000$
6 дней назад
Site Reliability Engineer - Vice President
3 дня назад
Senior Software Engineer, Platform Infrastructure (AWS/Kubernetes)
190 000 - 220 000$
Luxury Presence
2 дня назад
Senior DevOps Engineer (AI)
4 дня назад
Data Platform Engineer (Azure)
6 дней назад