3 дня назад
Senior Sustaining and Forward Deployed Engineer (AWS & Databricks)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Sustaining and Forward Deployed Engineer (AWS & Databricks): Operating and improving production systems, data pipelines, and customer deployments with an accent on incident response, reliability engineering, and hands-on automation. Focus on debugging AWS and Databricks workloads, solving complex production failures, and translating customer escalations into durable product and platform improvements.
Location: United States; remote work is offered.
Company
develops a healthcare data platform that helps health plans create trusted, connected data foundations for decision-making, automation, and AI use cases.
What you will do
- Lead production incident triage, mitigation, recovery, and root cause analysis for complex customer-impacting issues.
- Own post-launch reliability, operational quality, defect resolution, monitoring, alerting, and runbook improvements.
- Work directly with strategic customers on deployments, integrations, escalations, and production-grade technical challenges.
- Troubleshoot AWS infrastructure, Databricks jobs and clusters, Spark-based data pipelines, performance issues, scalability problems, and data correctness defects.
- Write production-quality Python code and automation that reduces operational toil and improves reliability and observability.
- Provide technical leadership, mentor engineers, and collaborate with Product, Engineering, Data, and Customer teams.
Requirements
- 10+ years of experience in software engineering, SRE, sustaining engineering, or production operations.
- Deep hands-on experience operating production systems in AWS and troubleshooting Databricks and large-scale data platforms.
- Proficiency in Python and experience building production services or operational tooling.
- Strong knowledge of distributed systems, incident management, root cause analysis, monitoring, alerting, observability, CI/CD, and infrastructure as code.
- Ability to own problems from detection through permanent resolution and work backward from customer impact to root cause across systems and codebases.
- Excellent communication skills and the ability to operate effectively during high-pressure, ambiguous, and customer-impacting incidents.
Nice to have
- Experience in healthcare, health insurance, or regulated data environments.
- Familiarity with Kubernetes/EKS, EMR, Lambda, Spark internals, Snowflake, FHIR, MDM systems, or entity resolution.
- Experience in SWAT, escalation engineering, tiger-team, or SRE/on-call programs.
Culture & Benefits
- Work-from-anywhere flexibility within the role’s stated location framework.
- Unlimited paid time off.
- Comprehensive health coverage with multiple plan options.
- Equity for every employee and eligibility for performance bonuses.
- Home office setup allowance and monthly cell phone allowance.
- Growth-focused environment with an emphasis on collaboration, inclusion, and thoughtful use of AI and automation.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior Site Reliability Engineer (AWS)
10 дней назад
Staff Site Reliability Engineer (Cybersecurity)
199 750 - 270 000$
Datadog
4 дня назад
Senior Software Engineer - Incident Insights & Readiness (SRE)
192 000 - 240 000$
4 дня назад
Staff Site Reliability Engineer (Remote, AWS)
170 000 - 210 000$
5 дней назад
Senior Site Reliability Engineer (AWS/AI)
140 000 - 160 000$
5 дней назад