Senior Site Reliability Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior Site Reliability Engineer (SRE/AWS/Kubernetes): Lead Site Reliability Engineering initiatives by defining SLIs, SLOs, error budgets, and reliability standards for clinical-grade applications and platform services. Focus on production incident management, observability and distributed tracing, Kubernetes performance and resiliency engineering, and automation with Infrastructure as Code to reduce MTTR/MTTD.
Location: Remote
Salary: $125,000 - $145,000
Company
builds an AI-native platform for medical imaging and clinical decision-driven care.
What you will do
- Define and improve SLIs, SLOs, error budgets, and reliability standards for clinical-grade applications and platform services.
- Own production incident management, including high-severity response, root cause analysis, blameless post-mortems, and corrective action planning.
- Design and maintain observability across logs, metrics, alerting, and distributed tracing using modern monitoring platforms.
- Optimize Kubernetes-based environments (service mesh, ingress, autoscaling, performance tuning, and capacity planning).
- Build automation, Infrastructure as Code (IaC), runbooks-as-code, and self-healing systems to reduce manual support effort.
- Drive disaster recovery, chaos engineering, load testing, and resiliency initiatives; participate in an on-call rotation.
Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Production Engineering supporting high-availability production systems.
- Deep expertise in Kubernetes and Linux, plus cloud platforms such as AWS (EKS, IAM, VPC), Azure, or GCP.
- Hands-on experience with Python, Go, or Bash, and Infrastructure as Code (Terraform, CloudFormation), including SLIs/SLOs in production.
- Experience with observability/monitoring tools (Datadog, Prometheus, Grafana, CloudWatch) and distributed tracing, plus large-scale troubleshooting and performance optimization.
- Experience supporting healthcare IT/enterprise imaging environments and interoperability standards such as DICOM and HL7.
- Strong understanding of HIPAA and HITRUST and regulated cloud environments; networking (TCP/IP, HTTP, gRPC, DNS) and databases (PostgreSQL, MySQL, Oracle).
Culture & Benefits
- Flexible remote schedules.
- Generous PTO and paid holidays.
- Tiered benefits options with eligibility starting the month after hire, including health & wellness coverage.
- 401k benefits and an annual discretionary bonus eligibility.
- Compensation reviews and career growth opportunities.
Hiring process
- Interviews and evaluation of technical fit for SRE, reliability, and regulated healthcare environments.
- Final selection includes review of experience and certifications.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →