Sr. Site Reliability Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Sr. Site Reliability Engineer (Kubernetes/Cloud): Building and operating reliable, highly available infrastructure for an AI-native medical imaging platform with an accent on observability, incident response, and cloud-native resilience. Focus on designing self-healing systems, optimizing Kubernetes environments, and strengthening disaster recovery and compliance for clinical-grade applications.
Location: Remote; U.S. employment indicators include 401(k) benefits and E-Verify participation.
Salary: $125,000–$145,000 annually, plus an annual discretionary bonus.
Company
develops MosaicOS™, an AI-native clinical technology platform for medical imaging and radiology operations.
What you will do
- Define and improve SLIs, SLOs, error budgets, and reliability standards for clinical-grade applications and platform services.
- Lead production incident response, root cause analysis, blameless post-mortems, and corrective action planning.
- Design observability solutions covering logs, metrics, alerting, and distributed tracing.
- Optimize Kubernetes environments, including service mesh, ingress, autoscaling, performance, and capacity planning.
- Develop Infrastructure as Code, automation, runbooks-as-code, and self-healing systems.
- Drive disaster recovery, chaos engineering, load testing, CI/CD reliability, and resiliency initiatives with engineering teams.
Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Production Engineering for highly available production systems.
- Deep expertise in Kubernetes, Linux, and AWS, Azure, or GCP cloud platforms.
- Hands-on experience with Python, Go, or Bash; Terraform or CloudFormation; and production reliability practices.
- Experience with Datadog, Prometheus, Grafana, CloudWatch, distributed tracing, incident management, and performance optimization.
- Experience with healthcare IT or enterprise imaging systems, including PACS, EMRs, VNAs, DICOM, or HL7.
- Knowledge of HIPAA, HITRUST, networking, databases, service mesh, CI/CD, container security, and resilient cloud-native architectures; participation in an on-call rotation is required.
Nice to have
- Bachelor’s degree in Information Technology, Computer Science, Engineering, or a related field.
- Master’s degree or certifications such as AWS Certified Solutions Architect, AWS Certified DevOps Engineer, CKA, or CKAD.
Culture & Benefits
- Flexible remote schedules.
- Competitive benefits with tiered coverage options beginning the month after hire.
- Generous PTO and paid holidays.
- Compensation reviews and career growth opportunities.
- Health and wellness coverage, 401(k) benefits, family planning support, and telehealth options, subject to eligibility.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →