6 дней назад
Site Reliability Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AWS/Kubernetes): Operating and improving a stable, scalable AWS and Kubernetes platform with an accent on Terraform infrastructure as code, Postgres reliability, and safer CI/CD delivery. Focus on building observability and incident-response practices, testing database recovery procedures, and reducing operational failure points in a HIPAA-compliant healthcare technology environment.
Location: In-office in Lehi, Utah, United States
Company
is a healthcare technology company building a secure, HIPAA-compliant AI platform for home health intake, clinical documentation, coding, and quality assurance.
What you will do
- Operate and improve AWS and Kubernetes environments, including Terraform modules, state management, review standards, and drift control.
- Improve deployment reliability through safer CI/CD checks, repeatable releases, reliable rollbacks, environment consistency, and release observability.
- Own production Postgres reliability, including backups, restore testing, monitoring, performance, capacity, access, and safe schema migrations.
- Mature observability with dashboards, metrics, logs, traces, actionable alerts, service-level indicators, and service-level objectives.
- Participate in the engineering on-call rotation, lead incident recovery, and conduct blameless post-incident reviews with tracked corrective actions.
- Partner with Security and product engineers on least-privilege access, secrets management, encryption, vulnerability remediation, disaster recovery, and auditable infrastructure changes.
Requirements
- At least 5 years of experience in site reliability, platform, infrastructure, or production engineering.
- Strong hands-on experience with AWS, production Kubernetes, Terraform, and infrastructure as code.
- Production Postgres experience covering backup and restore, performance, capacity, monitoring, and safe migrations.
- Experience building and operating CI/CD and deployment systems, plus modern observability across metrics, logs, and traces.
- Strong programming or scripting skills for automation and operational tools, with a record of diagnosing production failures and leading incidents through recovery.
- Ability to work independently, write clear automation and runbooks, and collaborate effectively with application engineers.
Nice to have
- Kubernetes security controls, policy as code, or software supply-chain security experience.
- Healthcare or other regulated-environment experience.
- Experience supporting SOC 2 controls or audit evidence.
- Experience improving cloud cost and capacity efficiency.
Culture & Benefits
- Competitive salary and meaningful equity.
- 401(k) and medical, dental, vision, and HSA insurance.
- High ownership with the opportunity to shape the security function from day one.
- Direct collaboration with founders, engineering leadership, and product engineers.
- Fast-paced, product-driven engineering culture focused on improving healthcare operations.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 часа назад
Staff DevOps Engineer (AWS)
6 дней назад
Infrastructure Engineer (AI)
200 000 - 275 000$
6 дней назад
DevOps Team Leader (AWS)
6 дней назад
Senior Site Reliability Engineer (Fintech)
160 000 - 200 000$
1 день назад
Engineer III, Site Reliability (SRE)
7 часов назад