обновлено 5 дней назад
Site Reliability Engineer (Azure)
115 000 - 120 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Azure): Maintaining reliable multi-tenant Azure workloads and insurance payment platforms with an accent on incident response, cloud infrastructure, observability, and automation. Focus on diagnosing failures across application, database, and infrastructure layers, building operational tooling, and improving production readiness through runbooks and AI-assisted workflows.
Location: Remote within the United States
Salary: $115,000–$120,000 USD per year
Company
develops insurance payment platforms that support digital communications, inbound payment processing, and outbound disbursements for insurance carriers.
What you will do
- Own production reliability for multi-tenant Azure workloads, including AKS, App Services, virtual machines, networking, and data stores.
- Participate in rotating on-call coverage, triage alerts, resolve incidents, communicate status, and improve response with runbooks and automation.
- Build internal tooling with Python, Taskfiles, GitLab pipelines, and small services for onboarding, environment setup, migrations, and operations.
- Improve observability through Prometheus and Grafana dashboards, alerts, log search, and failure-mode runbooks.
- Diagnose reliability issues across application code, databases, and infrastructure with development and cloud teams.
- Support ClaimsPay migrations, client cutovers, production readiness reviews, and AI-assisted investigation and automation.
Requirements
- At least five years of experience in software engineering, SRE, DevOps, or system administration.
- Experience with process automation and scripting or software development; Python is preferred, with Bash and PowerShell useful.
- Proficiency with Azure identity, networking, compute, storage, and monitoring services.
- Experience administering Linux and Windows servers and troubleshooting MySQL and/or SQL Server.
- Ability to isolate issues across application, database, and infrastructure layers in development and production.
- Bachelor’s degree in computer science or engineering, or equivalent experience.
Nice to have
- Kubernetes or AKS experience.
- CI/CD experience with GitLab CI, TeamCity, Octopus Deploy, or similar tools.
- Prometheus, Grafana, Elasticsearch/Kibana, or equivalent observability experience.
- Terraform or Ansible experience and deeper database operations knowledge.
- Experience with AI coding assistants, payments, fintech, or regulated multi-tenant SaaS environments.
Culture & Benefits
- Target workload balance of approximately 50% operations and on-call support and 50% engineering and automation.
- Documentation-driven work using Confluence, GitLab, and Jira, with an emphasis on durable fixes.
- Medical, dental, and vision insurance.
- 401(k) plan and work-life balance support.
- Career advancement opportunities within the organization.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
11 дней назад
Site Reliability Engineer (AI Infrastructure)
175 000 - 265 000$
7 дней назад
Sr. Manager, Site Reliability
59 550 - 110 594GBP
5 дней назад
Site Reliability Engineer (SRE) II
124 800 - 187 200$
9 дней назад
Site Reliability Engineer Engineer
8 часов назад
DevOps/SRE Engineer
9 дней назад
Site Reliability Engineer
123 000 - 150 000$