4 дня назад
Site Reliability Engineer (Azure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Azure): Designing and operating reliable, secure, and scalable cloud infrastructure for a decision intelligence SaaS platform with an accent on Azure automation, observability, infrastructure as code, and security controls. Focus on defining SLOs, improving GitOps and CI/CD workflows, resolving incidents, optimizing cloud costs, and deploying agentic workflows for vulnerability triage and incident response.
Location: Remote in the United States within the US Eastern Time zone, with authorization to work in the United States without current or future sponsorship, or remote within Sweden; Göteborg preferred
Company
is a PE-backed SaaS company providing a decision intelligence platform that combines decisioning, process automation, and machine learning for mission-critical organizations.
What you will do
- Design, implement, and manage tooling for developing, deploying, monitoring, and supporting Azure cloud infrastructure and services.
- Define and track SLIs and SLOs, use error budgets, and improve service reliability and delivery velocity.
- Monitor cloud resources and applications, troubleshoot incidents, participate in on-call rotation, and collaborate with Engineering and IT.
- Build and maintain infrastructure as code with Bicep and OpenTofu while advancing GitOps-based workflows.
- Implement Azure security controls, least-privilege access, just-in-time elevation, infrastructure drift detection, and security scanning in CI/CD pipelines.
- Automate release change tracking, vulnerability triage, incident response, cloud operations, and cost optimization through agentic workflows.
Requirements
- 5+ years of experience in cloud operations, site reliability engineering, or DevOps.
- Strong knowledge of core cloud services and experience with PowerShell plus scripting or development languages such as SQL, Python, C#, or JavaScript.
- Experience with infrastructure as code, including OpenTofu, Bicep/ARM, or Pulumi.
- Experience with an observability platform and OpenTelemetry-based instrumentation.
- Strong DevOps fundamentals, CI/CD experience, Linux, Docker, and Kubernetes knowledge.
- Willingness to own security-adjacent work as part of core SRE responsibilities.
Nice to have
- Azure experience with Functions, Container Solutions, SQL Database, Storage, and Key Vault.
- Experience with Entra ID, Defender for Cloud, Sentinel, and Azure Policy.
- Experience in regulated environments such as SOC 2, ISO 27001, or HIPAA.
- Networking, DNS, VPC, database operations, Windows Server, IIS, or configuration management experience with Ansible, Puppet, or Chef.
Culture & Benefits
- Competitive compensation and benefits.
- Flexible work environment.
- Opportunity to build and shape a premium support function with measurable customer impact.
- Close collaboration across Support, Engineering, Product, and Customer Success.
- Professional growth within a scaling SaaS organization.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Site Reliability Engineer I (Azure)
Baseten
8 дней назад
Site Reliability Engineer (AI)
165 000 - 330 000$
5 дней назад
Staff Site Reliability Engineer (SaaS)
6 дней назад
Site Reliability Engineering Team Lead (Principal SRE, Automotive AI)
132 000 - 211 400$
8 дней назад
Senior Site Reliability Engineer (Fintech)
160 000 - 200 000$
Replit
10 дней назад
Engineering Manager (SRE)
250 000 - 325 000$