2 дня назад
Senior Site Reliability Engineer (Azure/AWS)
130 000 - 160 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Azure/AWS): Owning the availability, performance, and capacity of production SaaS services across Azure, AWS, and FedRAMP High environments with an accent on observability, automation, incident response, and infrastructure as code. Focus on defining SLOs, tuning Datadog and Azure Monitor, automating remediation, troubleshooting distributed systems, and controlling observability costs.
Location: U.S. Remote. Candidates must not require any type of U.S. work authorization now or in the future, including F1-OPT, F1-CPT, H-1B, TN, L-1, and J1.
Salary: $130,000–$160,000 base annually, plus equity and bonus opportunities.
Company
develops a cloud-native Identity Security Platform for securing human and machine identities across cloud infrastructure, traditional systems, SaaS applications, and AI environments.
What you will do
- Own availability, performance, capacity, and reliability for production SaaS services running in Azure and AWS, including FedRAMP High environments.
- Define SLIs, SLOs, and error budgets with engineering and product teams to prioritize reliability work.
- Build and tune Datadog and Azure Monitor observability, including dashboards, synthetic checks, anomaly detection, alert routing, and cost controls.
- Automate remediation workflows, maintain Terraform infrastructure and Azure DevOps pipelines, and reduce manual operational work.
- Lead high-severity incident response, support escalations, post-incident reviews, and customer-facing root cause analyses.
- Improve on-call operations, administer web application firewalls, and ensure new services ship with monitoring, runbooks, and SLOs.
Requirements
- 8+ years of experience in Site Reliability Engineering, DevOps, cloud operations, or production engineering for a SaaS product.
- Hands-on Azure experience with AKS, App Service, Azure SQL, Redis, Service Bus, Front Door, Storage, cloud networking, and cloud security.
- Production experience with Datadog or a comparable observability platform, including metrics, logs, APM, queries, dashboards, and monitor design.
- Production Kubernetes experience, plus Terraform and CI/CD pipeline creation and troubleshooting; Azure DevOps is preferred.
- Strong scripting and systems knowledge with PowerShell, Python, YAML, JSON, DNS, TLS, load balancing, reverse proxies, firewalls, and packet-level troubleshooting.
- Experience with incident response, disaster recovery, technical communication, and participation in weekend and emergency on-call rotations; up to 10% travel.
Nice to have
- Experience with FedRAMP, NIST 800-53, SOC 2, ISO 27001, PCI, or other regulated environments.
- AWS, CloudFormation, SES, Jenkins, SaltStack, Consul, ELK, CloudWatch Logs Insights, or web application firewall administration.
- Experience with Microsoft Entra ID, SAML, OIDC, multi-region architectures, disaster recovery testing, PagerDuty, Jira Service Management, or chaos exercises.
Culture & Benefits
- Work with a global engineering organization focused on identity security and protecting enterprise systems.
- Culture centered on innovation, collaboration, respect, ownership, adaptability, and integrity.
- Healthcare insurance, pension or retirement matching, life insurance, employee assistance programs, paid time off, and company holidays.
- Competitive salary, performance-based bonus opportunities, equity, and career development.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Okta
5 дней назад
Staff Site Reliability Engineer, Networking (AWS/FedRAMP)
174 000 - 238 000$
Okta
4 дня назад
Staff Site Reliability Engineer, Networking (AWS/Networking)
174 000 - 238 000$
3 дня назад
Site Reliability Engineer (Azure/Terraform)
105 600 - 145 200$
3 дня назад
Lead Site Reliability Engineer (Azure)
113 000 - 142 300CAD
5 дней назад
Senior Site Reliability Engineer (AI)
191 000 - 226 000$
Okta
5 дней назад
Senior Manager, Site Reliability Engineering (Federal)
207 000 - 284 900$