12 дней назад
SRE RunOps Engineer 2 (AWS/Azure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
SRE RunOps Engineer 2 (AWS/Azure): Maintaining and improving the reliability, availability, and performance of the 7NOW delivery platform with an accent on incident response, automation, observability, and cloud infrastructure. Focus on designing resilient production systems, optimizing capacity and costs, strengthening CI/CD and disaster recovery, and solving complex issues across distributed services.
Location: Onsite in Irving, Texas, United States at 3200 Hackberry Road.
Company
operates a large convenience, restaurant, and fuel retail network and develops digital delivery services through the 7NOW platform.
What you will do
- Monitor, troubleshoot, and resolve production incidents across 7NOW platform services.
- Design automation and infrastructure-as-code solutions using scripting languages and cloud tooling.
- Participate in a 24/7 on-call rotation and improve observability, monitoring, SLIs, and SLOs.
- Perform root cause analysis, post-incident reviews, capacity planning, and performance tuning.
- Maintain CI/CD pipelines, runbooks, operational documentation, backup procedures, and disaster recovery testing.
- Optimize AWS or Azure infrastructure costs and collaborate with engineering, product, infrastructure, and security teams.
Requirements
- 4+ years of relevant work experience and a bachelor's degree in computer science, information technology, engineering, or equivalent practical experience.
- Proficiency in Python, Go, Bash, or PowerShell for automation and tooling.
- Experience with AWS or Azure, cloud infrastructure, Terraform, Ansible, or CloudFormation.
- Knowledge of Prometheus, Grafana, Datadog, New Relic, Splunk, CI/CD tools, Linux/Unix, networking, and distributed systems.
- Understanding of SQL and NoSQL databases, containerization, microservices, API design, security, vulnerability remediation, and backup strategies.
- Ability to troubleshoot complex multi-layer issues, document technical processes, and support changing priorities; AWS or Azure certifications are preferred.
Nice to have
- AWS or Azure cloud certification such as Solutions Architect Associate or SysOps Administrator.
- Experience using AI to improve observability, troubleshooting, and production issue resolution.
Culture & Benefits
- 24/7 operational support through an on-call rotation.
- Collaboration with software engineering, product management, infrastructure, and security teams.
- Mentoring, technical documentation, training sessions, and code reviews.
- General benefits information is provided for eligible US and Canada positions.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior Site Reliability Engineer (AWS)
6 дней назад
Site Reliability Engineering Team Lead (Principal SRE, Automotive AI)
132 000 - 211 400$
14 дней назад
Senior SRE (Site Reliability Engineer) – Modernized Application Operations
145 000 - 170 000$
14 дней назад
Software Development Engineer, SRE (US Federal)
137 000 - 205 400$
13 дней назад
Site Reliability Engineer (AI)
200 000 - 400 000$
13 дней назад
Site Reliability Engineering (SRE) Manager (Azure)
139 700 - 232 900$