обновлено 7 дней назад
Lead Enterprise Software Engineer (Application & IT Operations)
118 300 - 207 400$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Lead Enterprise Software Engineer (Application & IT Operations) (Cloud/AppOps): Managing the operational lifecycle of critical enterprise applications across Azure and AWS with an accent on reliability, observability, incident response, and release operations. Focus on designing monitoring and log pipelines, leading root cause analysis, automating runbooks, and improving resiliency, capacity, security, and disaster recovery.
Location: Hybrid role based in Riverwoods, Illinois, USA; candidates may be required to appear onsite at a office for interviews.
Salary: $118,300–$207,400 USD annually, plus bonus eligibility.
Company
provides enterprise-scale software and digital information solutions.
What you will do
- Own production reliability for critical applications, including SLOs, error budgets, capacity planning, and performance baselines.
- Lead major incident response, escalations, root cause analysis, stakeholder communication, and preventative improvements.
- Coordinate releases and changes, enforce readiness gates, validate deployments, and monitor post-release application health.
- Design and maintain dashboards, alert strategies, logging and tracing pipelines, and observability workflows.
- Develop runbooks and automation, reduce operational toil, and improve application support processes.
- Drive resiliency, disaster recovery, security, compliance, service reviews, operational KPIs, and mentoring for AppOps engineers.
Requirements
- 8–10 years of relevant experience and a bachelor’s degree in computer science, information systems, or a related field.
- Advanced experience operating applications on Azure and/or AWS, including networking, load balancers, DNS, certificates, storage, and messaging services.
- Hands-on experience with Datadog, Grafana/Prometheus, ELK/OpenSearch, OpenTelemetry, alert engineering, incident management, and RCA.
- Proficiency in PowerShell, Bash, or Python for runbooks, health checks, and remediation workflows; experience with Terraform and other infrastructure-as-code tools.
- Experience with CI/CD, blue/green, rolling and canary deployments, traffic management, multi-environment lifecycles, and Kubernetes.
- Ability to lead 24x7 reliability operations, change management, security and compliance activities, stakeholder communication, and mentoring.
Nice to have
- Azure Administrator/Architect or AWS SysOps/DevOps Professional certification.
- ITIL Foundation, Terraform, or SRE certification.
- Industry-recognized Kubernetes certification.
Culture & Benefits
- Hybrid work environment with an on-call rotation.
- Medical, dental, and vision plans.
- 401(k), FSA/HSA, commuter benefits, and tuition assistance.
- Vacation, sick time, and paid parental leave.
Hiring process
- Interviews are conducted without AI tools or external prompts and may include in-person interviews.
- Applicants may be required to appear onsite at a office.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Senior Site Reliability Engineer (Observability)
160 000 - 200 000$
5 дней назад
Site Reliability Engineer - Enterprise Technology
200 000 - 250 000$
4 дня назад
Sr Staff Site Reliability Engineer (AI)
207 400 - 259 200$
3 дня назад
Sr. Site Reliability Engineer (Kubernetes/AWS)
4 дня назад
Senior Site Reliability Engineer (Cloud-Native Infrastructure)
142 800 - 178 500$
1 день назад
Site Reliability Engineer II (AWS)
100 000 - 110 000$