Platform Operations Engineer (SRE)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Platform Operations Engineer (SRE): Designing and implementing cross-platform observability and incident management systems for ’s digital platform ecosystem with an accent on operational reliability and SRE principles. Focus on reducing operational toil through automation, optimizing CI/CD pipelines, and ensuring platform availability via SLOs and SLIs.
Location: Onsite at ’s Westerville, OH headquarters. Must be legally authorized to work in the US (no sponsorship provided).
Company
A $10.2 billion global critical infrastructure and data center technology company providing power, cooling, and IT infrastructure solutions.
What you will do
- Design, implement, and maintain end-to-end monitoring, alerting, and observability solutions across AI platforms and automation tools.
- Lead incident response for P1/P2 incidents, serving as incident commander and facilitating RCAs and blameless post-mortems.
- Define and enforce SLOs, SLIs, and error budgets to ensure platforms meet availability and performance targets.
- Eliminate manual operational toil through automation, scripting, and Infrastructure-as-Code (IaC) tools.
- Perform capacity planning and performance engineering to identify and resolve architectural bottlenecks.
- Partner with delivery teams to instrument CI/CD pipelines for reliability and manage progressive rollout strategies.
Requirements
- 5+ years of professional experience in platform operations, SRE, or DevOps.
- 3+ years of hands-on experience with enterprise observability platforms (e.g., Datadog, Grafana, Prometheus, Azure Monitor, Splunk).
- Strong knowledge of SRE principles, including SLOs, SLIs, and toil reduction practices.
- Proficiency with AWS, containerized environments (Docker, Kubernetes), and IaC tooling (Terraform, Ansible).
- Proficiency in multiple programming languages (Python, Ruby, Java, Javascript, C#, etc.) for automation.
- Must be legally authorized to work in the United States without sponsorship.
Nice to have
- SRE or Cloud certifications (Google SRE, AWS DevOps Professional, Azure).
- Experience with AIOps tooling or AI-assisted anomaly detection.
- Familiarity with Workato, UiPath, Power Automate, Compass AI, Writer AI, or Cursor.
- Experience with DevSecOps practices, including SAST/DAST scanning and compliance-as-code.
- Experience in Agile/Scrum environments and familiarity with ITIL frameworks.
Culture & Benefits
- Focus on Operational Excellence and a High-Performance Culture.
- Commitment to core principles: Safety, Integrity, Respect, Teamwork, and Diversity & Inclusion.
- Opportunity to work within a global organization operating in more than 130 countries.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →