Назад
Company hidden
1 день назад

Senior Site Reliability Specialist II

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Specialist II (Cloud Infrastructure/Kubernetes): Building reliable platform capabilities and improving the availability, scalability, performance, and resilience of critical production systems with an accent on cloud infrastructure, observability, automation, and distributed systems. Focus on leading cross-functional reliability initiatives, eliminating operational toil, improving incident response, and enabling engineering teams through self-service platforms and sound reliability practices.

Location: Remote in the United States

Company

hirify.global provides Critical Event Management technology that combines intelligent automation and risk intelligence to help enterprises and government organizations manage critical events and protect people and operations.

What you will do

  • Build platform capabilities and self-service tools that help engineering teams deliver reliable software safely and efficiently.
  • Lead complex initiatives across cloud infrastructure, Kubernetes, observability, networking, automation, and developer platforms.
  • Design solutions that improve platform availability, scalability, performance, resilience, recoverability, and operational readiness.
  • Coach engineering teams on observability, incident response, disaster recovery, capacity planning, production readiness, SLOs, SLIs, and error budgets.
  • Participate in on-call support, lead technical responses to high-severity incidents, and facilitate blameless post-incident reviews.
  • Establish engineering standards, document best practices, review architectures, and drive corrective actions through completion.

Requirements

  • Experience designing and operating complex production systems and cloud-native architectures.
  • Experience with distributed systems, container platforms, Infrastructure as Code, automation, and CI/CD.
  • Knowledge of observability, monitoring, logging, telemetry, incident response, and operational excellence.
  • Understanding of reliability engineering principles, including SLOs, SLIs, capacity planning, and performance optimization.
  • Ability to write software or automation in one or more modern programming languages.
  • Linux and networking fundamentals.

Nice to have

  • Experience working in regulated environments such as FedRAMP, DoD, IL4/IL5, SOC 2, or ISO 27001.

Culture & Benefits

  • Collaborative technical leadership environment focused on ownership, continuous improvement, operational excellence, and customer outcomes.
  • Opportunity to mentor engineers and influence architecture and engineering standards.
  • Remote work arrangement for candidates in the United States.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →