Назад
Company hidden
2 дня назад

Site Reliability Engineering Lead (f/m/d)

Формат работы
onsite
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
UK/Singapore/US +4 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineering Lead (AWS/Cloud Operations): Leading engineers responsible for reliable, secure, and performant cloud operations and shared platform services with an accent on team leadership, infrastructure as code, observability, and resiliency. Focus on incident response, reducing operational toil through automation, designing disaster recovery and business continuity strategies, and optimizing cloud costs.

Location: Bochum, Germany; on-site

Company

hirify.global develops AI code review and verification products that help enterprises build reliable, secure, maintainable, and compliant software.

What you will do

  • Lead and mentor a team of Cloud Engineers and Site Reliability Engineers.
  • Set and maintain high standards for system reliability, performance, and security.
  • Manage operational workload, on-call health, incident response, and automation to reduce manual toil.
  • Promote blameless post-mortems, continuous improvement, and architectural knowledge sharing.
  • Collaborate with value stream squads to operate a governed production platform that meets developer needs.
  • Define the squad vision around resiliency, business continuity, cost optimization, and platform engineering.

Requirements

  • 10+ years of software engineering experience with significant focus on SRE, cloud operations, or infrastructure engineering.
  • Deep understanding of DevOps and SRE practices, including management of mission-critical shared services such as Aurora databases, OpenSearch, and control planes.
  • Advanced AWS or comparable cloud provider expertise, including IAM, organizational units, and account vending at organizational scale.
  • Experience with infrastructure coding lifecycles and code review using Python, CDK, or Terraform.
  • Experience defining observability patterns for logging, tracing, and metrics, as well as disaster recovery and business continuity strategies.
  • Experience with Agile methodologies and cloud cost optimization, including rightsizing, Spot Instances, and Reserved Instances.

Culture & Benefits

  • Employee position with pension contributions, including a company-financed 3% gross salary benefit and a voluntary scheme with a 15% company contribution from social security savings.
  • Public transport reimbursement covering 60% of an annual subscription.
  • Annual discretionary company growth bonus.
  • Global workforce spanning 20+ countries and 35+ nationalities.
  • Annual company kick-off held at an international location.
  • Inclusive workplace with equal-opportunity employment practices and accommodation support.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →