Назад
Company hidden
10 часов назад

Site Reliability Engineer II

Формат работы
onsite
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
BAH
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer II (AWS/Kubernetes): Designing and delivering resilient reliability solutions for a personalized health platform with an accent on observability, incident response, automation, and platform performance. Focus on diagnosing production issues, optimizing capacity and fault tolerance, building operational tooling, and mentoring junior SREs.

Location: Tuzla, Bosnia and Herzegovina

Company

hirify.global provides a personalized health platform that combines health plan administration, wellbeing solutions, and care navigation for employers, health plans, and health systems.

What you will do

  • Design and implement resilient solutions for platform reliability problems with minimal rework.
  • Own observability strategies, monitor production systems, respond to incidents, and lead root cause analysis.
  • Analyze system and application metrics for performance tuning, capacity planning, and fault diagnosis.
  • Build automation for repetitive operational tasks and improve internal tooling to reduce toil.
  • Participate in planning, on-call rotations, incident triage, metrics reviews, and process improvements.
  • Collaborate with development teams, mentor junior engineers, and support technical interviews and assessments.

Requirements

  • Bachelor's degree in Computer Science, Engineering, or equivalent experience.
  • 3+ years of experience in SRE, DevOps, or infrastructure engineering.
  • Hands-on experience with AWS, Kubernetes, EKS, infrastructure as code, and monitoring patterns.
  • Strong understanding of observability principles, including SLIs, SLOs, error budgets, and structured alerting.
  • Production-quality programming proficiency in Python, Go, or Java.
  • Experience with New Relic, Datadog, Prometheus, CI/CD pipelines, deployment tooling, and incident management frameworks such as ITIL.

Nice to have

  • Experience in healthcare or another highly regulated industry.
  • GitLab or ArgoCD proficiency.
  • AWS cloud certification.

Culture & Benefits

  • Work on a health platform supporting members making important health decisions.
  • Contribute to a diverse, equitable, inclusive, and welcoming work environment.
  • Participate in a collaborative SRE organization with opportunities to mentor and raise the technical bar.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →