Назад
Company hidden
2 дня назад

Staff Site Reliability Engineer (AI/ML)

Формат работы
remote (только Poland)/onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Poland
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Site Reliability Engineer (Cloud-Native AI/ML Platforms): Building reliable, observable, and operationally mature cloud-native platforms and critical data and AI/ML services with an accent on SLOs, observability, automation, and incident response. Focus on designing distributed systems for reliability, eliminating operational toil, shaping cross-team architecture, and improving Kubernetes-based production environments.

Location: Onsite in Krakow, Poland; the position is also eligible for a remote work arrangement from home, subject to role-specific details.

Company

hirify.global is a science and technology company operating across life sciences, diagnostics, and biotechnology through a global portfolio of businesses.

What you will do

  • Establish and monitor SLOs and error budgets for critical services, using reliability, availability, performance, and cost data to guide engineering decisions.
  • Design and maintain observability for cloud-native applications and infrastructure, including monitoring, logging, tracing, dashboards, alerts, and runbooks.
  • Automate repetitive operational work and improve deployment, operations, and incident-response pipelines.
  • Lead incident detection, triage, mitigation, resolution, and blameless postmortems that produce lasting preventative measures.
  • Partner with development teams to build operability, reliability, and cost efficiency into new services and identify performance and architectural risks before production.
  • Set cross-team technical direction, establish reliability standards, document systems, and mentor engineers across operating companies.

Requirements

  • 5+ years of hands-on experience in Site Reliability Engineering, DevOps, or equivalent production reliability and operations work.
  • Strong practical knowledge of SRE principles, including SLOs, error budgets, toil reduction, and blameless incident management.
  • Experience with observability platforms such as Prometheus, Grafana, Datadog, ELK, OpenTelemetry, or Splunk.
  • Hands-on experience with AWS, Azure, or GCP; Docker, Kubernetes; Infrastructure as Code such as Terraform, OpenTofu, or Pulumi; and Python or Go.
  • Strong troubleshooting skills across distributed systems, microservices, CI/CD pipelines, and large-scale data infrastructure, with experience setting technical direction and mentoring engineers.
  • Ability to work onsite in Krakow, Poland; travel of up to 10% may be required.

Nice to have

  • Experience in life sciences, diagnostics, or biotechnology.
  • Experience with chaos engineering, resilience testing, capacity planning, or FinOps.
  • Experience working in a matrixed environment.

Culture & Benefits

  • Culture focused on continuous improvement, belonging, technical rigor, blameless learning, and collaboration.
  • Comprehensive benefits, including health care programs and paid time off.
  • Flexible remote work arrangements may be available for eligible roles.
  • Career development and internal growth opportunities across hirify.global's operating companies.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →