2 дня назад
Staff Site Reliability Engineer (AI/ML)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Site Reliability Engineer (Cloud-Native AI/ML Platforms): Building reliable, observable, and operationally mature cloud-native platforms and critical data and AI/ML services with an accent on SLOs, observability, automation, and incident response. Focus on designing distributed systems for reliability, eliminating operational toil, shaping cross-team architecture, and improving Kubernetes-based production environments.
Location: Onsite in Krakow, Poland; the position is also eligible for a remote work arrangement from home, subject to role-specific details.
Company
is a science and technology company operating across life sciences, diagnostics, and biotechnology through a global portfolio of businesses.
What you will do
- Establish and monitor SLOs and error budgets for critical services, using reliability, availability, performance, and cost data to guide engineering decisions.
- Design and maintain observability for cloud-native applications and infrastructure, including monitoring, logging, tracing, dashboards, alerts, and runbooks.
- Automate repetitive operational work and improve deployment, operations, and incident-response pipelines.
- Lead incident detection, triage, mitigation, resolution, and blameless postmortems that produce lasting preventative measures.
- Partner with development teams to build operability, reliability, and cost efficiency into new services and identify performance and architectural risks before production.
- Set cross-team technical direction, establish reliability standards, document systems, and mentor engineers across operating companies.
Requirements
- 5+ years of hands-on experience in Site Reliability Engineering, DevOps, or equivalent production reliability and operations work.
- Strong practical knowledge of SRE principles, including SLOs, error budgets, toil reduction, and blameless incident management.
- Experience with observability platforms such as Prometheus, Grafana, Datadog, ELK, OpenTelemetry, or Splunk.
- Hands-on experience with AWS, Azure, or GCP; Docker, Kubernetes; Infrastructure as Code such as Terraform, OpenTofu, or Pulumi; and Python or Go.
- Strong troubleshooting skills across distributed systems, microservices, CI/CD pipelines, and large-scale data infrastructure, with experience setting technical direction and mentoring engineers.
- Ability to work onsite in Krakow, Poland; travel of up to 10% may be required.
Nice to have
- Experience in life sciences, diagnostics, or biotechnology.
- Experience with chaos engineering, resilience testing, capacity planning, or FinOps.
- Experience working in a matrixed environment.
Culture & Benefits
- Culture focused on continuous improvement, belonging, technical rigor, blameless learning, and collaboration.
- Comprehensive benefits, including health care programs and paid time off.
- Flexible remote work arrangements may be available for eligible roles.
- Career development and internal growth opportunities across 's operating companies.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →