15 часов назад
Senior Observability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Observability Engineer (AI) (OpenTelemetry, Grafana, Langfuse): Building and operating a scalable enterprise observability platform across applications, infrastructure, cloud services, networks, databases, containers, and AI-powered systems with an accent on telemetry governance, resilient architecture, and cost-efficient logging. Focus on designing AI/LLM monitoring, optimizing telemetry pipelines, implementing automation and self-healing, and improving incident response through actionable alerts and root cause analysis.
Location: Warsaw, Poland; hybrid 70/30 work model with 70% in-office collaboration.
Company
is a biotechnology company developing mRNA medicines and the technology platform, infrastructure, and services that support their creation and delivery.
What you will do
- Own and evolve the enterprise observability platform, including strategy, governance, architecture standards, roadmap, capacity planning, and cost optimization.
- Build scalable monitoring and telemetry architectures for applications, infrastructure, databases, containers, cloud services, networks, distributed systems, and AI/LLM workloads.
- Develop metrics, traces, and logs pipelines across hybrid cloud and on-premises environments using OpenTelemetry, Prometheus, Grafana, VictoriaMetrics, and related technologies.
- Design AI observability capabilities for agents, LLM applications, and agentic workflows, including prompt, response, latency, error, reliability, and cost monitoring.
- Lead logging strategy with Grafana Loki, integrate observability with PagerDuty, ServiceNow, Jira, and incident workflows, and improve alert quality and incident response.
- Develop automation, self-healing, dashboards, runbooks, executive reporting, and continuous improvement practices using Python, Terraform, Ansible, and CI/CD.
Requirements
- 7+ years of experience in site reliability engineering, observability engineering, platform engineering, or related disciplines.
- Hands-on experience designing, implementing, and operating modern observability platforms.
- Strong knowledge of metrics, logs, traces, telemetry pipelines, SLOs, and SLIs.
- Experience supporting applications, infrastructure, containers, cloud services, distributed systems, and hybrid environments using AWS and/or Azure.
- Hands-on automation and infrastructure-as-code experience with Python, Terraform, Ansible, or Bash.
- Strong analytical, troubleshooting, problem-solving, communication, and stakeholder management skills.
Nice to have
- Experience in biotech, pharmaceutical, healthcare, or regulated environments such as GxP or HIPAA.
- Experience with AI observability tools such as Langfuse, Arize, Phoenix, or LangSmith.
- Experience monitoring AI agents, LLM applications, retrieval systems, or agentic workflows.
- Experience with large-scale logging strategies and integrations with PagerDuty, ServiceNow, or Jira.
- Relevant AWS, Azure, Kubernetes, observability, or related certifications.
Culture & Benefits
- Healthcare and voluntary benefit programs.
- Fitness, mindfulness, and mental health support.
- Family-building support, including fertility, adoption, and surrogacy benefits.
- Paid vacation, bank holidays, volunteer days, sabbatical, global recharge days, and year-end shutdown.
- Savings and investment programs, plus location-specific benefits.
- In-person collaboration supported by a 70/30 office-based work model.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →