5 дней назад
Senior Observability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Observability Engineer (AI): Building and governing an AI-enabled enterprise observability platform across applications, infrastructure, cloud services, networks, containers, databases, and LLM workloads with an accent on telemetry governance, resilient architecture, intelligent monitoring, and cost efficiency. Focus on designing telemetry and logging pipelines, implementing AI observability for agents and LLM applications, automating remediation, and improving incident detection and reliability.
Location: Cambridge, Massachusetts, United States, and Warsaw, Poland; 70% in-office under a hybrid work model
Company
is a biotechnology company developing mRNA medicines and supporting global operations through technology and business services.
What you will do
- Own the enterprise observability platform, including strategy, governance, capacity planning, telemetry coverage, cost optimization, and roadmap development.
- Design and operate scalable observability architectures for applications, infrastructure, cloud services, networks, containers, databases, distributed systems, and AI/LLM workloads.
- Build telemetry pipelines for metrics, traces, and logs across hybrid cloud and on-premises environments using OpenTelemetry, Prometheus, Grafana, VictoriaMetrics, Loki, and related technologies.
- Develop AI observability capabilities for agents, LLM applications, and agentic workflows, including prompt, response, latency, error, reliability, and cost monitoring.
- Automate operational processes and remediation with Python, Terraform, Ansible, CI/CD, and infrastructure-as-code practices.
- Integrate observability with PagerDuty, ServiceNow, Jira, and security workflows while supporting incident response, root cause analysis, dashboards, documentation, and training.
Requirements
- 7+ years of experience in site reliability engineering, observability engineering, platform engineering, or related technical disciplines.
- Hands-on experience designing, implementing, and operating modern observability platforms and working with metrics, logs, traces, telemetry pipelines, SLOs, and SLIs.
- Experience supporting applications, infrastructure, containers, cloud services, distributed systems, and hybrid environments using AWS and/or Azure.
- Experience with observability technologies such as OpenTelemetry, Prometheus, Grafana, VictoriaMetrics, Elastic, Datadog, or Dynatrace.
- Hands-on automation and infrastructure-as-code experience with Python, Terraform, Ansible, or Bash.
- Strong analytical, troubleshooting, communication, and stakeholder management skills; applications must be submitted in English with an English CV.
Nice to have
- Experience in biotech, pharmaceutical, healthcare, or other regulated environments, including GxP or HIPAA.
- Experience with AI observability tools such as Langfuse, Arize, Phoenix, or LangSmith.
- Experience monitoring AI agents, LLM applications, retrieval systems, or agentic workflows.
- Experience with enterprise logging strategies and integrations with PagerDuty, ServiceNow, or Jira.
- Relevant AWS, Azure, Kubernetes, or observability certifications.
Culture & Benefits
- Hybrid 70/30 work model emphasizing in-office collaboration, teamwork, and direct mentorship.
- Healthcare and voluntary benefit programs.
- Fitness, mindfulness, and mental health resources.
- Family-building support covering fertility, adoption, and surrogacy.
- Paid vacation, bank holidays, volunteer days, sabbatical, global recharge days, and year-end shutdown.
- Savings and investment programs with location-specific benefits.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
11 дней назад
Engineering Manager (AI)
210 000 - 300 000$
5 дней назад
Senior Platform Engineer (AI)
6 дней назад
Senior Platform Engineer (AI)
180 000 - 240 000$
6 дней назад
Sr. Specialist - Platform Operations (AI & Agentic Systems)
91 000 - 129 000$
11 дней назад
Principal Member of Technical Staff (AI Infrastructure)
200 000 - 350 000$
5 дней назад
Platform Engineer (AI)
160 000$