9 дней назад
Staff Observability Engineer (AI)
290 000 - 375 000PLN
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Observability Engineer (AI): Designing observability strategy, telemetry pipelines, and reliability solutions for customer-facing products, platform services, data pipelines, and AI-powered systems with an accent on OpenTelemetry, Amazon Bedrock, and end-to-end tracing. Focus on measuring AI quality and cost, building SLOs and operational tooling, and solving reliability challenges across distributed and agentic workflows.
Location: Remote Poland
Base salary: 290,000–375,000 PLN annually, with potential performance-based bonus and equity options.
Company
provides a real-time Customer Data Platform that helps organizations unify customer data and support personalized, privacy-conscious experiences and enterprise AI strategies.
What you will do
- Lead observability design, telemetry schemas, and OpenTelemetry pipeline architectures across products, platform services, data pipelines, and internal tools.
- Define service-level indicators, objectives, error budgets, and production-readiness requirements with engineering and product teams.
- Architect AI observability for Amazon Bedrock and agentic workflows, including model selection, latency, retries, token usage, tool execution, and cost attribution.
- Establish traceability across user requests, prompts, model calls, retrieved context, downstream services, and final responses while maintaining privacy and security controls.
- Build dashboards, alerts, SLOs, runbooks, and operational views for incident diagnosis, performance optimization, and cost control.
- Participate in an on-call rotation of approximately 20% and lead failure injection, production testing, and capacity planning.
Requirements
- 6+ years of experience in Site Reliability, Observability, or Platform Engineering supporting 24x7x365 production systems.
- Deep experience with OpenTelemetry and observability platforms such as Datadog, Sumo Logic, Prometheus, or Grafana.
- Experience with distributed systems, APIs, event-driven architectures, asynchronous workflows, and AI/ML or GenAI instrumentation.
- Experience with Amazon Bedrock or similar platforms, including model usage, latency, throttling, token metrics, and cost attribution.
- Proficiency in Java, Python, or Go, plus strong AWS expertise covering networking, IAM, security, and service quotas.
- Experience with IaC, CI/CD, containers, data modeling, telemetry pipelines, cardinality management, retention, responsible AI, and data privacy controls.
Nice to have
- Familiarity with agentic workflows, prompt engineering, vector databases such as Neptune, RAG architectures, LangChain, LlamaIndex, or SageMaker.
Culture & Benefits
- Remote-first working with support for a home office setup.
- Paid time off, extended paid parental leave, and company holidays.
- Health and wellness programs and competitive benefits.
- Professional development through on-demand courses and leadership programs.
- Volunteer support with 15 hours of paid work time annually.
- New hire equity grants and opportunities to participate in 's ownership program.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
11 дней назад
Staff Site Reliability Engineer (AI/ML)
10 дней назад
Site Reliability Engineer (AI)
200 000 - 400 000$
14 дней назад
Senior Site Reliability Engineer (AI)
185 500 - 232 000$
9 часов назад
Staff Service Reliability and Operational Intelligence Engineer (AI Ops)
152 000 - 228 000$
Vallettasoftware Software Development
9 дней назад
Senior DevOps / SRE Support Engineer (Kubernetes)
5 000 - 5 500$
Valletta.Software | AI-Care
3 часа назад
Senior DevOps / SRE Support Engineer (LATAM)
5 000 - 5 500$