6 дней назад
Manager, AI Ops Site Reliability Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Manager, AI Ops Site Reliability Engineer (LLM/Agent Engineering): Designing AI agent workflows and evaluation systems for operational incident investigation, with an accent on prompt and context engineering, retrieval, guardrails, and measurable accuracy. Focus on building ground-truth datasets, regression tests, hallucination detection, and agent-assisted operational workflows that reduce toil and MTTR.
Location: Thessaloniki, Chortiatis, Greece. Work arrangement: Hybrid. Periodic international and domestic travel required, less than 10%.
Company
is a global pharmaceutical company developing medicines and digital solutions to improve health outcomes and transform drug discovery and development.
What you will do
- Design AI agent workflows for incident intake and investigation, including prompts, context, retrieval, tool selection, and decision logic.
- Build and own evaluation harnesses using historical incidents, ground-truth datasets, accuracy metrics, regression tests, and hallucination detection.
- Define reasoning guardrails, confidence thresholds, approval rules, and human-deference conditions for operations agents.
- Convert operational runbooks and investigations into retrievable, machine-usable knowledge in collaboration with domain experts.
- Redesign operational workflows with agent assistance, pilot them, and measure toil and MTTR reduction.
- Continuously tune agent quality against production results while collaborating with SRE, operations, platform, and AI Ops teams.
Requirements
- Bachelor's degree in a technical field or equivalent practical experience.
- 4+ years of experience in software, data, or ML/AI engineering, including production delivery of LLM or agent-based applications.
- Experience with prompt and context engineering and structured AI evaluations, including ground-truth datasets, metrics, and regression testing.
- Working knowledge of agent frameworks and tooling such as Amazon Bedrock AgentCore or LangGraph, RAG, knowledge-base design, and strong Python skills.
- IT operations literacy and the ability to assess the plausibility of agent outputs in operational workflows.
- Strong analytical and communication skills, with experience collaborating with SREs and presenting accuracy results.
Nice to have
- Experience with LLM-as-judge evaluation.
- Familiarity with ServiceNow, Dynatrace, or observability data.
- Experience working in an agile environment.
Culture & Benefits
- Flexible workplace culture designed to support work-life harmony.
- Collaborative environment focused on continuous improvement and digital transformation.
- Periodic domestic and international travel of less than 10%.
- Inclusive workplace with support for reasonable disability accommodations.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Lead AI Engineer
9 дней назад
Senior AI Data Scientist, Agentic Automation (Marketing)
10 дней назад
Software Engineer II – ΑΙ Meshing
11 дней назад
Senior AI Data Scientist — Agentic Process Development
11 дней назад
Sr. Infrastructure Engineer (AWS/Kubernetes)
10 дней назад