3 дня назад
LLM Operations Engineer (GenAI/RAG)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
LLM Operations Engineer (GenAI/RAG): Building and operating production-grade LLM, GenAI, and RAG platforms for clinical machine-learning workflows with an accent on observability, evaluation, deployment, and regulated medical-device operations. Focus on shipping agentic pipelines, designing cost and latency frameworks, integrating LLM Ops with MLOps, and ensuring reliable performance across proprietary sensor and device data.
Location: Hybrid in Stockholm, Berlin, or London
Company
develops preventative healthcare services that combine proprietary non-invasive technology with direct clinical care and data-driven insights.
What you will do
- Own the operational lifecycle of LLM, GenAI, and RAG systems across deployment, monitoring, evaluation, and iteration.
- Set up MLflow Tracing observability for prompts, tool calls, retrievals, latency, and cost across production pipelines.
- Build an evaluation suite with LLM judges, custom scorers, and human feedback through review applications.
- Ship RAG or agentic pipelines with prompt and application versioning, A/B testing, safe rollout, and rollback.
- Develop cost, latency, and GPU-capacity frameworks for third-party APIs, Databricks External Models, and self-hosted serving.
- Integrate LLM Ops with the existing MLOps platform to support reliable clinical workflows within a regulated medical-device QMS.
Requirements
- Strong MLOps fundamentals across experiment tracking, training, monitoring, and production platformisation.
- Fluent Python skills, core machine-learning knowledge, and experience delivering end-to-end production ML systems.
- Hands-on experience with LLM or GenAI applications, including prompt engineering, RAG, agents, or chains.
- Experience with orchestration frameworks such as LangChain, LangGraph, or comparable tools.
- Working knowledge of PyTorch, distributed systems, ML orchestration, vector databases, embedding models, and human-feedback loops.
- Ability to navigate complex systems spanning the medical domain, regulation, firmware, hardware, and non-deterministic AI outputs.
Nice to have
- Experience with Kubernetes and Terraform for infrastructure as code or self-hosted models.
- Exposure to LLM evaluation and observability, including tracing, LLM-as-judge, guardrails, and safety scorers.
- Databricks MLflow 3 for GenAI experience.
- Experience with agentic or AI-assisted coding workflows while retaining ownership of the resulting code.
Culture & Benefits
- Work in a tech-enabled, human-centred engineering environment.
- Collaborate closely with the MLOps team rather than operating as a separate LLM Ops track.
- Contribute to preventative healthcare and clinical workflows involving proprietary sensor and device data.
- Operate in a fast-moving tools and platform ecosystem.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →