4 дня назад
Lead Machine Learning Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Lead Machine Learning Engineer (AI) (LLM/Agentic Systems): Building and operating an evaluation platform for agentic AI systems, spanning offline benchmarks, human review, LLM-as-judge workflows, and production monitoring with an accent on quality, safety, reproducibility, and task performance. Focus on designing data-intensive evaluation infrastructure, measuring hallucinations and goal completion, and driving scalable architecture across research, product, and platform teams.
Location: Mountain View or New York; hybrid with 10–12 days of in-office presence per month
Company
develops agentic AI systems and customer-support solutions powered by modern NLP, ML, and LLM technologies.
What you will do
- Own the technical roadmap and architecture for an evaluation platform covering offline benchmarking and online production monitoring of agentic and LLM-based systems.
- Design evaluation methodologies including golden and regression test sets, human-in-the-loop reviews, LLM-as-judge workflows, and automated metrics for task success, safety, and hallucinations.
- Build annotation and labeling pipelines, dataset versioning, data-quality checks, and reproducible benchmarking infrastructure.
- Create tooling that enables research and product teams to run and interpret experiments independently.
- Partner with Research, Product, and Platform teams to productize experiments into robust AI solutions.
- Mentor engineers, lead design reviews, communicate platform health and coverage, and guide expectations for model and agent releases.
Requirements
- Deep hands-on experience building and operating evaluation systems for modern ML, LLM, or agentic systems.
- Experience leading the technical direction of a project or small team, including architecture, design reviews, and long-term system ownership.
- Strong architectural skills and experience designing complex, data-intensive production systems.
- Production experience with Python and AWS, Kubernetes, and/or Docker.
- Experience with ML evaluation data pipelines, annotation workflows, dataset versioning, quality control, and reproducible benchmarking.
- Bachelor’s degree in Computer Science or a related field, plus demonstrated mentorship of junior and mid-level engineers.
Nice to have
- Experience building and evaluating agentic systems at scale.
- Experience with voice or audio quality evaluations.
- Production experience with LLM-centric inference, orchestration, evaluation, or monitoring services.
- Experience with large-scale ML experimentation, benchmarking, simulation, or conversational customer-support AI.
- Experience with model inference optimization, AWS, CI/CD, Kafka, or Athena.
Culture & Benefits
- Hybrid work structure balancing flexibility with in-person collaboration.
- Close collaboration across Research, Product, Platform, and Engineering.
- Opportunities to contribute to technical discussions, knowledge sharing, and architectural alignment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →