6 дней назад
Senior Software Engineer, ML Infrastructure (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Software Engineer, ML Infrastructure (AI): Building and operating ML infrastructure for conversational AI experiences, including ranking pipelines, model delivery, evaluation platforms, LLM agents, caching, and observability with an accent on reliability, latency, quality, and cost. Focus on designing distributed production systems, evaluating LLM behavior, diagnosing cross-service failures, and improving model lifecycle operations at scale.
Location: Cambridge, United States; hybrid working with office attendance generally required Monday through Thursday and flexible remote work on Fridays
Company
operates a TV streaming platform connecting consumers, content publishers, advertisers, and the broader television ecosystem.
What you will do
- Own fulfilment-ranking pipelines from training orchestration and quality gates through deployment and online validation.
- Build offline evaluation platforms combining LLM-as-judge harnesses with deterministic answer-quality checks.
- Develop LLM-agent capabilities including tool routing, retrieval, guardrails, and answer caching.
- Design caching and observability systems to improve latency, quality, reliability, and per-request cost visibility.
- Diagnose complex cross-service failures and improve the operability of distributed ML and LLM systems.
- Collaborate with machine-learning, product, data, and platform partners to turn ambiguous problems into measurable outcomes.
Requirements
- Strong production software-engineering experience designing, testing, operating, and debugging distributed services.
- Experience owning ML infrastructure or production model delivery across training, evaluation, versioning, deployment, monitoring, and rollback.
- Experience with latency, resilience, observability, cost optimization, caching, and production operations.
- Experience with LLMs, embeddings, semantic search, retrieval, or agent systems, including quality evaluation and failure-mode control.
- Commitment to automation, CI/CD, code quality, and evidence-led engineering decisions.
- Degree in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience.
Nice to have
- Experience with AWS or GCP, Kubernetes, SQL warehouses, Trino, Presto, Spark, Kafka, or other streaming systems.
- Experience with Go, Python, or Node.js.
- Experience using AI-assisted engineering tools such as coding harnesses, MCP servers, custom skills, or agent frameworks.
Culture & Benefits
- Inclusive, collaborative environment focused on pragmatic problem-solving and delivering working solutions.
- Global mental-health and financial-wellness support.
- Local benefits may include healthcare, life and disability coverage, commuter benefits, retirement options, and statutory leave.
- Employees receive time off in accordance with local leave policies and personal needs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
Senior Software Engineer, Machine Learning Platform (AI)
187 000 - 259 000$
10 дней назад
Senior ML Engineer (MLOps)
143 000 - 197 000$
9 дней назад
Principal ML Platform Engineer (AI)
175 000 - 325 000$
11 дней назад
Senior Forward Deployed Engineer (AI Studio)
8 дней назад
DevOps Engineer - Senior Vice President (MLOps/AI)
180 000 - 230 000$
6 дней назад
Machine Learning Cloud Infrastructure Engineer (AI)
150 000 - 175 000$