Назад
Company hidden
6 дней назад

Senior Software Engineer, ML Infrastructure (AI)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Software Engineer, ML Infrastructure (AI): Building and operating ML infrastructure for conversational AI experiences, including ranking pipelines, model delivery, evaluation platforms, LLM agents, caching, and observability with an accent on reliability, latency, quality, and cost. Focus on designing distributed production systems, evaluating LLM behavior, diagnosing cross-service failures, and improving model lifecycle operations at scale.

Location: Cambridge, United States; hybrid working with office attendance generally required Monday through Thursday and flexible remote work on Fridays

Company

hirify.global operates a TV streaming platform connecting consumers, content publishers, advertisers, and the broader television ecosystem.

What you will do

  • Own fulfilment-ranking pipelines from training orchestration and quality gates through deployment and online validation.
  • Build offline evaluation platforms combining LLM-as-judge harnesses with deterministic answer-quality checks.
  • Develop LLM-agent capabilities including tool routing, retrieval, guardrails, and answer caching.
  • Design caching and observability systems to improve latency, quality, reliability, and per-request cost visibility.
  • Diagnose complex cross-service failures and improve the operability of distributed ML and LLM systems.
  • Collaborate with machine-learning, product, data, and platform partners to turn ambiguous problems into measurable outcomes.

Requirements

  • Strong production software-engineering experience designing, testing, operating, and debugging distributed services.
  • Experience owning ML infrastructure or production model delivery across training, evaluation, versioning, deployment, monitoring, and rollback.
  • Experience with latency, resilience, observability, cost optimization, caching, and production operations.
  • Experience with LLMs, embeddings, semantic search, retrieval, or agent systems, including quality evaluation and failure-mode control.
  • Commitment to automation, CI/CD, code quality, and evidence-led engineering decisions.
  • Degree in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience.

Nice to have

  • Experience with AWS or GCP, Kubernetes, SQL warehouses, Trino, Presto, Spark, Kafka, or other streaming systems.
  • Experience with Go, Python, or Node.js.
  • Experience using AI-assisted engineering tools such as coding harnesses, MCP servers, custom skills, or agent frameworks.

Culture & Benefits

  • Inclusive, collaborative environment focused on pragmatic problem-solving and delivering working solutions.
  • Global mental-health and financial-wellness support.
  • Local benefits may include healthcare, life and disability coverage, commuter benefits, retirement options, and statutory leave.
  • Employees receive time off in accordance with local leave policies and personal needs.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →