Назад
Company hidden
13 дней назад

ML Research Engineer (AI)

235 000 - 295 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
ML Research Engineer (AI): Building evaluation and experimentation systems for a scientific agent, with an accent on reproducible replay, versioned datasets, trajectory analysis, and decision-quality metrics. Focus on designing controlled experiments, analyzing failure modes, and shipping production improvements to planning, tool use, and recovery policies.

Location: Brooklyn, NY, United States; on-site

Salary: $235,000–$295,000 per year

Company

Develops scientific agents that determine which experiments to run next and connect agent behavior with activity in the Workcell.

What you will do

  • Design reproducible replay systems for reconstructing historical agent runs and comparing models, prompts, and policies.
  • Build versioned evaluation datasets, evaluation harnesses, and metrics for decision quality, task success, error recovery, cost, and latency.
  • Instrument agent trajectories, analyze failure modes, and convert recurring failures into regression coverage.
  • Prototype, test, and ship control policies for revising approaches, pivoting, escalating to scientists, and terminating tasks.
  • Own experiment infrastructure, including data pipelines, artifact and dataset versioning, experiment tracking, and comparison tooling.
  • Partner with scientists and platform engineers to connect agent behavior with Workcell outcomes.

Requirements

  • 4–8 years of experience building production ML systems, research infrastructure, or data-intensive backends.
  • Excellent Python and strong software engineering practices, including testing, typing, packaging, code review, and CI.
  • Direct experience with ML evaluation and experimentation, including offline evaluation harnesses, dataset splits, metric design, and analysis of noisy or underpowered results.
  • Experience with schema design, data versioning, orchestration, and object storage.
  • Understanding of reproducible experimentation through pinned environments, seeded runs, versioned artifacts, and rerunnable results.
  • Ability to turn ambiguous behavior into controlled experiments and harden successful prototypes into reliable production systems.

Nice to have

  • Experience with sequential decision-making, including bandits, Bayesian optimization, reinforcement learning, or off-policy evaluation.
  • Experience with Langfuse, MLflow, Weights and Biases, DVC, or similar experimentation and observability tools.
  • Background in scientific computing, materials, chemistry, instrument data, lab automation, robotics, or self-driving labs.
  • Experience with large-scale inference, GPU scheduling, distributed execution, or observability tools such as OpenTelemetry, Prometheus, Grafana, or Datadog.

Culture & Benefits

  • Engineering-first, ML-strong, and research-capable environment.
  • Close collaboration with scientists and platform engineers.
  • Opportunity to contribute research ideas and take improvements from prototype through production.
  • Equal opportunity workplace.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →