13 дней назад
ML Research Engineer (AI)
235 000 - 295 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
ML Research Engineer (AI): Building evaluation and experimentation systems for a scientific agent, with an accent on reproducible replay, versioned datasets, trajectory analysis, and decision-quality metrics. Focus on designing controlled experiments, analyzing failure modes, and shipping production improvements to planning, tool use, and recovery policies.
Location: Brooklyn, NY, United States; on-site
Salary: $235,000–$295,000 per year
Company
Develops scientific agents that determine which experiments to run next and connect agent behavior with activity in the Workcell.
What you will do
- Design reproducible replay systems for reconstructing historical agent runs and comparing models, prompts, and policies.
- Build versioned evaluation datasets, evaluation harnesses, and metrics for decision quality, task success, error recovery, cost, and latency.
- Instrument agent trajectories, analyze failure modes, and convert recurring failures into regression coverage.
- Prototype, test, and ship control policies for revising approaches, pivoting, escalating to scientists, and terminating tasks.
- Own experiment infrastructure, including data pipelines, artifact and dataset versioning, experiment tracking, and comparison tooling.
- Partner with scientists and platform engineers to connect agent behavior with Workcell outcomes.
Requirements
- 4–8 years of experience building production ML systems, research infrastructure, or data-intensive backends.
- Excellent Python and strong software engineering practices, including testing, typing, packaging, code review, and CI.
- Direct experience with ML evaluation and experimentation, including offline evaluation harnesses, dataset splits, metric design, and analysis of noisy or underpowered results.
- Experience with schema design, data versioning, orchestration, and object storage.
- Understanding of reproducible experimentation through pinned environments, seeded runs, versioned artifacts, and rerunnable results.
- Ability to turn ambiguous behavior into controlled experiments and harden successful prototypes into reliable production systems.
Nice to have
- Experience with sequential decision-making, including bandits, Bayesian optimization, reinforcement learning, or off-policy evaluation.
- Experience with Langfuse, MLflow, Weights and Biases, DVC, or similar experimentation and observability tools.
- Background in scientific computing, materials, chemistry, instrument data, lab automation, robotics, or self-driving labs.
- Experience with large-scale inference, GPU scheduling, distributed execution, or observability tools such as OpenTelemetry, Prometheus, Grafana, or Datadog.
Culture & Benefits
- Engineering-first, ML-strong, and research-capable environment.
- Close collaboration with scientists and platform engineers.
- Opportunity to contribute research ideas and take improvements from prototype through production.
- Equal opportunity workplace.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →