3 дня назад
Evaluation Infrastructure Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Evaluation Infrastructure Engineer (AI): Building and operating the evaluation framework for computer-use agents across web applications, desktop software, and command-line environments with an accent on orchestration, reproducibility, observability, and cost-efficient scaling. Focus on integrating benchmarks, scheduling distributed evaluation workloads, instrumenting systems, and producing reliable results for model release decisions.
Location: Hybrid in Paris, with 3 office days per week; travel to London approximately every 4–6 weeks.
Company
H builds computer-use agents and the models behind them, offering a managed API and deploying agents into enterprise workflows.
What you will do
- Build and operate the evaluation framework for computer-use agents across web applications, desktop software, and command-line environments.
- Integrate benchmarks with researchers and forward deployed engineers, and shorten the path for adding new evaluations.
- Define benchmark standards and build checks that enforce reliable integrations.
- Improve scheduling, cluster utilization, observability, reproducibility, speed, and cost efficiency.
- Automate customer workflow measurements so results flow into evaluation harnesses and models.
- Own a major area such as run scaling, observability, or a related group of benchmarks.
Requirements
- 5+ years of backend development experience with production Python.
- Experience building test, QA, or evaluation tooling used by other engineering teams.
- Experience operating distributed systems on Kubernetes in a public cloud; AWS experience is particularly relevant.
- Experience shipping systems end to end, including REST or GraphQL APIs and integrations with external services.
- Knowledge of relational and non-relational databases and message queues such as SQS, RabbitMQ, or Kafka.
- Experience with metrics, tracing, and monitoring from the beginning of system development.
Nice to have
- Experience measuring LLM quality or building agents.
- Experience with Docker, virtual machines, Temporal, Dask, FastAPI, PostgreSQL, Grafana, or Datadog.
- Experience automating web or desktop software with Playwright, Selenium, or a custom browser extension.
- Experience setting engineering standards through code review, design review, or mentoring.
Culture & Benefits
- Close collaboration with researchers and forward deployed engineers.
- Regular exposure to customer workflows, from individual developers to large companies.
- Hybrid work in Paris with three office days per week.
- Approximately one London trip every 4–6 weeks.
- Competitive compensation package.
Hiring process
- 30-minute call with the Talent team.
- 60-minute technical challenge and 60-minute system design interview.
- 30-minute final conversation with the VP of Engineering, approximately 3.5 hours in total.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →