5 дней назад
Evaluation Researcher (AI)
200 000 - 600 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Evaluation Researcher (AI): Designing rigorous evaluations for AI-agent populations, predictive systems, and end-to-end behavioral simulations with an accent on measurement validity, statistical analysis, and real-world decision quality. Focus on building reusable evaluation infrastructure, testing calibration and behavioral fidelity, identifying failures hidden by aggregate metrics, and communicating uncertainty in evidence-based conclusions.
Location: New York City; in-person work five days a week in the office. Candidates must be located in the New York metropolitan area or be open to relocation.
Salary: $200,000–$600,000 per year, plus equity.
Company
builds simulations of human behavior using populations of AI agents to help companies and institutions test consequential decisions.
What you will do
- Own evaluation and measurement areas covering population construction, predictive systems, individual agent behavior, group dynamics, or end-to-end simulations.
- Define measurable constructs and decision criteria for realism, accuracy, calibration, usefulness, and decision quality.
- Design historical backtests, prospective studies, observational analyses, controlled experiments, and mixed-method evaluations.
- Build tests for profile coherence, population fidelity, prediction quality, longitudinal behavior, interaction effects, and end-to-end simulation outcomes.
- Develop reusable datasets, evaluation harnesses, libraries, graders, experiment schemas, leaderboards, and reporting systems.
- Inspect failures, quantify uncertainty, protect held-out evidence, and communicate positive, negative, null, and inconclusive findings.
Requirements
- Rigorous experience in ML evaluation, statistics, behavioral science, computational social science, economics, psychometrics, experimental design, or a related field.
- Experience building evaluations that influenced research direction, model capability, product decisions, or scientific conclusions.
- Strong understanding of observational data, sampling, statistical power, uncertainty, causal threats, selection, leakage, and condition shift.
- Ability to write code, analyze large datasets, build evaluation systems, and inspect individual examples alongside aggregate metrics.
- Ability to collaborate with system builders while maintaining independent judgment and to communicate technical evidence to technical and non-technical audiences.
- Willingness to work in person in New York City five days per week, or relocate to the New York metropolitan area.
Nice to have
- Experience in forecasting evaluation, econometrics, psychometrics, causal inference, survey methodology, experimental economics, measurement theory, or model behavior.
- Experience evaluating LLM agents, multi-agent systems, synthetic populations, recommender systems, probabilistic models, simulations, or decision-support tools.
- Experience with longitudinal records, transaction data, product analytics, field experiments, prospective studies, or operational validation.
- Experience building evaluation platforms, regression suites, experiment-tracking systems, shared datasets, model scorecards, or scientific reporting tools.
- Experience with automated graders, human evaluation, rubric design, inter-rater reliability, benchmark contamination, or adversarial evaluation.
Culture & Benefits
- Small, in-person team focused on urgency, high ownership, intellectual honesty, and following important work through to results.
- Emphasis on surfacing inconvenient evidence, revising conclusions, and reporting negative or inconclusive findings with the same care as positive results.
- Competitive base salary with equity participation.
- Comprehensive medical, vision, and dental coverage.
- Visa sponsorship and relocation support are available.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Machine Learning Engineer (AI)
250 000 - 270 000$
Anthropic
2 дня назад
Data Scientist, Product (AI)
285 000 - 380 000$
4 часа назад
Senior Machine Learning Researcher (Healthcare AI)
170 000 - 220 000$
7 дней назад
Senior Data Scientist (AI)
180 000 - 230 000$
Decagon
7 дней назад
Agent Data Scientist (AI)
165 000 - 215 000$
6 дней назад
Senior Data Scientist (AI)
150 000 - 250 000$