Staff Data Scientist (LLM/Agent Evaluation)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
TL;DR
Staff Data Scientist (LLM/Agent Evaluation): Lead measurement and modeling for an agentic platform, defining how agent quality is measured and improved with an accent on evaluation, experimentation, and causal/statistical rigor. Focus on building scalable analytics and modeling foundations that prove every agent improvement with evidence from production traces and automated graders.
Location: Onsite in New York City, five days per week
Salary: $180,000β$215,000 (plus equity)
Company
builds an AI operating layer for the industrial supply chain.
What you will do
- Define how agent quality is defined, measured, and improved, and set the bar for statistical and scientific rigor.
- Design, build, and maintain metrics, models, and reporting for agent quality, reliability, adoption, and unit economics.
- Build evaluation and experimentation foundations using production traces, rubrics, automated graders, regression suites, and causal/statistical analysis.
- Model agent failure modes, tool-use patterns, and cost/latency to drive product improvements and growth.
- Architect scalable analytics and modeling infrastructure with data integrity, governance, and accessibility; oversee the data warehouse.
- Partner with engineering, product, and operations to deliver actionable, statistically grounded insights and self-service analytics.
Requirements
- 7+ years in data science, machine learning, applied statistics, or quantitative research; 2+ years modeling/measuring LLM- or agent-based systems in production.
- Strong proficiency in Python and common ML/statistics libraries (e.g., scikit-learn, PyTorch, pandas, statsmodels).
- Strong proficiency in SQL.
- Experience designing and analyzing experiments (A/B testing) and applying statistical inference or causal methods.
- Experience with LLM evaluation and observability tools (e.g., Langfuse, Braintrust) and building automated evaluators.
- Excellent data storytelling and cross-department collaboration skills.
Nice to have
- Experience with notebook tools (e.g., Jupyter, Hex, Hyperquery) and modern data stack tools (e.g., dbt).
- Experience building internal agents or MCP servers for analytics workflows.
- Experience fine-tuning/distilling or rigorously evaluating LLMs; applying causal inference and experimentation at scale.
Culture & Benefits
- Start-up equity and competitive salary.
- 100% paid health, dental, and vision coverage.
- NYC perks: dinner provided via DoorDash, free DashPass, stocked kitchen, commuter benefits, and Gympass.
- Additional benefits including One Medical membership, HSA via Optum, Talkspace, HealthAdvocate, and Teledoc Health.
- Onsite collaboration in New York City, five days per week.
Hiring process
- Interviews focused on measurement/evaluation, experimentation, and statistical modeling approach.
- Technical discussions around LLM/agent evaluation, automated graders, and production analytics foundations.
- Cross-functional conversations with engineering, product, and operations stakeholders.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β