Назад
Company hidden
2 часа Π½Π°Π·Π°Π΄

Staff Data Scientist (LLM/Agent Evaluation)

180Β 000 - 215Β 000$
Π€ΠΎΡ€ΠΌΠ°Ρ‚ Ρ€Π°Π±ΠΎΡ‚Ρ‹
onsite
Π’ΠΈΠΏ Ρ€Π°Π±ΠΎΡ‚Ρ‹
fulltime
Π“Ρ€Π΅ΠΉΠ΄
senior
Английский
b2
Π‘Ρ‚Ρ€Π°Π½Π°
US
Вакансия ΠΈΠ· списка Hirify.GlobalВакансия ΠΈΠ· Hirify Global, списка ΠΌΠ΅ΠΆΠ΄ΡƒΠ½Π°Ρ€ΠΎΠ΄Π½Ρ‹Ρ… tech-ΠΊΠΎΠΌΠΏΠ°Π½ΠΈΠΉ
Для мэтча ΠΈ ΠΎΡ‚ΠΊΠ»ΠΈΠΊΠ° Π½ΡƒΠΆΠ΅Π½ Plus

ΠœΡΡ‚Ρ‡ & Π‘ΠΎΠΏΡ€ΠΎΠ²ΠΎΠ΄

Для мэтча с этой вакансиСй Π½ΡƒΠΆΠ΅Π½ Plus

ОписаниС вакансии

ВСкст:
/

TL;DR

Staff Data Scientist (LLM/Agent Evaluation): Lead measurement and modeling for an agentic platform, defining how agent quality is measured and improved with an accent on evaluation, experimentation, and causal/statistical rigor. Focus on building scalable analytics and modeling foundations that prove every agent improvement with evidence from production traces and automated graders.

Location: Onsite in New York City, five days per week

Salary: $180,000–$215,000 (plus equity)

Company

hirify.global builds an AI operating layer for the industrial supply chain.

What you will do

  • Define how agent quality is defined, measured, and improved, and set the bar for statistical and scientific rigor.
  • Design, build, and maintain metrics, models, and reporting for agent quality, reliability, adoption, and unit economics.
  • Build evaluation and experimentation foundations using production traces, rubrics, automated graders, regression suites, and causal/statistical analysis.
  • Model agent failure modes, tool-use patterns, and cost/latency to drive product improvements and growth.
  • Architect scalable analytics and modeling infrastructure with data integrity, governance, and accessibility; oversee the data warehouse.
  • Partner with engineering, product, and operations to deliver actionable, statistically grounded insights and self-service analytics.

Requirements

  • 7+ years in data science, machine learning, applied statistics, or quantitative research; 2+ years modeling/measuring LLM- or agent-based systems in production.
  • Strong proficiency in Python and common ML/statistics libraries (e.g., scikit-learn, PyTorch, pandas, statsmodels).
  • Strong proficiency in SQL.
  • Experience designing and analyzing experiments (A/B testing) and applying statistical inference or causal methods.
  • Experience with LLM evaluation and observability tools (e.g., Langfuse, Braintrust) and building automated evaluators.
  • Excellent data storytelling and cross-department collaboration skills.

Nice to have

  • Experience with notebook tools (e.g., Jupyter, Hex, Hyperquery) and modern data stack tools (e.g., dbt).
  • Experience building internal agents or MCP servers for analytics workflows.
  • Experience fine-tuning/distilling or rigorously evaluating LLMs; applying causal inference and experimentation at scale.

Culture & Benefits

  • Start-up equity and competitive salary.
  • 100% paid health, dental, and vision coverage.
  • NYC perks: dinner provided via DoorDash, free DashPass, stocked kitchen, commuter benefits, and Gympass.
  • Additional benefits including One Medical membership, HSA via Optum, Talkspace, HealthAdvocate, and Teledoc Health.
  • Onsite collaboration in New York City, five days per week.

Hiring process

  • Interviews focused on measurement/evaluation, experimentation, and statistical modeling approach.
  • Technical discussions around LLM/agent evaluation, automated graders, and production analytics foundations.
  • Cross-functional conversations with engineering, product, and operations stakeholders.

Π‘ΡƒΠ΄ΡŒΡ‚Π΅ остороТны: Ссли Ρ€Π°Π±ΠΎΡ‚ΠΎΠ΄Π°Ρ‚Π΅Π»ΡŒ просит Π²ΠΎΠΉΡ‚ΠΈ Π² ΠΈΡ… систСму, ΠΈΡΠΏΠΎΠ»ΡŒΠ·ΡƒΡ iCloud/Google, ΠΏΡ€ΠΈΡΠ»Π°Ρ‚ΡŒ ΠΊΠΎΠ΄/ΠΏΠ°Ρ€ΠΎΠ»ΡŒ, Π·Π°ΠΏΡƒΡΡ‚ΠΈΡ‚ΡŒ ΠΊΠΎΠ΄/ПО, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡ‚Π΅ этого - это мошСнники. ΠžΠ±ΡΠ·Π°Ρ‚Π΅Π»ΡŒΠ½ΠΎ ΠΆΠΌΠΈΡ‚Π΅ "ΠŸΠΎΠΆΠ°Π»ΠΎΠ²Π°Ρ‚ΡŒΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡˆΠΈΡ‚Π΅ Π² ΠΏΠΎΠ΄Π΄Π΅Ρ€ΠΆΠΊΡƒ. ΠŸΠΎΠ΄Ρ€ΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β†’