7 часов назад
Member of Technical Staff, Evaluations (AI)
175 000 - 250 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff, Evaluations (AI) (Python/LLM): Building reliable evaluation harnesses, automated reporting pipelines, analytical methods, and calibrated scoring systems for AI models and agents working on electrical and hardware engineering tasks with an accent on experimental rigor, synthetic datasets, and failure-mode analysis. Focus on isolating regressions through ablations, analyzing agent trajectories, validating LLM judges, and measuring complex engineering capabilities.
Location: New York, NY; San Francisco, or Los Angeles, United States
Salary: $175,000–$250,000 gross annual base salary, plus equity and benefits.
Company
Builds Atlas, an AI platform that applies physics-grounded intelligence to verify, debug, and optimize hardware across its lifecycle.
What you will do
- Own the design, reliability, and automation of evaluation harnesses and recurring reporting pipelines.
- Design ablations to measure the effects of model, harness, tool, data, and context changes on performance, cost, and runtime.
- Analyze agent trajectories, cluster failure modes, and identify root causes such as context limitations, tool inefficiency, and gaps in engineering understanding.
- Develop verifiable scoring systems and calibrate LLM-based judges against expert assessments.
- Partner with Operations and Electrical Engineering to create synthetic datasets covering demanding hardware-engineering tasks.
- Translate customer and partner definitions of quality into measurable evaluation artifacts and communicate results clearly.
Requirements
- Strong Python skills for reliable research or production infrastructure, data pipelines, and tooling.
- Experience designing evaluations, benchmarks, or metrics for ML systems, ideally large language models or agentic systems.
- Hands-on experience with LLM agents, harnesses, tool use, and trajectory analysis.
- Strong knowledge of experimental design and statistics, with sound judgment about trustworthy metrics.
- Excellent written and verbal communication for technical and non-technical audiences.
- Ability to work from one of the listed United States locations and meet applicable U.S. export-control authorization requirements without sponsorship for an export license.
Nice to have
- Electrical or hardware engineering experience, including circuit design, RF/EM, signal integrity, PCB layout, or test and measurement.
- Experience calibrating LLM-as-judge systems, curating synthetic datasets, or evaluating multimodal and spatiotemporal reasoning.
- Experience building trusted dashboards and reporting, or working with customers and external technical partners.
Culture & Benefits
- Full-time role reporting directly to the CTO.
- Medical, vision, and dental insurance with monthly premiums fully covered for employees and dependents.
- 401(k) retirement plan and unlimited PTO.
- Daily lunch provided through local restaurants.
- Relocation support provided.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 часов назад
Senior AI Engineer (LLM Evals)
250 000 - 300 000$
7 часов назад
Staff Engineer — Agentic AI
160 000 - 250 000$
3 дня назад
Backend/ML Engineer (AI)
13 000 - 150 000$
3 дня назад
Principal Software Engineer (AI)
160 200 - 425 000$
3 дня назад
Staff AI/ML Engineer (LLMs)
250 000 - 350 000$
Anthropic
4 дня назад
AI Engineer (GTM Claudification)
320 000 - 405 000$