Назад
Company hidden
20 часов назад

Assessment Design Lead (AI)

100 000 - 140 000$
Формат работы
remote (Global)
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US/Japan/SK +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Assessment Design Lead (AI): Designing reliable speaking assessments for an AI-powered language-learning platform with an accent on construct definition, CEFR alignment, rubrics, and validity. Focus on building validity evidence, auditing fairness across learner populations, and partnering with ML Engineers on automated scoring and calibration.

Location: Remote

Salary: $100K–$140K plus equity

Company

hirify.global is an AI-powered language-learning company building a conversation-first tutor with instant feedback and structured lessons across 15+ languages.

What you will do

  • Define assessment goals, frequency, and measurement approaches across curriculum mastery, proficiency, and placement tests, with an initial focus on the Proficiency Test.
  • Translate fluency, pronunciation, grammar, and task achievement into measurable constructs aligned with CEFR or comparable hirify.globaling standards.
  • Create item blueprints, rubrics, rater guidelines, and specifications that support both item writing and ML-based grading.
  • Own content validity and the quality bar, including score meaning, mastery thresholds, bias audits, and resistance to gaming.
  • Design and run validity studies against CEFR-anchored exams and expert human ratings.
  • Partner with Product and ML Engineering on automated scoring, calibration, feedback generation, and model evaluation.

Requirements

  • 4+ years designing rubrics, blueprints, and item specifications for a shipped language assessment product or equivalent psychometric work.
  • Deep knowledge of CEFR, ACTFL, IELTS, TOEFL, or comparable language proficiency frameworks.
  • Ability to assess fairness across L1 backgrounds and accents, including differential item functioning.
  • Practical understanding of inter-rater reliability, classical test theory, and basic IRT concepts.
  • Ability to translate qualitative constructs into precise requirements for ML scoring pipelines and make clear content-validity decisions.
  • Comfort operating in a 0-to-1 environment and using AI tools with sound judgment about output quality.

Nice to have

  • Background in speech or pronunciation science and L2-specific error taxonomies.
  • Familiarity with adaptive testing or IRT-adjacent concepts.
  • Experience at a large-scale language testing organization or high-rigor assessment environment.
  • Advanced degree in psychometrics, measurement, applied linguistics, SLA, or a related quantitative field.
  • Authored technical or validity reports, or published assessment research.

Culture & Benefits

  • Fully remote work with an asynchronous product experience and a distributed team.
  • Opportunity to collaborate with users and colleagues across San Francisco, Ljubljana, Seoul, Tokyo, and Taipei, with travel opportunities.
  • Work on a language-learning product serving more than 40 countries and 15+ languages.
  • Equity included in the compensation package.
  • Collaborative environment with significant ownership in a fast-growing Series C company.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →