Назад
Company hidden
4 часа назад

QA Lead (AI)

160 000 - 300 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior/lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
QA Lead (AI): Building evaluation systems and end-to-end testing infrastructure for AI-generated pharmaceutical marketing content with an accent on non-deterministic model quality, compliance-critical validation, and Playwright automation. Focus on detecting model and prompt regressions, testing agent failure modes, hardening background jobs, and establishing CI/CD quality gates.

Location: New York City, on-site

Salary: $160,000–$300,000 per year

Company

hirify.global builds AI-powered commercial tools that help life sciences organizations create and launch pharmaceutical marketing campaigns.

What you will do

  • Design evaluation systems for probabilistic AI and LLM outputs using evidence-based quality scoring.
  • Build safeguards for model and prompt changes, including regression, cost, latency, and agent-failure detection.
  • Test compliance-critical workflows, including claims verification against approved source material and safety disclosures.
  • Build and maintain Playwright end-to-end coverage across login, content creation, review, backend services, and frontend flows.
  • Perform manual and exploratory testing, investigate production signals, and convert failures into regression tests.
  • Establish CI/CD quality gates, harden long-running background jobs, and define testing standards as the first QA hire.

Requirements

  • Strong Python experience and a track record of building test infrastructure that runs automatically in CI/CD.
  • Strong end-to-end and UI test automation experience, especially with Playwright.
  • Hands-on manual and exploratory QA experience, including edge-case testing and release sign-off.
  • Experience testing non-deterministic, ML, or LLM-based systems, or the ability to build this capability from scratch.
  • Knowledge of golden datasets, LLM-as-judge methods, human-label calibration, statistical quality analysis, and error analysis.
  • Independent ownership, clear communication, and the ability to explain technical risk to non-technical stakeholders.

Nice to have

  • Experience with LangChain, LangGraph, LangSmith, Langfuse, Arize Phoenix, Braintrust, RAGAS, Promptfoo, OpenAI Evals, or similar tools.
  • Comfort with TypeScript and React for meaningful UI testing and bug reproduction.
  • Adversarial or red-team testing experience, including safety and regulatory validation.
  • Experience with async Python services, task queues, regulated industries, or MLOps/LLMOps quality SLOs.

Culture & Benefits

  • Visa sponsorship is available for O-1, H-1B, and TN visas.
  • Relocation support is available.
  • Health, dental, and vision insurance.
  • Ground-floor equity opportunity and unlimited PTO.
  • 401(k) with a 4% company match, in-office lunches five days per week, and a professional growth stipend.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →