Назад
Company hidden
2 дня назад

QA Engineer (Gen AI)

Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

QA Engineer (Gen AI) (LLM evaluation and document AI): Building automated quality assurance systems for AI agents, structured extraction, and document-processing pipelines with an accent on regression testing, data quality, and evaluation metrics. Focus on identifying LLM failure modes, validating ground-truth datasets, calibrating automated scoring, and testing multi-page business documents across AWS-based infrastructure.

Location: Austin, Texas, United States; applicants must be authorized to work for any U.S. employer. Visa sponsorship is not available.

Company

hirify.global is an AI-native software platform connecting U.S.-based manufacturers with critical suppliers for secure and resilient supply chains serving commercial and Department of Defense customers.

What you will do

  • Design and run automated regression suites for LLM and AI-agent evaluation.
  • Identify, track, and mitigate hallucinations, bias, factual inconsistencies, logical errors, and other model failure modes.
  • Create data-quality checks and maintain versioned ground-truth datasets for document parsing and structured extraction.
  • Evaluate model outputs field by field across PDFs, scans, spreadsheets, and repeated or nested records using tolerance-aware scoring.
  • Build monitoring tools and dashboards for LLM quality, validate automated scoring against human judgment, and turn production failures into regression cases.
  • Collaborate with ML engineers, data scientists, labeling specialists, and product teams to define quality benchmarks and testable field requirements.

Requirements

  • 3+ years of software testing and quality assurance experience, including 2+ years focused on ML evaluation, NLP, LLMs, or VLMs.
  • Strong understanding of LLM data-quality challenges, failure modes, and automated AI/ML testing.
  • Python experience with PyTest, Hypothesis, or similar testing frameworks.
  • Experience with LLM evaluation metrics and tools such as DeepEval, MLflow, or LangSmith.
  • Experience creating or using labeled evaluation datasets and applying precision, recall, F1, exact and fuzzy matching, numeric tolerance, and record-alignment techniques.
  • U.S. work authorization is required; visa sponsorship is not available.

Nice to have

  • SQL proficiency, including seeding test data across PostgreSQL environments.
  • Experience with RAG, RAGAS, HITL evaluation, LLM APIs, MLOps, or ML CI/CD pipelines.
  • Experience evaluating document AI and OCR pipelines, including layout, tables, scans, and low-quality source material.
  • Familiarity with AWS, Kubernetes, Datadog, .NET, Linux, EF Core, DDL, or controlled model and prompt comparison studies.

Culture & Benefits

  • Contract opportunity within an AI-native product company supporting manufacturing, commercial, and defense customers.
  • Full-time employee benefits may include medical, dental, and vision coverage.
  • Paid time off and company holidays are provided to eligible full-time employees.
  • 401(k) matching is available to eligible full-time employees.
  • Equal opportunity employment is provided across protected characteristics.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →