QA Engineer (Gen AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
QA Engineer (Gen AI) (LLM evaluation and document AI): Building automated quality assurance systems for AI agents, structured extraction, and document-processing pipelines with an accent on regression testing, data quality, and evaluation metrics. Focus on identifying LLM failure modes, validating ground-truth datasets, calibrating automated scoring, and testing multi-page business documents across AWS-based infrastructure.
Location: Austin, Texas, United States; applicants must be authorized to work for any U.S. employer. Visa sponsorship is not available.
Company
is an AI-native software platform connecting U.S.-based manufacturers with critical suppliers for secure and resilient supply chains serving commercial and Department of Defense customers.
What you will do
- Design and run automated regression suites for LLM and AI-agent evaluation.
- Identify, track, and mitigate hallucinations, bias, factual inconsistencies, logical errors, and other model failure modes.
- Create data-quality checks and maintain versioned ground-truth datasets for document parsing and structured extraction.
- Evaluate model outputs field by field across PDFs, scans, spreadsheets, and repeated or nested records using tolerance-aware scoring.
- Build monitoring tools and dashboards for LLM quality, validate automated scoring against human judgment, and turn production failures into regression cases.
- Collaborate with ML engineers, data scientists, labeling specialists, and product teams to define quality benchmarks and testable field requirements.
Requirements
- 3+ years of software testing and quality assurance experience, including 2+ years focused on ML evaluation, NLP, LLMs, or VLMs.
- Strong understanding of LLM data-quality challenges, failure modes, and automated AI/ML testing.
- Python experience with PyTest, Hypothesis, or similar testing frameworks.
- Experience with LLM evaluation metrics and tools such as DeepEval, MLflow, or LangSmith.
- Experience creating or using labeled evaluation datasets and applying precision, recall, F1, exact and fuzzy matching, numeric tolerance, and record-alignment techniques.
- U.S. work authorization is required; visa sponsorship is not available.
Nice to have
- SQL proficiency, including seeding test data across PostgreSQL environments.
- Experience with RAG, RAGAS, HITL evaluation, LLM APIs, MLOps, or ML CI/CD pipelines.
- Experience evaluating document AI and OCR pipelines, including layout, tables, scans, and low-quality source material.
- Familiarity with AWS, Kubernetes, Datadog, .NET, Linux, EF Core, DDL, or controlled model and prompt comparison studies.
Culture & Benefits
- Contract opportunity within an AI-native product company supporting manufacturing, commercial, and defense customers.
- Full-time employee benefits may include medical, dental, and vision coverage.
- Paid time off and company holidays are provided to eligible full-time employees.
- 401(k) matching is available to eligible full-time employees.
- Equal opportunity employment is provided across protected characteristics.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →