Назад
Company hidden
3 часа назад

AI Engineer

120 000 - 170 000$
Формат работы
remote (только USA)/hybrid
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Engineer (LLM Evaluation and Responsible AI): Building and improving production AI systems for public safety agencies with an accent on automated evaluation pipelines, prompt engineering, statistical analysis, and model reliability. Focus on designing regression frameworks, analyzing model failures, benchmarking LLM changes, and monitoring AI quality in real-world emergency communications applications.

Location: Remote within the United States; hybrid opportunity if based in Denver, Colorado

Salary: $120K–$170K annually, plus bonus and benefits

Company

hirify.global develops CommsCoach, an AI-powered platform that supports 9-1-1 and emergency communications centers through automated quality assurance, training, and real-time call evaluation.

What you will do

  • Design, build, and maintain automated evaluation pipelines for production LLM applications.
  • Develop prompt engineering strategies, compare LLMs, and measure changes using quantitative evaluation methods.
  • Build offline evaluation datasets and regression testing frameworks to track AI performance over time.
  • Analyze production AI behavior with Python, SQL, and statistical techniques, including error analysis and experiment design.
  • Create dashboards and reporting for AI quality, reliability, and performance metrics.
  • Partner with engineering, product, data science, and data engineering teams to deploy and monitor improvements and establish Responsible AI practices.

Requirements

  • Must pass FBI fingerprinting and background checks in multiple states.
  • U.S. citizenship is strongly preferred.
  • 3+ years of experience in software engineering, machine learning, data science, or a related technical field.
  • Strong Python development and SQL skills, including analysis of large datasets.
  • Experience designing evaluation metrics, interpreting AI model performance, and applying statistical methods such as hypothesis testing and experiment design.
  • Experience supporting production LLM or generative AI applications, prompt engineering, and systematic prompt evaluation.

Nice to have

  • Experience with AI evaluation or observability platforms such as Langfuse, LangSmith, MLflow, or Label Studio.
  • Experience with AWS services including Bedrock, Lambda, S3, Glue, or SageMaker.
  • Experience building dashboards with Tableau, Sisense, Power BI, or similar tools.
  • Knowledge of Responsible AI principles and evaluation methodologies.

Culture & Benefits

  • Health, dental, and vision benefits.
  • Flexible time off.
  • Opportunity to build AI systems supporting first responders and emergency communications professionals.
  • Work on LLM technologies, Responsible AI, and technically challenging public safety problems.
  • Collaborate across AI, engineering, product, and data science functions.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →