Назад
Company hidden
4 дня назад

Principal Data Scientist (NLP)

140 000 - 200 733$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Data Scientist (NLP): Building production NLP enrichment pipelines for scientific full-text, extracting entities, classifications, claims, and summaries at scale with an accent on comparing classical NLP, retrieval, and LLM-based approaches. Focus on evaluation design, production Python, high-volume LLM cost and concurrency management, reliable data pipeline orchestration, and agentic AI applications.

Location: Hoboken (HQ), New Jersey, USA

Salary: USD 140,000–200,733.33 per year

Company

hirify.global is a long-established global company focused on science, education, publishing, and transforming knowledge into practical impact.

What you will do

  • Design and build production NLP enrichment pipelines that extract entities, classifications, claims, and summaries from millions of scientific journal articles.
  • Compare classical NLP, embedding-based retrieval, LLM prompting, and fine-tuned smaller models using quality, evaluation, cost, and operational criteria.
  • Build golden evaluation sets with subject-matter experts and vendors, select metrics, and balance speed, quality, and cost.
  • Write production-quality Python and manage concurrency and costs for high-volume LLM workloads.
  • Collaborate with data engineers to orchestrate reliable, idempotent, retryable pipeline stages using Airflow and Dagster.
  • Contribute to agentic AI applications and collaborate with editors, product managers, and engineers on modeling and product decisions.

Requirements

  • Deep production experience with Python at scale, including clear tradeoff decisions between asyncio, threads, and queues.
  • Strong experience with modern and classical NLP, including LLMs, transformers, embeddings, retrieval, NER, classification, and sequence labeling.
  • Experience building evaluations, comparing modeling approaches, estimating costs, and responding to performance drift.
  • Track record of shipping production systems that deliver value to real users.
  • Ability to work with editors, product managers, engineers, subject-matter experts, and vendors.

Nice to have

  • Experience working with scientific or scholarly text.
  • Familiarity with AWS services including S3, Batch, Lambda, and SageMaker.
  • Experience with Parquet or Iceberg data lake patterns.
  • Experience running LLMs under production cost and latency budgets.
  • Exposure to agentic AI applications, tool use, multi-step reasoning, guardrails, and trajectory evaluation.

Culture & Benefits

  • Culture emphasizing bold ideas, diverse perspectives, learning, and measurable impact.
  • Meeting-free Friday afternoons for focused work and professional development.
  • Employee programs supporting community, continual learning, and internal mobility.
  • Comprehensive benefits package and competitive compensation.
  • Equal opportunity employer committed to reasonable accommodations.

Hiring process

  • Attach a resume or CV when applying.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →