Назад
Company hidden
2 дня назад

Data Engineer II- Life Sciences

Формат работы
hybrid
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Data Engineer II- Life Sciences (Python/PySpark): Building and maintaining production data pipelines that ingest, transform, validate, and reconcile clinical trial data across diverse source formats with an accent on distributed processing, SQL analysis, and data quality. Focus on translating clinical domain rules into reliable pipeline logic, responding to enterprise SLA requirements, and improving observability for job health and data anomalies.

Location: New York, United States; Workplace: Hybrid

Company

hirify.global provides healthcare information and data products designed to improve doctor interactions, patient outcomes, and the drug development lifecycle.

What you will do

  • Build and maintain Python and PySpark pipelines for clinical trial data intake, scoring, and status logic.
  • Transform CSV, JSON, Parquet, API, and customer data into hirify.global's internal data models.
  • Write and optimize SQL for large datasets, pipeline validation, data investigation, and customer-facing analysis.
  • Translate clinical subject-matter expertise into data rules and explain pipeline results to non-engineering stakeholders.
  • Implement data quality checks, validation, reconciliation, monitoring, alerting, and dashboards.
  • Respond to customer-driven changes while maintaining enterprise SLA reliability and engineering standards.

Requirements

  • Production experience building and maintaining data pipelines in Python.
  • Hands-on experience with PySpark or a comparable distributed data processing framework.
  • Strong SQL skills with large, messy, multi-source datasets.
  • Knowledge of testing, code review, documentation, and CI/CD practices.
  • Experience collaborating with cross-functional and non-technical stakeholders.
  • Ability to work in a hybrid role based in New York.

Nice to have

  • Experience with Argo, Airflow, Databricks, dbt, or similar orchestration tools.
  • Clinical trial, healthcare, regulated-domain, entity matching, or data mastering experience.
  • Familiarity with AWS services such as S3, Lambda, or ECS.

Culture & Benefits

  • Work in a healthcare data environment focused on health equity and improved patient outcomes.
  • Collaborate with clinical subject-matter experts, analysts, customer-facing teams, and engineers.
  • Operate in an environment where reliability, data quality, and on-time delivery are central priorities.
  • Inclusive workplace with equal employment opportunity and reasonable accommodation support.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →