2 дня назад
Data Engineer II- Life Sciences
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Data Engineer II- Life Sciences (Python/PySpark): Building and maintaining production data pipelines that ingest, transform, validate, and reconcile clinical trial data across diverse source formats with an accent on distributed processing, SQL analysis, and data quality. Focus on translating clinical domain rules into reliable pipeline logic, responding to enterprise SLA requirements, and improving observability for job health and data anomalies.
Location: New York, United States; Workplace: Hybrid
Company
provides healthcare information and data products designed to improve doctor interactions, patient outcomes, and the drug development lifecycle.
What you will do
- Build and maintain Python and PySpark pipelines for clinical trial data intake, scoring, and status logic.
- Transform CSV, JSON, Parquet, API, and customer data into 's internal data models.
- Write and optimize SQL for large datasets, pipeline validation, data investigation, and customer-facing analysis.
- Translate clinical subject-matter expertise into data rules and explain pipeline results to non-engineering stakeholders.
- Implement data quality checks, validation, reconciliation, monitoring, alerting, and dashboards.
- Respond to customer-driven changes while maintaining enterprise SLA reliability and engineering standards.
Requirements
- Production experience building and maintaining data pipelines in Python.
- Hands-on experience with PySpark or a comparable distributed data processing framework.
- Strong SQL skills with large, messy, multi-source datasets.
- Knowledge of testing, code review, documentation, and CI/CD practices.
- Experience collaborating with cross-functional and non-technical stakeholders.
- Ability to work in a hybrid role based in New York.
Nice to have
- Experience with Argo, Airflow, Databricks, dbt, or similar orchestration tools.
- Clinical trial, healthcare, regulated-domain, entity matching, or data mastering experience.
- Familiarity with AWS services such as S3, Lambda, or ECS.
Culture & Benefits
- Work in a healthcare data environment focused on health equity and improved patient outcomes.
- Collaborate with clinical subject-matter experts, analysts, customer-facing teams, and engineers.
- Operate in an environment where reliability, data quality, and on-time delivery are central priorities.
- Inclusive workplace with equal employment opportunity and reasonable accommodation support.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Data Engineer (AI)
112 400 - 174 500$
6 дней назад
Mid-Level/Lead Data Engineer (Python, AWS, Spark)
85 500 - 130 000$
6 дней назад
Data Engineer II (Healthcare)
7 дней назад
Data Engineer (AI)
150 000 - 200 000$
2 дня назад
Data Engineer (Databricks)
110 000 - 126 000$
3 дня назад
Data Engineer (AI)
89 886 - 175 444$