Назад
Company hidden
6 часов назад

Data Engineer, Forward Deployed (AI)

Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Data Engineer, Forward Deployed (AI): Building scalable Lakehouse pipelines for high-frequency time-series, industrial, laboratory, and historian data with an accent on AWS, Databricks/PySpark, and AI-ready data architecture. Focus on contextualising industrial signals, synchronising hybrid cloud and on-premise data, and building parallel training and low-latency inference pipelines for forecasting, anomaly detection, optimisation, and LLM search.

Location: Houston, United States; work format: Hybrid

Company

hirify.global builds Orbital, a physics-informed foundation model for energy operations across oil and gas, refineries, and petrochemicals.

What you will do

  • Architect and maintain scalable Lakehouse pipelines for high-frequency time-series, laboratory, historian, IoT, and industrial data.
  • Ingest and contextualise data from OPC UA servers, process historians, LIMS systems, sensors, alarms, events, and P&IDs.
  • Build real-time streaming and batch ingestion pipelines across Databricks, Lakeflow, Lakebase, AWS, and hybrid cloud/on-premise environments.
  • Synchronise historian archives, unstructured files, and AWS S3/EBS storage while managing secure, high-throughput transfers.
  • Implement schema-change detection, signal-drift handling, lineage, and audit trails across Spark, Databricks, and AWS pipelines.
  • Build parallel training and inference pipelines supporting time-series forecasting, anomaly detection, optimisation, and retrieval-augmented LLM workloads.

Requirements

  • Deep expertise in PostgreSQL, including partitioning, indexing, query optimisation, and storage design.
  • Strong Python skills for data processing, scripting, and pipeline orchestration.
  • Hands-on AWS experience, including EKS, S3, EBS, IAM, KMS, and CloudWatch.
  • Proven experience with Databricks and PySpark for large-scale distributed data processing.
  • Experience with industrial time-series data, control systems, DCS/SCADA logs, or process historians.
  • Experience managing unstructured data synchronisation across hybrid cloud and on-premise environments.

Nice to have

  • Data engineering experience in oil and gas or energy environments.
  • Knowledge of Kafka, Flink, Spark Streaming, or MLOps stacks for data versioning and lineage.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →