6 часов назад
Data Engineer, Forward Deployed (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Data Engineer, Forward Deployed (AI): Building scalable Lakehouse pipelines for high-frequency time-series, industrial, laboratory, and historian data with an accent on AWS, Databricks/PySpark, and AI-ready data architecture. Focus on contextualising industrial signals, synchronising hybrid cloud and on-premise data, and building parallel training and low-latency inference pipelines for forecasting, anomaly detection, optimisation, and LLM search.
Location: Houston, United States; work format: Hybrid
Company
builds Orbital, a physics-informed foundation model for energy operations across oil and gas, refineries, and petrochemicals.
What you will do
- Architect and maintain scalable Lakehouse pipelines for high-frequency time-series, laboratory, historian, IoT, and industrial data.
- Ingest and contextualise data from OPC UA servers, process historians, LIMS systems, sensors, alarms, events, and P&IDs.
- Build real-time streaming and batch ingestion pipelines across Databricks, Lakeflow, Lakebase, AWS, and hybrid cloud/on-premise environments.
- Synchronise historian archives, unstructured files, and AWS S3/EBS storage while managing secure, high-throughput transfers.
- Implement schema-change detection, signal-drift handling, lineage, and audit trails across Spark, Databricks, and AWS pipelines.
- Build parallel training and inference pipelines supporting time-series forecasting, anomaly detection, optimisation, and retrieval-augmented LLM workloads.
Requirements
- Deep expertise in PostgreSQL, including partitioning, indexing, query optimisation, and storage design.
- Strong Python skills for data processing, scripting, and pipeline orchestration.
- Hands-on AWS experience, including EKS, S3, EBS, IAM, KMS, and CloudWatch.
- Proven experience with Databricks and PySpark for large-scale distributed data processing.
- Experience with industrial time-series data, control systems, DCS/SCADA logs, or process historians.
- Experience managing unstructured data synchronisation across hybrid cloud and on-premise environments.
Nice to have
- Data engineering experience in oil and gas or energy environments.
- Knowledge of Kafka, Flink, Spark Streaming, or MLOps stacks for data versioning and lineage.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Software Engineer (AI)
150 000 - 200 000$
5 часов назад
Staff Data Engineer (AI)
4 дня назад
Senior Backend Engineer, Data Systems (AI Infrastructure)
180 000 - 240 000$
2 дня назад
Backend Software Engineer (AI Infrastructure)
200 000 - 250 000$
Scale AI
4 дня назад
Field Engineer, Data Engine (AI)
140 000 - 260 000$
5 дней назад
Forward Deployed Engineer (AI)
145 000 - 180 000$