Назад
Company hidden
7 часов назад

Data Infrastructure & ML Engineer (Hybrid Role)

12 213 307 - 18 319 961$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Data Infrastructure & ML Engineer (Hybrid Role) (Data Pipelines/Python/ML): Building scalable ETL/ELT pipelines, database systems, and Python-based data processing workflows for analytics and machine learning with an accent on data traceability, distributed storage, and production reliability. Focus on designing partitioning and sharding strategies, validating semi-structured data, and preparing scalable foundations for model training and deployment.

Location: Hybrid role in Beverly, Massachusetts, United States

Salary: $122,133.07–$183,199.61 annual base salary, plus eligibility for a team incentive bonus and benefits for regular employees working 20+ hours per week.

Company

hirify.global develops semiconductor manufacturing equipment and technologies.

What you will do

  • Design and build end-to-end ETL/ELT pipelines for tool-generated logs, JSON, and other semi-structured data.
  • Implement data validation, monitoring, error handling, traceability, and governance to ensure reliable data quality.
  • Design scalable database schemas and manage single-node and distributed database environments.
  • Implement tablespaces, partitioning, sharding, and query optimization for large-scale datasets.
  • Develop Python workflows using dataframes and libraries such as Pandas, NumPy, and Plotly.
  • Prepare datasets and scalable data foundations for machine learning model training and deployment in collaboration with data scientists and engineers.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, plus 5+ years of experience.
  • Strong experience with database design, SQL-based systems, distributed systems, partitioning, and sharding.
  • Proven experience building ETL/ELT data pipelines.
  • Strong Python proficiency for data processing, including experience with dataframes.
  • Experience handling log-based and semi-structured data such as JSON.
  • Understanding of data traceability, validation, governance, performance, and reliability.

Nice to have

  • Experience with time-series or log analytics systems.
  • Exposure to real-time or streaming architectures such as Kafka.
  • Experience with Azure, AWS, or GCP.
  • Familiarity with machine learning workflows and lifecycle.
  • Semiconductor or high-throughput systems experience.

Culture & Benefits

  • Collaboration across engineering and data teams.
  • Focus on problem-solving, analytical thinking, data integrity, performance, and reliability.
  • Team incentive bonus eligibility.
  • Comprehensive benefits package for regular employees working 20+ hours per week.
  • Equal opportunity employment and reasonable accommodation for qualified candidates and employees with disabilities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →