7 часов назад
Data Infrastructure & ML Engineer (Hybrid Role)
12 213 307 - 18 319 961$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Data Infrastructure & ML Engineer (Hybrid Role) (Data Pipelines/Python/ML): Building scalable ETL/ELT pipelines, database systems, and Python-based data processing workflows for analytics and machine learning with an accent on data traceability, distributed storage, and production reliability. Focus on designing partitioning and sharding strategies, validating semi-structured data, and preparing scalable foundations for model training and deployment.
Location: Hybrid role in Beverly, Massachusetts, United States
Salary: $122,133.07–$183,199.61 annual base salary, plus eligibility for a team incentive bonus and benefits for regular employees working 20+ hours per week.
Company
develops semiconductor manufacturing equipment and technologies.
What you will do
- Design and build end-to-end ETL/ELT pipelines for tool-generated logs, JSON, and other semi-structured data.
- Implement data validation, monitoring, error handling, traceability, and governance to ensure reliable data quality.
- Design scalable database schemas and manage single-node and distributed database environments.
- Implement tablespaces, partitioning, sharding, and query optimization for large-scale datasets.
- Develop Python workflows using dataframes and libraries such as Pandas, NumPy, and Plotly.
- Prepare datasets and scalable data foundations for machine learning model training and deployment in collaboration with data scientists and engineers.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, plus 5+ years of experience.
- Strong experience with database design, SQL-based systems, distributed systems, partitioning, and sharding.
- Proven experience building ETL/ELT data pipelines.
- Strong Python proficiency for data processing, including experience with dataframes.
- Experience handling log-based and semi-structured data such as JSON.
- Understanding of data traceability, validation, governance, performance, and reliability.
Nice to have
- Experience with time-series or log analytics systems.
- Exposure to real-time or streaming architectures such as Kafka.
- Experience with Azure, AWS, or GCP.
- Familiarity with machine learning workflows and lifecycle.
- Semiconductor or high-throughput systems experience.
Culture & Benefits
- Collaboration across engineering and data teams.
- Focus on problem-solving, analytical thinking, data integrity, performance, and reliability.
- Team incentive bonus eligibility.
- Comprehensive benefits package for regular employees working 20+ hours per week.
- Equal opportunity employment and reasonable accommodation for qualified candidates and employees with disabilities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Data Infrastructure Engineer (AI)
170 000 - 360 000$
12 часов назад
Senior Data Scientist (AI)
225 000 - 250 000$
3 часа назад
Senior Data Engineer (Python/SQL)
10 часов назад
ML Evals Engineer (AI)
180 000 - 350 000$
15 часов назад
Senior Data Scientist (AI/ML)
95 000 - 166 000$
6 дней назад