7 дней назад
Senior ML Engineer / Applied ML Engineer (Entity Resolution)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior ML Engineer / Applied ML Engineer (Entity Resolution): Building and evolving production-ready entity resolution systems that apply machine learning, LLMs, and search algorithms to hundreds of millions to billions of records with an accent on distributed matching, large-scale data pipelines, and analytical SQL. Focus on designing clustering and deduplication algorithms, optimizing candidate generation and scoring, and balancing accuracy, recall, runtime, and infrastructure cost.
Location: Warsaw, Poland; hybrid
Company
connects talent with product careers at high-growth companies, and the client is a global management consultancy building an internal data platform for entity resolution and advanced analytics.
What you will do
- Design and improve entity matching, clustering, deduplication, and distributed matching algorithms.
- Build production ML systems using rule-based, statistical, ML-assisted, embedding, nearest-neighbor, and similarity-based methods.
- Develop and optimize Spark and SQL pipelines processing hundreds of millions to billions of records.
- Build analytical workflows with Airflow and dbt, and data models in Snowflake or similar warehouses.
- Measure match rate, precision, recall, accuracy, runtime, and cost while iterating on matching approaches.
- Debug data quality issues and ensure distributed pipelines are fault-tolerant, observable, and cost-efficient.
Requirements
- 5+ years of software development experience using Python, ideally including team leadership.
- Strong SQL skills and experience with Snowflake or a similar analytical database.
- Production experience with Apache Spark and Databricks using PySpark or Scala.
- Strong understanding of distributed systems, algorithms, partitioning, shuffles, joins, and scalability trade-offs.
- Experience building complex data pipelines or data systems and AI/ML-assisted systems involving embeddings, inference, or re-ranking.
- Experience with or willingness to learn fuzzy and semantic matching methods such as edit distance, token similarity, BM25, vector embeddings, cosine similarity, and approximate nearest-neighbor search.
Nice to have
- Experience with Airflow, entity resolution, deduplication, record linkage, search systems, vector databases, or large-scale data quality frameworks.
- Experience operating pipelines over 100M+ records.
- Experience with Kubernetes, containerized or serverless functions, Azure or other cloud platforms, and GitHub Actions.
Culture & Benefits
- Work on a high-impact data innovation and productization initiative within a top-tier management consultancy.
- Contribute to building a comprehensive business directory using AI/ML enrichment and semantic matching.
- Use modern cloud-based infrastructure and large-scale data technologies.
- Work in a collaborative, innovative environment supporting operations across EMEA.
Hiring process
- recruiter interview.
- Cultural fit and technical interviews.
- Optional final interview followed by an offer.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →