Назад
Company hidden
7 дней назад

Senior ML Engineer / Applied ML Engineer (Entity Resolution)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Poland
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior ML Engineer / Applied ML Engineer (Entity Resolution): Building and evolving production-ready entity resolution systems that apply machine learning, LLMs, and search algorithms to hundreds of millions to billions of records with an accent on distributed matching, large-scale data pipelines, and analytical SQL. Focus on designing clustering and deduplication algorithms, optimizing candidate generation and scoring, and balancing accuracy, recall, runtime, and infrastructure cost.

Location: Warsaw, Poland; hybrid

Company

hirify.global connects talent with product careers at high-growth companies, and the client is a global management consultancy building an internal data platform for entity resolution and advanced analytics.

What you will do

  • Design and improve entity matching, clustering, deduplication, and distributed matching algorithms.
  • Build production ML systems using rule-based, statistical, ML-assisted, embedding, nearest-neighbor, and similarity-based methods.
  • Develop and optimize Spark and SQL pipelines processing hundreds of millions to billions of records.
  • Build analytical workflows with Airflow and dbt, and data models in Snowflake or similar warehouses.
  • Measure match rate, precision, recall, accuracy, runtime, and cost while iterating on matching approaches.
  • Debug data quality issues and ensure distributed pipelines are fault-tolerant, observable, and cost-efficient.

Requirements

  • 5+ years of software development experience using Python, ideally including team leadership.
  • Strong SQL skills and experience with Snowflake or a similar analytical database.
  • Production experience with Apache Spark and Databricks using PySpark or Scala.
  • Strong understanding of distributed systems, algorithms, partitioning, shuffles, joins, and scalability trade-offs.
  • Experience building complex data pipelines or data systems and AI/ML-assisted systems involving embeddings, inference, or re-ranking.
  • Experience with or willingness to learn fuzzy and semantic matching methods such as edit distance, token similarity, BM25, vector embeddings, cosine similarity, and approximate nearest-neighbor search.

Nice to have

  • Experience with Airflow, entity resolution, deduplication, record linkage, search systems, vector databases, or large-scale data quality frameworks.
  • Experience operating pipelines over 100M+ records.
  • Experience with Kubernetes, containerized or serverless functions, Azure or other cloud platforms, and GitHub Actions.

Culture & Benefits

  • Work on a high-impact data innovation and productization initiative within a top-tier management consultancy.
  • Contribute to building a comprehensive business directory using AI/ML enrichment and semantic matching.
  • Use modern cloud-based infrastructure and large-scale data technologies.
  • Work in a collaborative, innovative environment supporting operations across EMEA.

Hiring process

  • hirify.global recruiter interview.
  • Cultural fit and technical interviews.
  • Optional final interview followed by an offer.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →