28 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Data Scientist (Machine Learning)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Π’Π΅ΠΊΡΡ:
TL;DR
Data Scientist (Machine Learning): Building and evaluating entity-matching models for messy, large-scale company data with an accent on embeddings, LLMs, NLP, classification, and experimental rigor. Focus on designing ranking and similarity approaches, analyzing model behavior and error cases, and balancing model quality, inference cost, and scalability.
Location: Lima, Peru
Company
designs, builds, and scales AI-powered solutions by combining data, artificial intelligence, cloud, design, and development.
What you will do
- Build and evaluate machine learning approaches for company and entity matching.
- Develop embedding- and LLM-based matching, scoring, ranking, and similarity methodologies.
- Work with messy, multilingual data including names, aliases, domains, websites, firmographic attributes, and data hierarchies.
- Define benchmark datasets, baselines, metrics, test sets, and error-analysis processes.
- Design experiments, analyze model behavior and trade-offs, and compare LLM-assisted solutions with lower-cost alternatives.
- Communicate recommendations to engineering and business stakeholders and document successful and unsuccessful experiments.
Requirements
- 5+ years of professional Data Science or Machine Learning experience.
- Strong applied machine learning fundamentals and experience with supervised and unsupervised learning, classification, NLP, embeddings, semantic similarity, and LLMs.
- Excellent Python and SQL skills.
- Working knowledge of neural networks and transformer architectures.
- Hands-on experience with TensorFlow, PyTorch, PyCaret, or equivalent machine learning frameworks.
- Strong English communication skills.
Nice to have
- Entity resolution, record linkage, deduplication, ranking, or similarity-scoring experience.
- Retrieval, clustering, or candidate-generation experience.
- Experience with Spark, Snowflake, Databricks, or BigQuery.
- Experience with company, domain, website, or firmographic data.
- Experience working with multilingual datasets.
Culture & Benefits
- High-performance culture built around empowering excellence, collaborative teamwork, respect, transparency, and efficient communication.
- Opportunity to work on AI-native transformation and scalable products with cross-functional teams.
- Environment focused on learning quickly, taking ownership, and modern ways of working.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β