23 часа назад
Data Scientist (AI)
hhВакансия с HeadHunter. Контакт ведёт на hh.ru
Мэтч & Сопровод
Покажет вашу совместимость и напишет письмо
Описание вакансии
Текст:
TL;DR
Data Scientist (AI) (Python/PySpark/Databricks): Building a large-scale media intelligence platform and AI-powered audience analytics solutions with an accent on production machine learning, entity resolution, semantic similarity, and LLM-augmented pipelines. Focus on designing end-to-end data science projects, orchestrating reproducible MLOps workflows, processing billion-row datasets, and deploying reliable models.
Location: Poland; fully remote, office-based, or hybrid work options are available.
Company
An international software development and outsourcing company delivering data analytics and AI-powered solutions for customer behavior analysis, marketing optimization, and business decision-making.
What you will do
- Own end-to-end delivery of significant data science projects, from problem scoping and methodology design through production deployment.
- Design and implement production-quality Python and PySpark solutions on Databricks for billion-row datasets.
- Develop machine learning and AI workflows for entity resolution, probabilistic record linkage, embedding-based matching, semantic similarity, and LLM-augmented pipelines.
- Apply DataOps and MLOps practices, including experiment tracking, pipeline orchestration, model monitoring, and reproducibility.
- Create reusable tools, libraries, technical documentation, user stories, and solution designs; conduct code reviews and lead technical workshops.
- Collaborate with product, engineering, operations, and data engineering teams on scalable pipeline design and cross-functional reviews; mentor junior data scientists.
Requirements
- 3+ years of hands-on data science experience owning complex, multi-sprint projects.
- Bachelor's degree in Statistics, Data Science, Computer Science, Mathematics, or another quantitative field; a Master's degree is preferred.
- Advanced Python skills with production-quality, well-tested, and well-documented code, plus strong SQL and PySpark experience.
- Hands-on Databricks experience, including Workflows, Delta Lake, and job orchestration.
- Working knowledge of AWS or Google Cloud Platform and strong foundations in regression, classification, clustering, model evaluation, and experimental design.
- Experience with MLOps, Airflow-based orchestration, reproducible deployment, RAG, LLM applications, vector databases, and semantic search; English at Intermediate+ level or above.
Nice to have
- Knowledge graph construction, entity resolution, semantic data modeling, probabilistic record linkage, identity graphs, or embedding-based entity matching.
- Experience with causal inference, A/B testing, synthetic control, uplift modeling, deduplication, data enrichment, or web-to-TV linkage.
- Background in media, ad tech, audience measurement, TV viewership, digital audience modeling, CTV/OTT, or privacy-constrained identity resolution.
- Familiarity with Nielsen, Comscore, LiveRamp, or The Trade Desk.
Culture & Benefits
- Opportunities to work with clients in FinTech, Healthcare, Retail, Telecom, and other industries.
- Possibility to change projects and develop expertise in different business domains.
- Mentoring and onboarding systems supporting professional, financial, and career development.
- Corporate training portal and compensation for professional certifications such as AWS and PMP.
- Private health insurance and sports compensation, depending on the type of employment.
- Employment contract or B2B engagement options, plus referral and corporate activity programs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →