Назад
Company hidden
6 часов назад

Machine Learning Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning Engineer (AI) (data acquisition and distributed systems): Building data collection, extraction, filtering, synthetic data generation, and analysis pipelines to prepare high-quality datasets for domain-specific AI models with an accent on scalable data processing and model-assisted data quality. Focus on designing distributed systems for terabytes of data, developing crawling and ingestion infrastructure, and deploying Kubernetes-based services for indexing, storage, and synchronization.

Location: On-site in Dublin, California, United States

Company

AI product development focused on enterprise generative AI and domain-specific models.

What you will do

  • Design and develop data processing pipelines for data extraction, filtering, labeling, and analysis.
  • Implement machine learning models that improve data quality and diversity, including quality classifiers, document layout models, and code verification models.
  • Lead data acquisition engineering projects covering web crawling, data ingestion, and processing.
  • Develop scalable distributed systems capable of handling terabytes of data.
  • Architect data indexing and search algorithms and maintain backend storage services using key-value databases and synchronization systems.
  • Deploy solutions in Kubernetes infrastructure-as-code environments and perform routine system checks.

Requirements

  • BS, MS, or PhD in Computer Science or a related field.
  • Proficiency in at least one deep learning framework, such as PyTorch.
  • Experience training machine learning models for text or vision problems.
  • Strong expertise in stateful distributed systems, data processing, and large-scale data pipelines.
  • Proficiency in Python or another programming language commonly used in machine learning, with the ability to write clean, maintainable code.
  • Experience with distributed workloads and infrastructure such as multiprocessing, Ray, Docker, and Kubernetes, plus strong problem-solving skills when addressing data anomalies and bias.

Nice to have

  • Active GitHub contributions.
  • Experience building large-scale datasets and bespoke data processing libraries.
  • Familiarity with data crawling, collection, or processing tools such as Scrapy, Selenium, VPNs, Hadoop, and Datasketch.
  • Multilingual ability and familiarity with state-of-the-art techniques for preparing AI training data.

Culture & Benefits

  • Full-time work within an environment focused on diversity and inclusion.
  • Mentorship, knowledge exchange, and constructive feedback.
  • Continuous learning and professional development opportunities.
  • Encouragement of curiosity, creativity, and exploration beyond conventional approaches.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →