6 часов назад
Machine Learning Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Machine Learning Engineer (AI) (data acquisition and distributed systems): Building data collection, extraction, filtering, synthetic data generation, and analysis pipelines to prepare high-quality datasets for domain-specific AI models with an accent on scalable data processing and model-assisted data quality. Focus on designing distributed systems for terabytes of data, developing crawling and ingestion infrastructure, and deploying Kubernetes-based services for indexing, storage, and synchronization.
Location: On-site in Dublin, California, United States
Company
AI product development focused on enterprise generative AI and domain-specific models.
What you will do
- Design and develop data processing pipelines for data extraction, filtering, labeling, and analysis.
- Implement machine learning models that improve data quality and diversity, including quality classifiers, document layout models, and code verification models.
- Lead data acquisition engineering projects covering web crawling, data ingestion, and processing.
- Develop scalable distributed systems capable of handling terabytes of data.
- Architect data indexing and search algorithms and maintain backend storage services using key-value databases and synchronization systems.
- Deploy solutions in Kubernetes infrastructure-as-code environments and perform routine system checks.
Requirements
- BS, MS, or PhD in Computer Science or a related field.
- Proficiency in at least one deep learning framework, such as PyTorch.
- Experience training machine learning models for text or vision problems.
- Strong expertise in stateful distributed systems, data processing, and large-scale data pipelines.
- Proficiency in Python or another programming language commonly used in machine learning, with the ability to write clean, maintainable code.
- Experience with distributed workloads and infrastructure such as multiprocessing, Ray, Docker, and Kubernetes, plus strong problem-solving skills when addressing data anomalies and bias.
Nice to have
- Active GitHub contributions.
- Experience building large-scale datasets and bespoke data processing libraries.
- Familiarity with data crawling, collection, or processing tools such as Scrapy, Selenium, VPNs, Hadoop, and Datasketch.
- Multilingual ability and familiarity with state-of-the-art techniques for preparing AI training data.
Culture & Benefits
- Full-time work within an environment focused on diversity and inclusion.
- Mentorship, knowledge exchange, and constructive feedback.
- Continuous learning and professional development opportunities.
- Encouragement of curiosity, creativity, and exploration beyond conventional approaches.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Data Scientist / Machine Learning Engineer (Generative AI)
2 часа назад
Machine Learning Engineer (Generative ML)
173 000 - 259 000$
4 часа назад
Machine Learning Research Engineer (Robotics)
1 час назад
Machine Learning Scientist (AI)
2 часа назад
Machine Learning Research Scientist (AI)
100 000 - 300 000$
3 часа назад