Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Data Engineer | AI Data Platform (AI/multimodal data): Building infrastructure, data engines, and AI-native pipelines for large-scale multimodal datasets used to train frontier models with an accent on distributed processing, multimodal storage, and data quality. Focus on developing Spark, Flink, or Ray systems, integrating model-training feedback, and optimizing datasets for LLM and VLM workloads.
Location: Mountain View, New York City, or Redmond, United States
Salary: USD $142,800–$274,800 per year for IC5 roles across the U.S.; USD $165,600–$296,400 per year for IC6 roles. San Francisco Bay Area and New York City ranges may differ.
Company
Microsoft AI develops large-scale artificial intelligence systems and the data infrastructure required to train frontier models.
What you will do
- Build AI data infrastructure, data engines, and intelligent multimodal data processing systems.
- Develop AI-native data pipelines, storage systems, and table layers.
- Support datasets spanning text, image, video, and audio modalities.
- Drive the data, model, and evaluation iteration loop through targeted data improvements.
- Partner closely with model and training teams.
Requirements
- Master’s degree in computer science, mathematics, software engineering, computer engineering, or a related field with 4+ years of relevant experience; or bachelor’s degree with 6+ years of experience; equivalent experience is also accepted.
- Experience with distributed data platforms such as Spark, Flink, or Ray.
- Proficiency in Python and experience with SQL and Shell.
- Experience working with multimodal data, including text, image, video, or audio.
- Role locations are in the United States: Mountain View, New York City, or Redmond.
Nice to have
- Experience building datasets for LLM, VLM, or multimodal model pre-training or post-training.
- Experience with synthetic data generation, including text, image-text, rendering, diffusion-based, or multimodal trajectory generation.
- Familiarity with MMLU, MMBench, MM-BrowseComp, and evaluation-driven failure analysis.
- Experience with Lance, Iceberg, Paimon, or Parquet, including schema evolution, versioning, transactions, indexing, and performance optimization.
Culture & Benefits
- Work on large-scale AI data infrastructure supporting frontier model development.
- Close collaboration with model and training teams.
- Benefits and other compensation may be available depending on role eligibility.
- Applications are accepted on an ongoing basis until the position is filled.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior Data Engineer (AI)
108 600 - 183 018$
Scale AI
4 дня назад
Senior Data Engineer (AI)
180 000 - 225 000$
4 дня назад
Senior Data Platform Engineer (AI)
125 000 - 205 000$
Thinking Machines Lab
4 дня назад
Data Infrastructure Engineer (AI)
300 000 - 400 000$
Anthropic
8 дней назад
Analytics Data Engineer (AI)
320 000 - 405 000$
4 дня назад
AI Data Engineer (Microsoft Fabric)
82 500 - 137 500$