Назад
3 дня назад

Data Engineer | AI Data Platform

142 800 - 274 800$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Data Engineer | AI Data Platform (AI/multimodal data): Building infrastructure, data engines, and AI-native pipelines for large-scale multimodal datasets used to train frontier models with an accent on distributed processing, multimodal storage, and data quality. Focus on developing Spark, Flink, or Ray systems, integrating model-training feedback, and optimizing datasets for LLM and VLM workloads.

Location: Mountain View, New York City, or Redmond, United States

Salary: USD $142,800–$274,800 per year for IC5 roles across the U.S.; USD $165,600–$296,400 per year for IC6 roles. San Francisco Bay Area and New York City ranges may differ.

Company

Microsoft AI develops large-scale artificial intelligence systems and the data infrastructure required to train frontier models.

What you will do

  • Build AI data infrastructure, data engines, and intelligent multimodal data processing systems.
  • Develop AI-native data pipelines, storage systems, and table layers.
  • Support datasets spanning text, image, video, and audio modalities.
  • Drive the data, model, and evaluation iteration loop through targeted data improvements.
  • Partner closely with model and training teams.

Requirements

  • Master’s degree in computer science, mathematics, software engineering, computer engineering, or a related field with 4+ years of relevant experience; or bachelor’s degree with 6+ years of experience; equivalent experience is also accepted.
  • Experience with distributed data platforms such as Spark, Flink, or Ray.
  • Proficiency in Python and experience with SQL and Shell.
  • Experience working with multimodal data, including text, image, video, or audio.
  • Role locations are in the United States: Mountain View, New York City, or Redmond.

Nice to have

  • Experience building datasets for LLM, VLM, or multimodal model pre-training or post-training.
  • Experience with synthetic data generation, including text, image-text, rendering, diffusion-based, or multimodal trajectory generation.
  • Familiarity with MMLU, MMBench, MM-BrowseComp, and evaluation-driven failure analysis.
  • Experience with Lance, Iceberg, Paimon, or Parquet, including schema evolution, versioning, transactions, indexing, and performance optimization.

Culture & Benefits

  • Work on large-scale AI data infrastructure supporting frontier model development.
  • Close collaboration with model and training teams.
  • Benefits and other compensation may be available depending on role eligibility.
  • Applications are accepted on an ongoing basis until the position is filled.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →