Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Location: Mountain View, New York, or Redmond, United States. Employees living within 50 miles of a U.S. Microsoft office are expected to work from that office at least four days per week starting January 26, 2026.
Salary: USD $142,800–$274,800 per year for Data Engineering IC5 across the U.S.; USD $188,000–$304,200 in the San Francisco Bay Area and New York City metropolitan area. Data Engineering IC6: USD $165,600–$296,400 across the U.S.; USD $220,800–$331,200 in the San Francisco Bay Area and New York City metropolitan area.
Company
Microsoft AI develops frontier AI models and products, including humanist superintelligence systems designed to remain controllable, safety-aligned, and focused on human benefit.
What you will do
- Build EB-scale data infrastructure for collecting, ingesting, cleaning, curating, governing, querying, and analyzing multimodal training data.
- Develop intelligent processing systems for text, images, video, and documents, including labeling, taxonomy, semantic feature extraction, classification, captioning, and quality scoring.
- Build AI-native pipelines using LLMs, VLMs, and agents for filtering, deduplication, annotation, synthetic data generation, orchestration, and anomaly detection.
- Design AI-optimized storage and table layers with Lance, Iceberg, Paimon, and Parquet, including schema evolution, indexing, transactions, and versioning.
- Create rare, high-value datasets through real-world collection, human annotation, and synthetic generation for domains such as healthcare, education, retail inspection, OCR, and GUI interaction.
- Partner with model researchers and training engineers to use evaluation results and failure analysis to improve data and model performance.
Requirements
- Master’s degree in computer science, mathematics, software engineering, computer engineering, or a related field with 4+ years of relevant experience; alternatively, a bachelor’s degree with 6+ years of experience or equivalent experience.
- Experience with distributed data platforms such as Spark, Flink, or Ray.
- Proficiency in Python and experience with SQL and Shell.
- Experience working with multimodal data, including text, images, video, or audio.
- Experience in data engineering, data science, software development, data modeling, or business analytics.
Nice to have
- Experience building datasets for LLM, VLM, or multimodal model pre-training and post-training.
- Experience with synthetic data generation, evaluation benchmarks, and failure-analysis-driven data improvements.
- Knowledge of Lance, Iceberg, Paimon, Parquet, ETL systems, data warehouses, and large-scale dataset processing.
- Understanding of LLMs, speech/audio models, vision models, multimodal models, agents, and modern LLM toolchains.
Culture & Benefits
- Startup-like Superintelligence Team within Microsoft AI with a collaborative, fast-paced working environment.
- Work on AI systems intended to advance science, education, and global well-being.
- Opportunity to partner with product teams whose models reach billions of users.
- Benefits and additional compensation may be available depending on the role and location.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →