7 дней назад
Member of Technical Staff — ML Infra (Data)
200 000 - 300 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff — ML Infra (Data) (Spark/Ray/Dask, multimodal AI data): Building and operating petabyte-scale pipelines for ingesting, processing, filtering, and curating video, audio, and text training data with an accent on throughput, reliability, and data quality. Focus on productionizing research code, optimizing compute, I/O, and storage bottlenecks, and designing versioned, traceable datasets for foundation model training.
Location: In-person in Seattle, five days a week
Salary: $200,000–$300,000 base salary per year, plus meaningful equity.
Company
is a research company building photorealistic, real-time AI avatars with emotional intelligence and full-duplex audiovisual interaction.
What you will do
- Design, build, and operate large-scale pipelines for ingesting, processing, filtering, and curating multimodal video, audio, and text training data.
- Productionize research-grade data processing code while preserving correctness and processing fidelity.
- Optimize throughput and efficiency across compute, I/O, and storage, identifying and eliminating bottlenecks.
- Build data quality systems for deduplication, filtering, validation, and quality scoring.
- Manage petabyte-scale datasets, including storage architecture, versioning, lineage tracking, and cost efficiency.
- Work with researchers to translate data requirements into scalable processing systems and efficient research tooling.
Requirements
- Production experience building and operating large-scale data pipelines where naive approaches do not scale.
- Strong proficiency with distributed data processing frameworks such as Spark, Ray, Dask, or similar.
- Strong software engineering fundamentals, including clean, testable, and maintainable code.
- Familiarity with ML data pipelines and the effect of data quality and format on model training.
- Ability to turn research prototypes into production-ready pipelines in days rather than weeks.
- Ability to work in person in Seattle five days per week.
Nice to have
- Experience with multimodal data, including video and audio formats, codecs, FFmpeg, or decord.
- Experience building pipelines for large-scale model pre-training or fine-tuning.
- Familiarity with DVC, Delta Lake, Apache Iceberg, streaming pipelines, or online data processing.
- Prior work at an AI lab, video platform, or other data-intensive company.
- Contributions to open-source data tooling.
Culture & Benefits
- Small research-focused team working on unsolved real-time AI problems.
- Visa sponsorship is available from day one, including O-1, H-1B, and green card support.
- AI-native tooling with unlimited tokens.
- HSA plan with approximately $2,000 in annual company contributions.
- 15 days of PTO, public holidays, and a full week of office closure at year-end.
- Daily lunch, drinks, snacks, commuter benefits, and a 401(k) plan.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Staff ML Infra Engineer (Search & Discovery)
174 000 - 299 000$
6 дней назад
Lead Data Engineer (AI/ML)
188 251 - 230 084$
6 дней назад
Data Software Engineer III - ML Ops
108 160 - 162 240$
7 дней назад
Senior Bioinformatics Research Engineer (Machine Learning)
161 925 - 227 325$
7 дней назад
AI/ML Engineer
107 000 - 233 000$
Snowflake
7 дней назад
Senior ML Platform Engineer (AI)
236 000 - 339 250$