2 часа назад
Member of Technical Staff, Data Infrastructure (AI)
200 000 - 350 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff, Data Infrastructure (AI): Building and operating scalable infrastructure for distributed LLM training pipelines and petabyte-scale data catalogs with an accent on high-throughput ingestion, data orchestration, storage, and retrieval. Focus on developing web crawling and real-time processing systems, improving dataset quality and versioning, and ensuring privacy-compliant data collection.
Location: Bay Area, United States; in-office
Annual base salary: $200,000–$350,000 USD, plus equity and benefits.
Company
AI company developing diffusion-based large language models and large-scale AI infrastructure.
What you will do
- Design, build, and operate scalable, fault-tolerant infrastructure for distributed LLM research, including compute, data orchestration, and storage.
- Develop high-throughput data ingestion, processing, and transformation systems for training data catalogs, deduplication, quality checks, and search.
- Build web crawling, data ingestion, and real-time processing systems for model training operations.
- Develop tools for efficient data storage, retrieval, and versioning across distributed systems.
- Work with researchers to accelerate experiments, develop datasets, and improve infrastructure efficiency.
- Ensure data collection follows privacy regulations and ethical AI practices.
Requirements
- BS, MS, PhD, or equivalent experience in Computer Science, Machine Learning, or a related field.
- 3+ years of experience building large-scale data processing pipelines, particularly for AI/ML applications.
- Strong Python skills and experience with Apache Spark, Beam, or Airflow.
- Experience with web scraping, crawling technologies, Common Crawl, SQL, and NoSQL databases.
- Understanding of machine learning fundamentals and experience with PyTorch or TensorFlow.
- Ability to work in the Bay Area in an office-based role.
Nice to have
- Experience with synthetic data generation and data augmentation.
- Experience with LLMs, tokenization, embeddings, and model architectures.
- Experience managing human annotation workflows and quality control.
- Experience with vector databases and embedding-based retrieval.
- Experience with distributed computing and large-scale storage such as HDFS, S3, or BigQuery.
Culture & Benefits
- Collaborate with AI researchers and inventors of diffusion models.
- Competitive salary, equity, and opportunities to shape foundational AI technology.
- Flexible vacation and paid time off.
- Health, dental, and vision insurance, plus 401(k) match.
- Catered meals and commuter subsidies.
- Collaborative and inclusive work environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 часов назад
Staff Engineer, Machine Learning Life Sciences (AI)
148 530 - 204 250$
2 часа назад
AI and ML Engineer (NLP/LLM)
180 000 - 260 000$
6 дней назад
Staff Data Scientist (AI/ML)
170 000 - 248 000$
3 часа назад
Member of Technical Staff, Machine Learning Engineer (AI)
240 000 - 350 000$
2 дня назад
Senior Data Scientist (AI/ML)
124 000 - 329 200$
4 дня назад
Data Infrastructure Engineer (AI)
170 000 - 360 000$