Назад
Company hidden
2 часа назад

Member of Technical Staff, Data Infrastructure (AI)

200 000 - 350 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff, Data Infrastructure (AI): Building and operating scalable infrastructure for distributed LLM training pipelines and petabyte-scale data catalogs with an accent on high-throughput ingestion, data orchestration, storage, and retrieval. Focus on developing web crawling and real-time processing systems, improving dataset quality and versioning, and ensuring privacy-compliant data collection.

Location: Bay Area, United States; in-office

Annual base salary: $200,000–$350,000 USD, plus equity and benefits.

Company

AI company developing diffusion-based large language models and large-scale AI infrastructure.

What you will do

  • Design, build, and operate scalable, fault-tolerant infrastructure for distributed LLM research, including compute, data orchestration, and storage.
  • Develop high-throughput data ingestion, processing, and transformation systems for training data catalogs, deduplication, quality checks, and search.
  • Build web crawling, data ingestion, and real-time processing systems for model training operations.
  • Develop tools for efficient data storage, retrieval, and versioning across distributed systems.
  • Work with researchers to accelerate experiments, develop datasets, and improve infrastructure efficiency.
  • Ensure data collection follows privacy regulations and ethical AI practices.

Requirements

  • BS, MS, PhD, or equivalent experience in Computer Science, Machine Learning, or a related field.
  • 3+ years of experience building large-scale data processing pipelines, particularly for AI/ML applications.
  • Strong Python skills and experience with Apache Spark, Beam, or Airflow.
  • Experience with web scraping, crawling technologies, Common Crawl, SQL, and NoSQL databases.
  • Understanding of machine learning fundamentals and experience with PyTorch or TensorFlow.
  • Ability to work in the Bay Area in an office-based role.

Nice to have

  • Experience with synthetic data generation and data augmentation.
  • Experience with LLMs, tokenization, embeddings, and model architectures.
  • Experience managing human annotation workflows and quality control.
  • Experience with vector databases and embedding-based retrieval.
  • Experience with distributed computing and large-scale storage such as HDFS, S3, or BigQuery.

Culture & Benefits

  • Collaborate with AI researchers and inventors of diffusion models.
  • Competitive salary, equity, and opportunities to shape foundational AI technology.
  • Flexible vacation and paid time off.
  • Health, dental, and vision insurance, plus 401(k) match.
  • Catered meals and commuter subsidies.
  • Collaborative and inclusive work environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →