Назад
Company hidden
2 часа назад

Member of Technical Staff, Data (AI)

200 000 - 350 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff, Data (AI): Building scalable data pipelines, synthetic data generation systems, and evaluation frameworks for training high-quality datasets for diffusion-based LLMs with an accent on petabyte-scale processing, web crawling, and data curation. Focus on designing distributed storage and retrieval systems, measuring data quality and diversity, and ensuring privacy-compliant data collection.

Location: Bay Area, in office

Salary: $200,000–$350,000 annual base salary, plus equity and benefits.

Company

hirify.global develops diffusion-based large language models, including Mercury, for faster and more efficient AI inference.

What you will do

  • Develop data mixes for LLM training using open-source datasets, synthetic data, and curated human feedback.
  • Design and implement petabyte-scale data processing pipelines.
  • Build systems for web crawling, data ingestion, and real-time processing.
  • Develop distributed tools for data storage, retrieval, and versioning.
  • Create evaluation frameworks for data diversity, quality, and representativeness.
  • Ensure data collection complies with privacy regulations.

Requirements

  • BS, MS, or PhD in Computer Science, Machine Learning, or a related field, or equivalent experience.
  • 3+ years of experience building large-scale data processing pipelines for AI or ML applications.
  • Strong Python skills and experience with Apache Spark, Beam, or Airflow.
  • Knowledge of synthetic data generation, data augmentation, web scraping, crawling technologies, and Common Crawl.
  • Understanding of machine learning fundamentals and experience with PyTorch or TensorFlow.
  • Experience with SQL and NoSQL databases for structured and unstructured data.

Nice to have

  • Experience with LLMs, tokenization, embeddings, and model architectures.
  • Experience managing human annotation workflows and quality control.
  • Experience with vector databases and embedding-based retrieval systems.
  • Knowledge of ethical AI practices, distributed computing, and large-scale storage such as HDFS, S3, or BigQuery.

Culture & Benefits

  • Collaboration with AI researchers and the inventors of diffusion models.
  • Opportunity to influence foundational AI technology and product direction.
  • Competitive salary, equity, and flexible vacation and paid time off.
  • Health, dental, and vision insurance with a 401(k) match.
  • Catered meals and commuter subsidies.
  • Collaborative and inclusive work environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →