Назад
Company hidden
3 дня назад

Research Engineer, Data Infrastructure (AI)

200 000 - 350 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
UK/US/India
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Engineer, Data Infrastructure (AI): Building scalable data infrastructure and high-throughput pipelines for acquiring, processing, filtering, deduplicating, and augmenting massive datasets used to pretrain foundational models, with an accent on dataset quality, reproducibility, and data-mixture optimization. Focus on designing ablation experiments, co-designing data versioning and loading systems with research teams, and managing novel external data sources.

Location: On-site in San Francisco, California. hirify.global also has in-person offices in London and Bangalore. Visa sponsorship support is available on a case-by-case basis.

Salary: $200,000–$350,000 per year, plus equity.

Company

hirify.global develops AI model architectures and experiences that enable foundation models to learn from and interact with the world like humans.

What you will do

  • Build and operate scalable infrastructure for acquiring, ingesting, and combining massive text datasets.
  • Design reproducible high-throughput pipelines for preprocessing, filtering, deduplication, and data augmentation.
  • Run ablation experiments to evaluate how data sources, processing choices, and mixture weights affect model quality.
  • Partner with research and infrastructure teams on data loading, versioning, and experimentation systems.
  • Establish data-quality standards and connect dataset characteristics with model behavior.
  • Source novel datasets and manage relationships and budgets with external data vendors and partners.

Requirements

  • Hands-on experience with ML data infrastructure, including training data pipelines, dataset versioning, large-scale data loading, and connections between data systems and model training and inference.
  • Strong software engineering skills, including clean, well-tested code and fluency with modern tools.
  • Experience building and evaluating datasets for generative models and working knowledge of model training and inference.
  • Experience with large-scale parallel data processing using tools such as Ray, Spark, or Kubernetes.
  • Experience with language model pretraining.
  • Ability to work in person from the San Francisco office.

Culture & Benefits

  • In-person, collaborative work environment with a focus on rapid execution and high-quality engineering.
  • Fully covered medical, dental, and vision insurance for employees and families.
  • 401(k), flexible PTO, and parental leave.
  • Monthly commuter allowance and daily lunch, dinner, and snacks.
  • Competitive base salary and equity package.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →