Назад
4 дня назад

Data Infrastructure Engineer (AI)

300 000 - 400 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Data Infrastructure Engineer (AI): Building and operating scalable infrastructure for distributed LLM training pipelines, multimodal data catalogs, and petabyte-scale processing systems with an accent on distributed compute, data orchestration, storage, and data quality. Focus on designing high-throughput ingestion and transformation systems, scaling deduplication and search, and ensuring traceability, monitoring, and reproducibility across the data lifecycle.

Location: On-site in San Francisco, California; the listing also mentions New York.

Annual salary: $300,000–$400,000 USD, depending on background, skills, and experience.

Company

Thinking Machines Lab builds AI systems designed to extend human will and judgment, including frontier models, model customization tools, and human-AI interfaces.

What you will do

  • Design, build, and operate fault-tolerant infrastructure for LLM research, including distributed compute, data orchestration, and multimodal storage.
  • Develop high-throughput systems for data ingestion, processing, transformation, training data catalogs, deduplication, quality checks, and search.
  • Build traceability, reproducibility, and quality-control systems across the data lifecycle.
  • Implement monitoring and alerting for platform reliability and performance.
  • Collaborate with research teams to improve data quality, accelerate experiments, and shorten training cycles.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, or a related field.
  • Proficiency in at least one backend language: Python or Rust.
  • Fluency with distributed compute frameworks such as Apache Spark or Ray.
  • Strong familiarity with cloud infrastructure, data lake architectures, and batch and streaming pipelines.
  • Experience with Kafka, dbt, Terraform, Airflow, web crawlers, deduplication, data mining, search, file formats, and storage systems such as Parquet and Delta Lake.
  • Ability to own projects end-to-end, work across technology stacks, and collaborate with research and cross-functional teams.

Culture & Benefits

  • Small, high-impact engineering team working closely with researchers and subject matter experts.
  • Generous health, dental, and vision benefits.
  • Unlimited paid time off and paid parental leave.
  • Documentation, testing, and teammate-enablement are emphasized.
  • Visa sponsorship is available, with support through the visa process for qualified candidates.
  • Relocation support is available as needed.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →