Назад
1 день назад

Software Engineer (AI Data Infrastructure)

350 000 - 475 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer (AI Data Infrastructure) (Web Crawling/Distributed Systems): Building and scaling web-crawling systems and ingestion pipelines for frontier-model pretraining data with an accent on internet-scale collection, extraction, deduplication, filtering, and petabyte-scale processing. Focus on designing reliable distributed infrastructure, applying machine learning to crawl selection and data quality, and setting technical direction for crawling systems.

Location: Hybrid role based in San Francisco, California, United States

Salary: $350,000–$475,000 USD annual salary

Company

Thinking Machines Lab builds AI systems that extend human judgment and communication, including frontier models, model customization tools, and human-AI interfaces.

What you will do

  • Design and scale web crawlers and ingestion infrastructure for Inkling's pretraining data.
  • Build large-scale extraction, deduplication, and data-quality filtering pipelines.
  • Develop specialized crawlers for high-value and hard-to-reach data sources.
  • Work with pretraining and data teams to evaluate how crawled data affects model performance.
  • Improve the reliability and efficiency of crawling and ingestion infrastructure at petabyte scale.
  • Set technical direction and help other engineers develop expertise in crawling systems.

Requirements

  • 8+ years of experience designing, building, and scaling web crawlers, scrapers, or large-scale distributed data-acquisition systems.
  • Demonstrated ownership of crawler or data-acquisition infrastructure at internet scale.
  • Strong software engineering skills in Python, Go, or Rust, with practical distributed-systems experience.
  • Working knowledge of practical and legal web data collection, including robots.txt, rate limiting, and licensing.
  • Ability to work in a hybrid role based in San Francisco, California.

Nice to have

  • Experience applying machine learning to crawl selection, extraction, or data-quality classification at internet scale.
  • Experience setting technical direction for crawling, data-acquisition, or search infrastructure teams.
  • Experience designing petabyte-scale storage and processing systems.
  • Open-source contributions to crawling, scraping, or data-infrastructure tools.
  • Background in search-engine crawling or indexing, or on a frontier AI lab's data-acquisition team.

Culture & Benefits

  • Health, dental, and vision benefits.
  • Unlimited paid time off and paid parental leave.
  • Relocation support as needed.
  • Visa sponsorship is available, with support through the visa process for qualified candidates.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →