Назад
21 час назад

Web Crawling Research Engineer (AI)

350 000 - 475 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
lead
Страна
US
vacancy_detail.hirify_telegram_tooltipВакансия из Telegram канала -

Мэтч & Сопровод

Покажет вашу совместимость и напишет письмо

Описание вакансии

TL;DR
Web Crawling Research Engineer (AI): Build and scale internet-scale web-crawling systems, ingestion infrastructure, and pipelines that transform raw crawls into pretraining data with an accent on distributed collection, extraction, deduplication, and data quality. Focus on designing specialized crawlers, improving petabyte-scale reliability and efficiency, assessing effects on model performance, and setting technical direction.

Web Crawling Research Engineer

Company

Thinking Machines Lab

Conditions

6 days agoSalary: 350K - 475K

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build and own web-crawling systems, from distributed internet-scale collection through filtering, deduplication, and data retention decisions. You will develop the crawler, its scalable infrastructure, and pipelines that transform raw crawls into usable pretraining data. You will create specialized crawlers, improve reliability and efficiency at petabyte scale, assess how data changes affect model performance, and help set technical direction.

Requirements

  • 8+ years designing, building, and scaling web crawlers, scrapers, or large-scale distributed data-acquisition systems
  • Track record of owning crawler or data-acquisition infrastructure at internet scale
  • Strong software engineering skills in Python, Go, or Rust
  • Experience with distributed systems
  • Working knowledge of practical and legal considerations for large-scale web data collection, including robots.txt, rate limiting, and licensing

Responsibilities

  • Design and scale web crawler and ingestion infrastructure for pretraining data
  • Build pipelines for large-scale extraction, deduplication, and data-quality filtering
  • Build specialized crawlers for high-value or hard-to-reach data sources
  • Work with the pretraining team to assess how crawled-data changes affect model performance
  • Improve the reliability and efficiency of crawling and ingestion infrastructure at petabyte scale
  • Set technical direction and bring other engineers up to speed

Benefits

  • Health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support as needed

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →

Текст вакансии взят без изменений

Источник -