Назад
Company hidden
4 дня назад

Staff Software Engineer (Agentic Search, Crawler)

Формат работы
remote (Global)/hybrid
Тип работы
fulltime
Английский
b2
Страна
UK/US/Netherlands +1 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Software Engineer (Agentic Search, Crawler): Building web-scale crawling and content acquisition infrastructure for an agentic AI search engine with an accent on distributed systems, data ingestion, and crawl freshness. Focus on designing URL discovery, scheduling, deduplication, extraction, and orchestration systems that operate reliably across billions of URLs and high-throughput pipelines.

Location: Hybrid in Amsterdam, London, Tel Aviv, or New York; remote work is available across Europe and the US. Relocation support is available.

Company

A publicly traded AI infrastructure company building a full-stack cloud platform for data, model training, GPU orchestration, inference optimization, and production deployment.

What you will do

  • Design, implement, and operate web-scale crawling systems for acquiring and continuously refreshing content.
  • Build ingestion workflows for web content, structured feeds, and partner data sources.
  • Develop URL discovery, crawl scheduling, prioritization, recrawl, deduplication, extraction, and orchestration systems.
  • Define observability and quality metrics covering crawl coverage, freshness, throughput, and content quality.
  • Operate high-throughput infrastructure while managing resource usage, bandwidth, reliability, and cost.
  • Collaborate with indexing and ML teams to align acquired content with retrieval and ranking requirements.

Requirements

  • Experience building backend or distributed systems in production.
  • Deep expertise in Go or C++; Python or Java experience is valuable.
  • Professional experience with web crawling, scraping, content extraction, and web protocols including HTTP, DNS, and TLS.
  • Experience with large-scale systems handling 10k+ RPS, billions of URLs, or high-throughput data pipelines.
  • Strong understanding of scalability, fault tolerance, resource management, production operations, and distributed-system debugging.

Nice to have

  • Experience building streaming data pipelines and event-driven systems.
  • Kafka, Pulsar, NATS, RabbitMQ, or similar messaging platforms.
  • Distributed schedulers, queues, and asynchronous processing systems.
  • Spark, Flink, Beam, or MapReduce.
  • Experience with ad tech, social networks, search engines, or other large-scale content platforms.

Culture & Benefits

  • Opportunity to join a founding Search engineering team and shape a next-generation search platform for AI systems.
  • Ownership of architecture and technical strategy for a critical part of the AI stack.
  • Competitive base salary and annual bonus.
  • Healthcare and local benefits package.
  • Resources and stability provided by a publicly traded AI infrastructure company.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →