Назад
5 дней назад

Staff Software Engineer, Observability & Profiling (AI)

325 000 - 390 000GBP
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK/US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Software Engineer, Observability & Profiling (AI): Building scalable observability systems for Anthropic’s multi-cluster AI infrastructure with an accent on telemetry pipelines, continuous profiling, eBPF-based tracing, and cross-signal diagnostics. Focus on designing high-throughput ingest and storage, investigating kernel, network, and accelerator behavior, and applying AI-assisted tooling to improve incident detection and resolution.

Location: London, UK; hybrid policy requires staff to work from one of the offices at least 25% of the time.

Annual salary: £325,000–£390,000 GBP.

Company

Anthropic builds reliable, interpretable, and steerable AI systems intended to be safe and beneficial for users and society.

What you will do

  • Design and build scalable telemetry ingest and storage pipelines for metrics, logs, traces, and error data across multi-cluster infrastructure.
  • Develop low-overhead observability solutions, instrumentation libraries, SDKs, and eBPF-based auto-instrumentation.
  • Own observability platforms and drive migrations and architectural improvements that increase reliability, reduce cost, and support growth.
  • Build cross-signal correlation, unified query interfaces, and AI-assisted diagnostic tools to reduce detection and resolution times.
  • Turn continuous profiling and utilization telemetry into optimization insights across CPU, memory, and accelerator fleets.
  • Partner with Research, Inference, Product, and Infrastructure teams to improve operational visibility and incident response.

Requirements

  • Hands-on experience building and operating large-scale observability or monitoring infrastructure.
  • End-to-end experience with observability signals, from instrumentation and ingest through querying and analysis.
  • Understanding of high-throughput telemetry pipelines and large-scale operational data storage and querying tradeoffs.
  • Ability to investigate below the application layer, including the kernel, network stack, and hardware.
  • Strong communication skills and the ability to work independently and with cross-functional teams.
  • Bachelor’s degree or equivalent education, training, or relevant experience.

Nice to have

  • 10+ years of relevant industry experience.
  • Production experience with eBPF observability, continuous profiling, kernel and syscall debugging, or performance engineering.
  • Experience profiling accelerator workloads or operating high-cardinality metrics systems and large-scale telemetry storage backends.
  • Experience with OpenTelemetry, collector pipelines, and tail-based sampling.
  • Interest in applying AI and LLMs to root cause analysis, anomaly detection, or intelligent alerting.

Culture & Benefits

  • Collaborative work across research, engineering, policy, and business functions.
  • Flexible working hours and an office-focused hybrid environment.
  • Generous vacation and parental leave.
  • Competitive compensation, benefits, and optional equity donation matching.
  • Visa sponsorship is available, with immigration lawyer support, subject to role and candidate eligibility.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →