Назад
Company hidden
3 дня назад

Senior Observability Engineer (AI)

Формат работы
remote (только Europe)/hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK/Europe
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Observability Engineer (AI): Building and operating shared observability services for bare-metal servers, Kubernetes, GPU clusters, networks, and cloud-native applications with an accent on Grafana, VictoriaMetrics, Prometheus, and OpenTelemetry. Focus on designing telemetry pipelines, improving incident detection and response, automating infrastructure with GitOps, and developing AI-assisted monitoring capabilities.

Location: London, UK or remote in the EU; hybrid setup

Company

hirify.global is building a full-stack AI cloud spanning data centers, hardware, and a cloud platform for AI teams.

What you will do

  • Design, scale, and operate shared metrics, logs, and tracing services across infrastructure and workloads.
  • Manage declarative Kubernetes deployments with Argo CD, GitLab CI/CD, and GitHub CI/CD.
  • Automate host and virtual machine configuration with Ansible and SaltStack.
  • Build telemetry pipelines and dashboards for bare-metal systems, networks, GPU clusters, AI training runs, and inference services.
  • Define OpenTelemetry instrumentation, context propagation, and Collector pipelines for Go, Node, and Python services.
  • Develop dashboards and alerts, improve PagerDuty workflows, participate in on-call rotations, and reduce incident diagnosis and recovery time.

Requirements

  • Hands-on experience with Grafana and production telemetry systems, including VictoriaMetrics, VictoriaLogs, or comparable platforms.
  • Strong Linux fundamentals, production Kubernetes experience, and experience monitoring distributed and cloud-native systems.
  • Practical experience with OpenTelemetry instrumentation, metrics, traces, structured logs, context propagation, and Collector pipelines.
  • Experience with declarative deployments, CI/CD, and configuration management using tools such as Argo CD, GitLab CI/CD, Ansible, and SaltStack.
  • Experience investigating production incidents, maintaining alerts, and sharing operational responsibilities.
  • Ability to explain technical decisions, guide teams, and support shared observability practices.

Nice to have

  • Experience operating large-scale observability systems across multiple data centers.
  • Experience monitoring GPU clusters and AI/ML workloads, including NVIDIA DCGM.
  • Experience with hardware and network telemetry such as IPMI/Redfish, NetFlow/sFlow, SNMP, gNMI, or eBPF.
  • Experience across bare-metal, virtual machine, and Kubernetes environments.
  • Experience applying LLMs, ML, or agent-based systems to automation and incident investigation.

Culture & Benefits

  • Full-time, permanent employment.
  • Cash and equity compensation with local benefits.
  • Work alongside engineers, researchers, and partners across the global AI ecosystem.
  • International environment with more than 40 nationalities.

Hiring process

  • Applications are submitted through the Careers page.
  • Hiring proceeds without an artificial deadline and moves forward when the right candidate is identified.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →