3 дня назад
Senior Observability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Observability Engineer (AI): Building and operating shared observability services for bare-metal servers, Kubernetes, GPU clusters, networks, and cloud-native applications with an accent on Grafana, VictoriaMetrics, Prometheus, and OpenTelemetry. Focus on designing telemetry pipelines, improving incident detection and response, automating infrastructure with GitOps, and developing AI-assisted monitoring capabilities.
Location: London, UK or remote in the EU; hybrid setup
Company
is building a full-stack AI cloud spanning data centers, hardware, and a cloud platform for AI teams.
What you will do
- Design, scale, and operate shared metrics, logs, and tracing services across infrastructure and workloads.
- Manage declarative Kubernetes deployments with Argo CD, GitLab CI/CD, and GitHub CI/CD.
- Automate host and virtual machine configuration with Ansible and SaltStack.
- Build telemetry pipelines and dashboards for bare-metal systems, networks, GPU clusters, AI training runs, and inference services.
- Define OpenTelemetry instrumentation, context propagation, and Collector pipelines for Go, Node, and Python services.
- Develop dashboards and alerts, improve PagerDuty workflows, participate in on-call rotations, and reduce incident diagnosis and recovery time.
Requirements
- Hands-on experience with Grafana and production telemetry systems, including VictoriaMetrics, VictoriaLogs, or comparable platforms.
- Strong Linux fundamentals, production Kubernetes experience, and experience monitoring distributed and cloud-native systems.
- Practical experience with OpenTelemetry instrumentation, metrics, traces, structured logs, context propagation, and Collector pipelines.
- Experience with declarative deployments, CI/CD, and configuration management using tools such as Argo CD, GitLab CI/CD, Ansible, and SaltStack.
- Experience investigating production incidents, maintaining alerts, and sharing operational responsibilities.
- Ability to explain technical decisions, guide teams, and support shared observability practices.
Nice to have
- Experience operating large-scale observability systems across multiple data centers.
- Experience monitoring GPU clusters and AI/ML workloads, including NVIDIA DCGM.
- Experience with hardware and network telemetry such as IPMI/Redfish, NetFlow/sFlow, SNMP, gNMI, or eBPF.
- Experience across bare-metal, virtual machine, and Kubernetes environments.
- Experience applying LLMs, ML, or agent-based systems to automation and incident investigation.
Culture & Benefits
- Full-time, permanent employment.
- Cash and equity compensation with local benefits.
- Work alongside engineers, researchers, and partners across the global AI ecosystem.
- International environment with more than 40 nationalities.
Hiring process
- Applications are submitted through the Careers page.
- Hiring proceeds without an artificial deadline and moves forward when the right candidate is identified.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Senior Cloud Operations Engineer (AWS)
9 дней назад
Senior Systems Engineer (AWS)
Teleport
3 дня назад
Senior Solutions Engineer (AI Infrastructure)
5 дней назад
AI Platform Engineer
9 дней назад
DevOps Engineer (Kubernetes)
Runpod
10 дней назад
Datacenter Infrastructure Specialist (AI)
105 000 - 140 000$