Назад
1 месяц назад

Senior Product Manager, Observability (AI Cloud)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Product Manager, Observability (AI Cloud): Owns the observability platform for a globally distributed GPU fleet, including telemetry pipelines, log and metrics aggregation, trace collection, APIs, and dashboards with an accent on fleet health, incident response, and alerting at scale. Focus on defining observability architecture and UX trade-offs, reducing alert noise, expanding telemetry coverage across infrastructure, and improving time-to-detect, time-to-resolve, and platform reliability.

Location: UK

Company

Nscale is building a vertically integrated GenAI cloud platform spanning data centres, software, and applications for the AI stack.

What you will do

  • Own the roadmap for the observability platform, including telemetry pipelines, log and metrics aggregation, trace collection, customer-facing APIs, and dashboards.
  • Define how logs, metrics, and traces are captured from physical infrastructure, aggregated, and surfaced for fleet management and incident response.
  • Set and optimise the alerting strategy by improving signal quality, reducing noise, and routing actionable alerts.
  • Prioritise telemetry requirements as the fleet expands across new hardware, sites, and deployment types.
  • Use incident reviews and site operations insights to turn manual effort and visibility gaps into platform capabilities.
  • Define reliability and effectiveness metrics while mentoring junior product managers and improving product documentation and decision-making.

Requirements

  • 5–8 years of product management experience, including ownership of observability, infrastructure, or operations-facing products.
  • Experience building observability stacks that capture and surface logs, metrics, and traces at scale.
  • Hands-on experience with Prometheus, Loki, Mimir, Datadog, Grafana, or OpenTelemetry.
  • Experience with data centre or infrastructure deployment tooling, provisioning workflows, networking automation, or zero-touch deployment pipelines.
  • Ability to work with operators and delivery teams, including SREs, infrastructure engineers, project controllers, and data centre technicians.
  • Strong technical fluency in telemetry pipelines, time-series storage, alerting systems, observability integrations, and architecture trade-offs.

Nice to have

  • Experience with bare-metal provisioning tools such as OpenStack Ironic or MAAS, or network automation tools such as NetBox or Nautobot.
  • Degree in computer science or engineering, or prior experience as an engineer, SRE, or infrastructure operator.
  • Familiarity with GPU infrastructure, accelerated computing, data centre operations, or hyperscaler-scale deployments.
  • Experience with Jira Service Management, ServiceNow, Zendesk, or Freshservice.
  • Experience in high-growth environments where the product is built alongside the fleet it monitors.

Culture & Benefits

  • Work in a culture focused on innovation, ownership, accountability, transparency, collaboration, adaptability, and resilience.
  • Contribute to an inclusive, diverse, and equitable workplace.
  • Work closely with Fleet Software, Network Engineering, Data Centre Operations, engineering, design, go-to-market, and customer teams.
  • Work on software that turns contracted capacity into live GPU infrastructure for an AI cloud platform.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →