Назад
Company hidden
10 часов назад

Senior Observability & Telemetry Engineer (GPU/AI Infrastructure)

Формат работы
remote (только Europe)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/US/Japan +4 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Observability & Telemetry Engineer (GPU/AI Infrastructure): Design and operate low-latency, high-scale telemetry pipelines for distributed GPU cloud infrastructure and edge deployments with an accent on metrics, logs, traces, and infrastructure performance visibility. Focus on building observability for GPU clusters, networking, storage, and AI workloads, implementing SLOs and anomaly detection, and diagnosing complex distributed-system bottlenecks.

Location: Europe / EMEA, remote

Company

hirify.global's Radian Arc, now part of InferX, provides an IaaS platform for cloud gaming, artificial intelligence, and machine learning applications within telecommunications carrier networks.

What you will do

  • Design and operate scalable, low-latency telemetry pipelines for metrics, logs, and traces across distributed GPU infrastructure and edge deployments.
  • Build telemetry storage, dashboards, monitoring tools, and customer-facing performance analysis capabilities.
  • Instrument GPU clusters, inference and training workloads, compute, storage, and networking systems.
  • Develop network telemetry collectors and exporters using Go or Python and protocols including gNMI, SNMP, and streaming telemetry.
  • Implement alerting, anomaly detection, SLOs, SLIs, and integrations with incident-management workflows.
  • Collaborate with platform, networking, storage, compute, and operations teams while mentoring engineers on observability practices.

Requirements

  • Experience operating large distributed infrastructure platforms and production-scale observability systems.
  • Strong programming skills in Go, Python, or Rust.
  • Experience with metrics, logging, tracing, alerting, dashboards, and large-scale time-series data platforms.
  • Experience with GPU cloud, HPC, or AI infrastructure and monitoring training or inference workloads.
  • Knowledge of Kubernetes, automation, CI/CD, complex networking environments, and telemetry protocols such as gNMI and SNMP.
  • Strong analytical skills with the ability to diagnose distributed-system performance issues and turn telemetry into actionable insights.

Culture & Benefits

  • Permanent, full-time contract with an ASAP start.
  • Compensation package based on expertise, experience, transferable skills, business needs, and market demands.
  • Flexible, international, and hybrid-friendly work environment.
  • Opportunity to join a fast-growing scale-up focused on AI cloud and GPU infrastructure.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →