10 часов назад
Senior Observability & Telemetry Engineer (GPU/AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Observability & Telemetry Engineer (GPU/AI Infrastructure): Design and operate low-latency, high-scale telemetry pipelines for distributed GPU cloud infrastructure and edge deployments with an accent on metrics, logs, traces, and infrastructure performance visibility. Focus on building observability for GPU clusters, networking, storage, and AI workloads, implementing SLOs and anomaly detection, and diagnosing complex distributed-system bottlenecks.
Location: Europe / EMEA, remote
Company
's Radian Arc, now part of InferX, provides an IaaS platform for cloud gaming, artificial intelligence, and machine learning applications within telecommunications carrier networks.
What you will do
- Design and operate scalable, low-latency telemetry pipelines for metrics, logs, and traces across distributed GPU infrastructure and edge deployments.
- Build telemetry storage, dashboards, monitoring tools, and customer-facing performance analysis capabilities.
- Instrument GPU clusters, inference and training workloads, compute, storage, and networking systems.
- Develop network telemetry collectors and exporters using Go or Python and protocols including gNMI, SNMP, and streaming telemetry.
- Implement alerting, anomaly detection, SLOs, SLIs, and integrations with incident-management workflows.
- Collaborate with platform, networking, storage, compute, and operations teams while mentoring engineers on observability practices.
Requirements
- Experience operating large distributed infrastructure platforms and production-scale observability systems.
- Strong programming skills in Go, Python, or Rust.
- Experience with metrics, logging, tracing, alerting, dashboards, and large-scale time-series data platforms.
- Experience with GPU cloud, HPC, or AI infrastructure and monitoring training or inference workloads.
- Knowledge of Kubernetes, automation, CI/CD, complex networking environments, and telemetry protocols such as gNMI and SNMP.
- Strong analytical skills with the ability to diagnose distributed-system performance issues and turn telemetry into actionable insights.
Culture & Benefits
- Permanent, full-time contract with an ASAP start.
- Compensation package based on expertise, experience, transferable skills, business needs, and market demands.
- Flexible, international, and hybrid-friendly work environment.
- Opportunity to join a fast-growing scale-up focused on AI cloud and GPU infrastructure.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Sr SRE & Automation Engineer (Customer Facing)
14 часов назад
Senior Observability Engineer (AI)
4 дня назад
Telemetry Engineer (OpenTelemetry)
100 000 - 150 000$
4 дня назад
Monitoring Engineer (Observability)
75 000 - 85 000$
5 часов назад
Senior Site Reliability Engineer (Kubernetes)
2 дня назад
Systems Engineer (AI Infrastructure)
140 000 - 225 000$