Назад
4 дня назад

Senior Observability Platform Engineer (AI)

160 000 - 230 000$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Observability Platform Engineer (AI): Designing, building, and scaling observability systems for GPU clusters, AI workloads, and supporting infrastructure with an accent on metrics, logs, traces, alerting, and scalable data pipelines. Focus on reducing signal noise and cardinality, improving reliability, integrating observability across Kubernetes-based platforms, and enabling fast debugging.

Location: US

Salary: $160,000–$230,000 USD per year, plus potential bonus, equity, and/or commission.

Company

Nscale provides high-performance, cost-effective GPU cloud infrastructure for AI startups and enterprise customers.

What you will do

  • Design, build, and operate scalable observability systems covering metrics, logs, traces, and alerting.
  • Make architectural decisions around observability tooling, data pipelines, storage, and retention.
  • Improve signal quality by reducing noise and cardinality and refining alerting practices.
  • Integrate observability into services and platforms in collaboration with SRE, infrastructure, and AI/ML teams.
  • Develop reusable patterns, libraries, and best practices across engineering teams.
  • Participate in incident response and postmortems, evaluate tools, and mentor engineers through reviews and knowledge sharing.

Requirements

  • 5+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles.
  • Experience operating and scaling observability systems in production.
  • Strong understanding of metrics, logs, traces, alerting, and SLOs.
  • Hands-on experience with several of Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, or Elastic.
  • Production programming experience with Python, Go, or a similar language.
  • Experience with Kubernetes-based infrastructure and Infrastructure-as-Code tools such as Terraform or Ansible.

Nice to have

  • Experience with observability data pipelines using Kafka, Vector, Fluent Bit, or similar tools.
  • Exposure to AI/ML infrastructure or GPU-based systems.
  • Familiarity with performance monitoring for distributed systems.
  • Experience improving developer experience through observability tooling.

Culture & Benefits

  • Hands-on role with influence over platform design and evolution.
  • Culture focused on innovation, ownership, accountability, openness, and transparency.
  • Medical, dental, and vision benefits may be available.
  • Flexible paid time off, parental leave, and retirement plan participation may be available.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →