Назад
Company hidden
12 дней назад

Software Engineer Focused on Observability (AI)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer Focused on Observability (AI): Building observability instrumentation, telemetry pipelines, and internal platforms for large-scale AI inference systems with an accent on metrics, logging, tracing, alerting, and distributed systems reliability. Focus on designing scalable monitoring infrastructure, reducing MTTR through root-cause analysis, and balancing telemetry signal, cost, noise, and performance impact.

Location: Sunnyvale, California, United States; hybrid

Company

hirify.global builds AI hardware and large-scale inference systems designed to deliver high-performance training and inference.

What you will do

  • Design and implement observability instrumentation across services and platforms.
  • Build and maintain scalable telemetry pipelines for metrics, logs, and distributed traces.
  • Develop internal observability platforms, libraries, and developer tooling.
  • Define and operationalize SLIs, SLOs, monitoring, and alerting strategies.
  • Partner with engineers to make systems debuggable by design and reduce MTTR during incidents.
  • Create actionable dashboards and alerts while balancing telemetry signal, cost, noise, and performance impact.

Requirements

  • Strong experience in backend or systems software engineering.
  • Proficiency in one or more of Go, C++, Rust, Java, or Python.
  • Solid understanding of distributed systems, networking fundamentals, concurrency, and performance tradeoffs.
  • Hands-on experience with metrics, logs, distributed tracing, production monitoring, and alerting.
  • Experience designing high-signal alerts, scalable telemetry pipelines, and service-level indicators and objectives.
  • Familiarity with OpenTelemetry, Prometheus, Grafana, Datadog, Elastic, Jaeger, Tempo, or similar tools.

Nice to have

  • Experience with high-performance computing, AI/ML systems, or inference platforms.
  • Hardware-aware observability involving accelerators, GPUs, or custom hardware.
  • Prior SRE or platform engineering experience.
  • Experience debugging large-scale production incidents or building internal developer platforms and shared libraries.

Culture & Benefits

  • Work on an AI platform designed beyond the constraints of GPUs.
  • Opportunities to publish and open-source AI research.
  • Work with one of the fastest AI supercomputers in the world.
  • Job stability combined with startup vitality.
  • Non-corporate culture focused on individual beliefs, learning, growth, and inclusion.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →