12 дней назад
Software Engineer Focused on Observability (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineer Focused on Observability (AI): Building observability instrumentation, telemetry pipelines, and internal platforms for large-scale AI inference systems with an accent on metrics, logging, tracing, alerting, and distributed systems reliability. Focus on designing scalable monitoring infrastructure, reducing MTTR through root-cause analysis, and balancing telemetry signal, cost, noise, and performance impact.
Location: Sunnyvale, California, United States; hybrid
Company
builds AI hardware and large-scale inference systems designed to deliver high-performance training and inference.
What you will do
- Design and implement observability instrumentation across services and platforms.
- Build and maintain scalable telemetry pipelines for metrics, logs, and distributed traces.
- Develop internal observability platforms, libraries, and developer tooling.
- Define and operationalize SLIs, SLOs, monitoring, and alerting strategies.
- Partner with engineers to make systems debuggable by design and reduce MTTR during incidents.
- Create actionable dashboards and alerts while balancing telemetry signal, cost, noise, and performance impact.
Requirements
- Strong experience in backend or systems software engineering.
- Proficiency in one or more of Go, C++, Rust, Java, or Python.
- Solid understanding of distributed systems, networking fundamentals, concurrency, and performance tradeoffs.
- Hands-on experience with metrics, logs, distributed tracing, production monitoring, and alerting.
- Experience designing high-signal alerts, scalable telemetry pipelines, and service-level indicators and objectives.
- Familiarity with OpenTelemetry, Prometheus, Grafana, Datadog, Elastic, Jaeger, Tempo, or similar tools.
Nice to have
- Experience with high-performance computing, AI/ML systems, or inference platforms.
- Hardware-aware observability involving accelerators, GPUs, or custom hardware.
- Prior SRE or platform engineering experience.
- Experience debugging large-scale production incidents or building internal developer platforms and shared libraries.
Culture & Benefits
- Work on an AI platform designed beyond the constraints of GPUs.
- Opportunities to publish and open-source AI research.
- Work with one of the fastest AI supercomputers in the world.
- Job stability combined with startup vitality.
- Non-corporate culture focused on individual beliefs, learning, growth, and inclusion.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
12 дней назад
Infrastructure Engineer (AI)
180 000 - 300 000$
12 дней назад
Lead Observability Engineer (AWS)
5 дней назад
Monitoring Engineer (Observability)
75 000 - 85 000$
12 дней назад
Infrastructure Engineer (AI)
200 000 - 250 000$
12 дней назад
Senior AI Platform Engineer (AI)
Ramp
12 дней назад
Production Engineer (AI Infrastructure)
240 000 - 330 000$