Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Observability Platform Engineer (AI): Designing, building, and scaling observability systems for GPU clusters, AI workloads, and supporting infrastructure with an accent on metrics, logs, traces, alerting, and scalable data pipelines. Focus on reducing signal noise and cardinality, improving reliability, integrating observability across Kubernetes-based platforms, and enabling fast debugging.
Location: US
Salary: $160,000–$230,000 USD per year, plus potential bonus, equity, and/or commission.
Company
Nscale provides high-performance, cost-effective GPU cloud infrastructure for AI startups and enterprise customers.
What you will do
- Design, build, and operate scalable observability systems covering metrics, logs, traces, and alerting.
- Make architectural decisions around observability tooling, data pipelines, storage, and retention.
- Improve signal quality by reducing noise and cardinality and refining alerting practices.
- Integrate observability into services and platforms in collaboration with SRE, infrastructure, and AI/ML teams.
- Develop reusable patterns, libraries, and best practices across engineering teams.
- Participate in incident response and postmortems, evaluate tools, and mentor engineers through reviews and knowledge sharing.
Requirements
- 5+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles.
- Experience operating and scaling observability systems in production.
- Strong understanding of metrics, logs, traces, alerting, and SLOs.
- Hands-on experience with several of Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, or Elastic.
- Production programming experience with Python, Go, or a similar language.
- Experience with Kubernetes-based infrastructure and Infrastructure-as-Code tools such as Terraform or Ansible.
Nice to have
- Experience with observability data pipelines using Kafka, Vector, Fluent Bit, or similar tools.
- Exposure to AI/ML infrastructure or GPU-based systems.
- Familiarity with performance monitoring for distributed systems.
- Experience improving developer experience through observability tooling.
Culture & Benefits
- Hands-on role with influence over platform design and evolution.
- Culture focused on innovation, ownership, accountability, openness, and transparency.
- Medical, dental, and vision benefits may be available.
- Flexible paid time off, parental leave, and retirement plan participation may be available.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Production Engineer (AI)
155 000 - 185 000$
5 дней назад
Software Engineer (AI Infrastructure)
145 000 - 220 000$
5 дней назад
Senior DevOps Engineer / Site Reliability Engineer (AI)
170 000 - 220 000$
4 дня назад
Platform Engineer (AI)
187 000 - 250 000$
4 дня назад
Software Engineer (AI)
163 000 - 178 000$
6 дней назад
Principal Software Engineer (Infra/Platform)
180 000 - 235 000$