Назад
обновлено 5 дней назад

Software Engineer (Observability)

165 000 - 330 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer (Observability) (Python/Rust/Go): Building high-throughput telemetry ingest and storage pipelines, instrumentation libraries, and diagnostic tools for multi-cloud AI infrastructure with an accent on metrics, logs, traces, and operational reliability. Focus on scaling observability platforms, improving alerting and SLO infrastructure, reducing telemetry costs, and applying AI to root-cause analysis and anomaly detection.

Location: Hybrid in San Francisco, Montreal, New York, Seattle, or Toronto

Salary: $165,000–$330,000 USD per year, plus equity

Company

Baseten provides applied AI research, infrastructure, and developer tooling that help AI companies bring machine learning models into production.

What you will do

  • Design and build scalable telemetry ingest and storage pipelines for metrics, logs, and traces across multi-cloud infrastructure.
  • Own and evolve observability platforms, migrations, and architectural improvements that increase reliability, reduce costs, and support organizational growth.
  • Build instrumentation libraries, SDKs, and integrations that help engineering teams emit high-quality telemetry.
  • Develop alerting and SLO infrastructure that enables teams to monitor reliability targets with minimal noise.
  • Partner with Inference, Product, and Infrastructure teams to improve operational visibility and incident response.

Requirements

  • Deep experience with at least one observability area: metrics, logging, tracing, or error analytics.
  • Understanding of high-throughput data pipelines, columnar storage engines, and telemetry ingestion and querying tradeoffs at scale.
  • Experience with platforms such as Prometheus, Grafana, ClickHouse, OpenTelemetry, or similar systems.
  • Strong proficiency in at least one of Python, Rust, or Go.
  • Excellent communication skills and the ability to work independently on ambiguous, high-impact infrastructure challenges.
  • Interest in applying AI and LLMs to root cause analysis, anomaly detection, or intelligent alerting.

Culture & Benefits

  • Competitive compensation with meaningful equity.
  • U.S. employees and dependents receive full medical, dental, and vision insurance coverage.
  • Flexible PTO and a company-wide Winter Break from Christmas Eve through New Year's Day.
  • Paid parental leave and a fertility and family-building stipend.
  • Company-facilitated 401(k) for U.S. employees.
  • Exposure to ML startups and opportunities for learning and networking.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →