Назад
Company hidden
12 часов назад

Senior Site Reliability Engineer (AI)

130 000 - 190 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (AI) (Python/AWS): Owning the reliability, performance, and observability of CloudZero's real-time ingestion platform processing billions of events across AWS, Azure, and GCP with an accent on event-driven systems, serverless infrastructure, and cross-team SLOs. Focus on building reliability tooling, automating deployments and recovery, instrumenting critical paths, and designing resilient systems without Kubernetes or container infrastructure.

Location: Hybrid in Boston, MA or San Francisco, CA

Salary: $130,000–$190,000 per year plus equity

Company

hirify.global is an AI ROI company providing a financial control plane that connects AI and cloud spending to business outcomes in real time.

What you will do

  • Own reliability for the real-time Kafka ingestion path, including cross-team SLOs, failure modes, and architectural improvements.
  • Define critical-path release standards, monitor error budgets, and build observability so failures are detected before customers are affected.
  • Develop production Python tooling such as load generators, fault-injection harnesses, SLO libraries, deployment safety checks, internal services, and automation.
  • Design and maintain CloudFormation and SAM modules for reliable, cost-efficient serverless infrastructure.
  • Automate deployments, scaling, backups, limit changes, and infrastructure operations without relying on cloud consoles.
  • Partner with product engineering teams to design resilient services, improve deployment pipelines, and drive adoption of reliability practices across 40+ engineers.

Requirements

  • Strong production Python experience, including systems that are owned, tested, and maintained at scale.
  • Experience operating asynchronous, event-driven distributed systems and reasoning about back-pressure, consumer lag, replay, poison messages, and partial failure.
  • Typically 5+ years building and operating distributed systems in AWS, with ownership of reliability outcomes.
  • Hands-on Infrastructure as Code experience with CloudFormation and SAM, or equivalent depth in Terraform or Pulumi.
  • Experience instrumenting production systems with monitoring tools such as Sumo Logic, Datadog, Prometheus, or Splunk, plus production debugging under pressure.
  • Ability to explain complex technical issues clearly, document systems thoroughly, and influence teams without direct authority.

Nice to have

  • Experience building chaos engineering or load-testing practices.
  • Internal developer portal experience with Cortex or Backstage.
  • Test automation, ephemeral test environments, or GitHub Actions at scale.
  • Experience building LLM-backed tooling used by engineers.

Culture & Benefits

  • Collaborative, fast-moving environment with emphasis on ownership, creativity, curiosity, and reliable system design.
  • Light on-call responsibilities through a weekly shared-infrastructure rotation; feature teams support their own services.
  • Opportunity to work on serverless infrastructure processing billions of events daily across AWS, Azure, and GCP.
  • Work focused on measurable customer impact and complex cloud and AI cost challenges.
  • Equity included in the compensation package.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →