12 часов назад
Senior Site Reliability Engineer (AI)
130 000 - 190 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (AI) (Python/AWS): Owning the reliability, performance, and observability of CloudZero's real-time ingestion platform processing billions of events across AWS, Azure, and GCP with an accent on event-driven systems, serverless infrastructure, and cross-team SLOs. Focus on building reliability tooling, automating deployments and recovery, instrumenting critical paths, and designing resilient systems without Kubernetes or container infrastructure.
Location: Hybrid in Boston, MA or San Francisco, CA
Salary: $130,000–$190,000 per year plus equity
Company
is an AI ROI company providing a financial control plane that connects AI and cloud spending to business outcomes in real time.
What you will do
- Own reliability for the real-time Kafka ingestion path, including cross-team SLOs, failure modes, and architectural improvements.
- Define critical-path release standards, monitor error budgets, and build observability so failures are detected before customers are affected.
- Develop production Python tooling such as load generators, fault-injection harnesses, SLO libraries, deployment safety checks, internal services, and automation.
- Design and maintain CloudFormation and SAM modules for reliable, cost-efficient serverless infrastructure.
- Automate deployments, scaling, backups, limit changes, and infrastructure operations without relying on cloud consoles.
- Partner with product engineering teams to design resilient services, improve deployment pipelines, and drive adoption of reliability practices across 40+ engineers.
Requirements
- Strong production Python experience, including systems that are owned, tested, and maintained at scale.
- Experience operating asynchronous, event-driven distributed systems and reasoning about back-pressure, consumer lag, replay, poison messages, and partial failure.
- Typically 5+ years building and operating distributed systems in AWS, with ownership of reliability outcomes.
- Hands-on Infrastructure as Code experience with CloudFormation and SAM, or equivalent depth in Terraform or Pulumi.
- Experience instrumenting production systems with monitoring tools such as Sumo Logic, Datadog, Prometheus, or Splunk, plus production debugging under pressure.
- Ability to explain complex technical issues clearly, document systems thoroughly, and influence teams without direct authority.
Nice to have
- Experience building chaos engineering or load-testing practices.
- Internal developer portal experience with Cortex or Backstage.
- Test automation, ephemeral test environments, or GitHub Actions at scale.
- Experience building LLM-backed tooling used by engineers.
Culture & Benefits
- Collaborative, fast-moving environment with emphasis on ownership, creativity, curiosity, and reliable system design.
- Light on-call responsibilities through a weekly shared-infrastructure rotation; feature teams support their own services.
- Opportunity to work on serverless infrastructure processing billions of events daily across AWS, Azure, and GCP.
- Work focused on measurable customer impact and complex cloud and AI cost challenges.
- Equity included in the compensation package.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Site Reliability Engineer
87 400 - 123 400$
4 дня назад
Senior Site Reliability Engineer (GovCloud)
117 000 - 209 330$
5 дней назад
Senior Site Reliability Engineer (AI Infrastructure)
215 000 - 275 000$
5 дней назад
Senior Software Engineer, Site Reliability Engineering (AWS)
153 000 - 210 000$
5 дней назад
Sr. Staff Production Engineer (Cloud Infrastructure)
143 500 - 205 000$
SandboxAQ
4 дня назад
Staff Platform Engineer (AI)
121 600 - 228 000$