обновлено 3 дня назад
Staff Site Reliability Engineer (AI)
220 000 - 260 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Site Reliability Engineer (AWS/AI): Leading the evolution of the platform infrastructure to ensure reliability, observability, and scalability with an accent on AWS migration and CI/CD automation. Focus on designing resilient distributed systems, improving incident response processes, and enhancing developer experience through ephemeral environments.
Location: On-site in the Soho office in New York City, United States
Salary: $220,000–$260,000 annually, plus equity
Company
builds an AI operating system for revenue and accounting teams, combining accounting expertise with agents and applications that automate revenue workflows with controls, auditability, and human oversight.
What you will do
- Set the direction for AWS infrastructure and evolve the platform, including migration from ECS/Fargate to a more scalable runtime.
- Build and improve CI/CD systems, GitHub Actions workflows, ephemeral environments, and preview deployments.
- Establish observability standards across metrics, logs, and tracing, including dashboards, alert hygiene, and SLO development.
- Define reliability standards, SLIs, SLOs, and error budgets across services.
- Lead high-severity incidents, postmortems, and actionable reliability improvements.
- Partner with engineering and product teams, automate operational work, and mentor engineers on resilient system design.
Requirements
- 10+ years of experience in SRE, infrastructure, or backend engineering.
- Production experience running AWS-based systems and leading platform-level changes.
- Strong software engineering experience in one or more modern programming languages.
- Expertise operating distributed systems at scale, with deep experience in AWS, observability tooling, and CI/CD.
- Ability to assess risk, rollback strategy, blast radius, and feedback loops while working across engineering teams.
- Must work on-site from the Soho office in New York City.
Culture & Benefits
- Senior individual contributor role with influence across engineering and product.
- Unlimited PTO and parental leave of up to 12 weeks.
- Up to 100% employer-covered monthly medical, dental, and vision premiums.
- Lunch provided through Sharebite, with dinner available for later office days.
- Tax-free commuter and parking benefits, voluntary insurance plans, an Employee Assistance Program, and free One Medical membership.
- 401(k) plan and equity compensation.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Staff Site Reliability Engineer (Cybersecurity)
199 750 - 270 000$
8 дней назад
Senior Site Reliability Engineer (Healthcare)
200 000 - 240 000$
8 дней назад
Site Reliability Engineering (SRE) Manager (Azure)
139 700 - 232 900$
8 дней назад
Senior Manager, Site Reliability Engineering (AI Ops)
222 000 - 300 500$
6 дней назад
Principal Site Reliability Engineer (Kubernetes)
190 000 - 220 000$
Replit
5 дней назад
Engineering Manager (SRE)
250 000 - 325 000$