2 месяца назад
Site Reliability Engineer
220 000 - 300 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer: Improving reliability and scaling cloud infrastructure platform with an accent on operational processes, system performance, and production incident response. Focus on designing architectural changes, fostering reliability culture, and managing large-scale Kubernetes clusters and databases.
Location
Location: Must be able to work in-person in New York or Stockholm offices
Salary
Salary: $220K – $300K per year
Company
builds the new infrastructure layer for AI, providing instant GPU access, sub-second container starts, and native storage for production-ready AI workloads at scale.
What you will do
- Identify architectural changes to improve reliability and performance
- Foster a culture of reliability across the engineering organization
- Define and implement operational processes such as deployments and upgrades
- Operate systems like Kubernetes, Postgres, Redis
- Participate in on-call rotations and respond to production incidents
Requirements
- 5+ years of experience writing high-quality production code
- 2+ years of on-call experience for critical production services
- Strong cloud skills with deep familiarity with at least one hyperscaler cloud (AWS preferred)
- Experience with auto scaling, fleet management, and capacity planning at scale
- Experience operating databases, monitoring, CI/CD, and other infrastructure at scale
- Experience owning and scaling Kubernetes clusters to thousands of nodes is a plus
- Experience with systems safety research and control theory is a plus
- Ability to work in-person in NYC or Stockholm offices
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
13 дней назад
Senior Site Reliability Engineer (Healthcare)
200 000 - 240 000$
14 дней назад
Senior SRE (Site Reliability Engineer) – Modernized Application Operations
145 000 - 170 000$
13 дней назад
Site Reliability Engineering (SRE) Manager (Azure)
139 700 - 232 900$
13 дней назад
Site Reliability Engineer (AI)
200 000 - 400 000$
Camunda
14 дней назад
Senior Site Reliability Engineer (Kubernetes)
149 800 - 241 500$