обновлено 1 день назад
Senior Site Reliability Engineer (AI Infrastructure Operations)
170 000 - 265 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (AI Infrastructure Operations) (AI infrastructure): Owning reliability for critical production services and automation across GPU cloud infrastructure with an accent on SLOs, observability, incident response, and scalable systems engineering. Focus on designing resilient distributed systems, leading complex incidents, building toil-reducing automation, and improving Kubernetes, Linux, networking, and bare-metal operations.
Location: Houston, San Francisco, or Seattle
Salary: $170,000–$265,000 USD base salary per year, with potential bonus and equity.
Company
Nscale operates a GPU cloud for AI-native startups and global enterprises, providing infrastructure from bare metal through platform services.
What you will do
- Own reliability for critical production services and set reliability direction across the platform.
- Define and maintain SLOs, incident processes, observability, alerting, and sustainable on-call practices.
- Lead architecture and design reviews to build reliability into systems from the beginning.
- Lead complex production incidents, identify root causes, and implement lasting fixes.
- Build tooling and automation that reduces operational toil across the SRE team.
- Mentor SREs through design reviews, pairing, and incident debriefs while raising engineering standards.
Requirements
- 6–10 years of experience in SRE, systems engineering, or software engineering with production ownership at scale.
- Strong software engineering skills in Python, Go, or a similar language.
- Deep knowledge of Linux, networking, distributed systems, Kubernetes, and virtualized or bare-metal environments.
- Experience running AI or GPU workloads or high-performance computing environments.
- Experience implementing SLOs, observability, large-scale alerting, incident management, and sustainable on-call practices.
- Track record of serving as a senior technical voice during incidents and design reviews and mentoring other engineers.
Nice to have
- Familiarity with high-performance networking, including InfiniBand and RDMA.
Culture & Benefits
- Shared on-call rotation with a focus on reducing operational load over time.
- Direct ownership of reliability practices across the platform.
- Flexible approach to organizing the workday.
- Competitive base salary with compensation reviewed every 12 months.
- Potential bonus, equity, medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Reddit
8 дней назад
Staff Site Reliability Engineer, Ads
217 000 - 303 900$
6 дней назад
Sr. Staff Lead Site Reliability Engineer (AWS)
220 000 - 330 000$
5 дней назад
DevOps & SRE Engineer (Kubernetes)
100 000 - 150 000$
2 дня назад
Senior Site Reliability Engineer (AI)
191 000 - 226 000$
6 дней назад
Sr. Staff Lead Site Reliability Engineer (R5803)
183 000 - 275 000$
6 дней назад
Software Engineer (Cloud Infrastructure/SRE)
147 900 - 220 000$