Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI/GPU) (AI infrastructure): Building automation, observability, and reliability tooling for production AI and GPU workloads with an accent on Linux, networking, distributed systems, and incident response. Focus on defining SLOs and SLIs, troubleshooting high-scale services, and improving availability, scalability, and efficiency through code.
Location: Houston, New York, San Francisco, or Seattle, United States
Salary: $130,000–$200,000 USD per year, with possible bonus and equity
Company
Nscale provides a GPU cloud infrastructure platform for AI-native startups and global enterprises, from bare metal through platform services.
What you will do
- Build and own automation and tooling that keeps the platform running and reduces operational toil.
- Define and maintain SLOs, SLIs, dashboards, metrics, logs, and alerting for service health.
- Lead incident response, troubleshoot production issues, perform root cause analysis, and run post-incident reviews.
- Investigate and resolve performance and reliability problems across Linux, networking, and distributed services.
- Partner with Engineering, Networking, and Infrastructure teams to improve reliability across the stack.
- Improve availability, scalability, and efficiency through code and participate in the on-call rotation.
Requirements
- 3–6 years of experience in SRE, systems engineering, or software engineering, including production experience in a data center or cloud environment.
- Strong programming skills in Python, Go, or a similar language, with a focus on automation.
- Solid knowledge of Linux, networking fundamentals, and distributed systems.
- Experience troubleshooting live production issues and owning fixes through post-incident reviews.
- Fluency with monitoring and observability, including metrics, logs, dashboards, and alerting.
- Ability to work in a fast-moving environment and participate in an on-call rotation.
Nice to have
- Experience with AI or GPU workloads or high-performance computing.
- Familiarity with InfiniBand or RDMA.
- Experience with Kubernetes and virtualized or bare-metal environments.
Culture & Benefits
- Ownership, accountability, and direct involvement with production infrastructure.
- Competitive base salary with equity, reviewed every 12 months.
- Early scope and a progression plan focused on developing selected skills.
- Flexible work approach with autonomy over the working day.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
12 дней назад
DevOps & SRE Engineer (Kubernetes)
100 000 - 150 000$
9 дней назад
Senior Site Reliability Engineer (AI)
191 000 - 226 000$
14 дней назад
Site Reliability Engineer (AWS)
180 000 - 220 000$
14 дней назад
Sr. Staff Lead Site Reliability Engineer (AWS)
220 000 - 330 000$
11 дней назад
Site Reliability Engineer II (AWS)
100 000 - 110 000$
13 дней назад
Software Engineer (Cloud Infrastructure/SRE)
147 900 - 220 000$