Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Operational Engineer (AI): Building and improving operational tooling, integrations, dashboards, and automations for GPU cloud engineering services with an accent on service reliability, incident response, and operational readiness. Focus on coordinating incidents, identifying gaps in monitoring and ownership, automating routine changes, and improving service-health, SLA, cost, and operational-risk reporting.
Location: US
Company
Nscale operates a high-performance, cost-efficient GPU cloud infrastructure for AI-native startups and global enterprises.
What you will do
- Onboard services to operational tooling, including alerting, on-call processes, dashboards, runbooks, Jira, and service-catalog records.
- Build and maintain integrations, dashboards, reports, and automations across Grafana, PagerDuty, Jira, Backstage, and cloud billing platforms.
- Improve operational readiness by identifying gaps in ownership, alerting, documentation, recovery procedures, and service dependencies.
- Coordinate incident response as incident commander, including communications, timelines, post-incident reviews, and follow-up actions.
- Make routine changes safer and more repeatable through clear processes and automation.
- Contribute to service-health, SLA, cost, patching, operational-risk, and service-ownership reporting.
Requirements
- 2–5 years of experience in operations engineering, SRE, cloud infrastructure, platform engineering, or a similar role.
- Experience operating or supporting production services.
- Experience with monitoring, alerting, on-call processes, runbooks, and engineering work management.
- Hands-on experience with some of Grafana, PagerDuty, Jira, Backstage, public cloud, dashboards, integrations, or workflow automation.
- Ability to use operational data to identify gaps and follow actions through to closure.
Culture & Benefits
- Culture centered on ownership, accountability, and speed.
- Close involvement with the infrastructure powering AI services.
- Focus on reducing manual toil and improving operational outcomes.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Sr. Site Reliability Engineer (AI)
5 дней назад
Head of Site Reliability Engineering (AI)
195 000 - 285 000$
Anthropic
5 дней назад
Incident Response Manager (AI)
290 000 - 365 000$
Replit
6 дней назад
Engineering Manager (SRE)
250 000 - 325 000$
9 дней назад
Associate, Operations Engineer (Fintech)
110 000 - 140 000$
4 дня назад