Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Operational Engineer (AI): Designing and delivering shared operational capabilities for a GPU cloud serving AI startups and enterprises with an accent on service readiness, observability, incident response, and cost-aware infrastructure. Focus on building cross-service integrations and auditable automation, leading incident analysis, and improving reliability through continuity and recovery testing.
Location: US
Company
GPU cloud infrastructure for AI-native startups and global enterprises, spanning bare metal infrastructure and platform services.
What you will do
- Design and deliver shared operational capabilities, including service-readiness checks, canaries, runbooks, health reporting, alerting workflows, and automation.
- Establish standards for service ownership, on-call readiness, dashboards, recovery procedures, and operational evidence.
- Partner with service teams to resolve recurring operational issues and cross-team blockers.
- Build integrations and workflows across Grafana, PagerDuty, Jira, Backstage, public-cloud platforms, and reporting systems.
- Lead incident response and technical analysis, turning incidents, change failures, and near misses into lasting engineering improvements.
- Automate reporting and support continuity testing, failure experiments, mentoring, and operational design reviews.
Requirements
- 6–10 years of experience in operations engineering, SRE, cloud infrastructure, platform engineering, or a similar production-focused engineering role.
- Experience designing and operating monitoring, alerting, on-call, runbook, service-readiness, and operational-reporting capabilities.
- Hands-on experience with Grafana, PagerDuty, Jira, Backstage, public cloud platforms, integrations, and automation.
- Experience using incident, change, service-health, continuity, patching, or cost data to drive operational improvements.
- Ability to lead cross-service initiatives, establish practical standards, and support teams through incidents.
Culture & Benefits
- Ownership, accountability, and fast execution are central to the working culture.
- Close involvement with the infrastructure that powers AI workloads.
- Opportunity to influence engineering standards across multiple services.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Sr. Site Reliability Engineer (AI)
8 дней назад
Senior SRE (Kubernetes)
150 000 - 170 000$
9 дней назад
Senior Manager, Site Reliability Engineering (AI Ops)
222 000 - 300 500$
11 дней назад
Site Reliability Engineer, Global Banking & Markets, Vice President
150 000 - 250 000$
9 дней назад
Site Reliability Engineering (SRE) Manager (Azure)
139 700 - 232 900$
6 дней назад
Staff Site Reliability Engineer (Cybersecurity)
199 750 - 270 000$