Назад
обновлено 2 часа назад

Senior Operational Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Operational Engineer (AI): Designing and delivering shared operational capabilities for a GPU cloud serving AI startups and enterprises with an accent on service readiness, observability, incident response, and cost-aware infrastructure. Focus on building cross-service integrations and auditable automation, leading incident analysis, and improving reliability through continuity and recovery testing.

Location: US

Company

GPU cloud infrastructure for AI-native startups and global enterprises, spanning bare metal infrastructure and platform services.

What you will do

  • Design and deliver shared operational capabilities, including service-readiness checks, canaries, runbooks, health reporting, alerting workflows, and automation.
  • Establish standards for service ownership, on-call readiness, dashboards, recovery procedures, and operational evidence.
  • Partner with service teams to resolve recurring operational issues and cross-team blockers.
  • Build integrations and workflows across Grafana, PagerDuty, Jira, Backstage, public-cloud platforms, and reporting systems.
  • Lead incident response and technical analysis, turning incidents, change failures, and near misses into lasting engineering improvements.
  • Automate reporting and support continuity testing, failure experiments, mentoring, and operational design reviews.

Requirements

  • 6–10 years of experience in operations engineering, SRE, cloud infrastructure, platform engineering, or a similar production-focused engineering role.
  • Experience designing and operating monitoring, alerting, on-call, runbook, service-readiness, and operational-reporting capabilities.
  • Hands-on experience with Grafana, PagerDuty, Jira, Backstage, public cloud platforms, integrations, and automation.
  • Experience using incident, change, service-health, continuity, patching, or cost data to drive operational improvements.
  • Ability to lead cross-service initiatives, establish practical standards, and support teams through incidents.

Culture & Benefits

  • Ownership, accountability, and fast execution are central to the working culture.
  • Close involvement with the infrastructure that powers AI workloads.
  • Opportunity to influence engineering standards across multiple services.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →