Назад
1 день назад

Senior Cloud Native Platform Engineer (AI Cloud)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Cloud Native Platform Engineer (AI Cloud): Building and improving shared Kubernetes-based platform capabilities for AI applications and services with an accent on reliability, observability, automation, and production operations. Focus on designing resilient deployment workflows, debugging complex issues across Linux, networking, storage, and cloud infrastructure, and reducing operational toil through engineering and automation.

Location: UK

Company

Nscale provides high-performance GPU cloud infrastructure for AI start-ups and enterprise customers.

What you will do

  • Build and improve shared cloud-native platform capabilities for internal engineering teams running AI applications and services.
  • Own platform areas including Kubernetes operations, workload runtime configuration, deployment workflows, observability, and environment automation.
  • Improve platform reliability, scalability, supportability, rollout safety, capacity awareness, and recovery procedures.
  • Develop automation, tooling, CI/CD workflows, integrations, and operational configurations that reduce manual effort.
  • Improve incident prevention, detection, response, recovery, runbooks, and operational practices.
  • Collaborate with software engineering, infrastructure, and SRE teams while contributing to design reviews, code reviews, standards, and mentoring.

Requirements

  • Hands-on experience operating and improving Kubernetes platforms in production.
  • Experience with infrastructure automation, CI/CD, configuration management, or GitOps workflows.
  • Strong understanding of reliability engineering, observability, incident response, failure analysis, and operational readiness.
  • Production-quality automation, tooling, or backend development experience with Go, Python, Bash, or similar languages.
  • Strong Linux and networking fundamentals, including processes, filesystems, cgroups, TCP/IP, DNS, routing, load balancing, and container networking.
  • Experience debugging complex production issues across multiple system layers and mentoring less experienced engineers.

Culture & Benefits

  • Opportunity to shape operating standards for a next-generation AI cloud platform.
  • Work on complex infrastructure challenges with substantial ownership and impact.
  • Contribute to scaling high-performance and sustainable data centre operations.
  • Culture focused on innovation, ownership, accountability, openness, and transparency.
  • Inclusive environment welcoming people from diverse backgrounds and offering reasonable accommodations.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →