Назад
Company hidden
5 дней назад

Senior Site Reliability Engineer (AI Platform)

69 768 - 96 900
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
c1
Страна
France/UK/US +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (AI Platform) (Kubernetes/Cloud Infrastructure): Building and operating highly available production infrastructure for AI workloads with an accent on Kubernetes platforms, distributed systems, observability, and cloud efficiency. Focus on leading complex production investigations, improving reliability through SLOs and guardrails, and designing scalable automation, CI/CD, and FinOps solutions.

Location: Paris, France. Positions listed within a specific city are only available in that location; the role may use a hybrid-remote or in-office schedule depending on the position.

Annual base salary: €69,768–€96,900 EUR.

Company

hirify.global provides search and product discovery technology, powering AI-enabled customer experiences and high-volume search applications.

What you will do

  • Own and evolve production infrastructure supporting AI-related workloads and services at scale.
  • Design and operate highly available Kubernetes-based platforms.
  • Drive reliability through SLOs, observability, capacity planning, and production guardrails.
  • Lead complex production investigations and implement durable architectural improvements.
  • Improve networking, databases, service communication, compute, CI/CD, progressive delivery, and developer experience.
  • Drive cloud efficiency and FinOps initiatives while improving on-call and incident response.

Requirements

  • Strong hands-on production experience with GCP, AWS, or Azure.
  • Strong experience designing and operating Kubernetes and cloud-native production systems at scale.
  • Strong understanding of distributed systems, networking, and reliability engineering.
  • Experience operating business-critical systems with demanding availability, scalability, and operational requirements.
  • Ability to independently own ambiguous, cross-team technical problems and deliver measurable outcomes.
  • Excellent written and spoken English required.

Nice to have

  • Go or Python engineering experience.
  • Experience with infrastructure supporting AI/ML workloads, model serving, GPUs, or other compute-intensive systems.
  • Experience with coding agents, agentic development workflows, AI-assisted debugging, and automation.

Culture & Benefits

  • Flexible workplace model with autonomy over where and when to work, depending on role eligibility.
  • Remote and hybrid-remote options are available for many team members.
  • High-trust environment focused on impact, contribution, and output.
  • Values include grit, trust, candor, care, and humility.
  • Inclusive and collaborative workplace with a global presence.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →