Назад
Company hidden
13 часов назад

Staff Site Reliability Engineer (AI)

Формат работы
remote (Global)
Тип работы
fulltime
Грейд
senior
Английский
c1
Страна
France/Poland/Ireland +8 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Staff Site Reliability Engineer (AI): Building and evolving the reliability foundation for an AWS-based platform that provisions, deploys, and operates AI agents supporting global payment infrastructure, with an accent on event-driven architecture, observability, and scalable infrastructure. Focus on defining SLOs and error budgets, replacing synchronous communication with durable asynchronous messaging, and solving complex distributed-systems and production-reliability challenges.

Location: Remote globally; the role is listed across Europe and can be performed from anywhere.

Company

hirify.global provides an AI-native global commerce operating system connecting enterprise merchants, banks, and wallets to payment, fraud prevention, KYC/KYB, and stablecoin infrastructure.

What you will do

  • Define the platform-wide reliability strategy, including SLOs, SLIs, error budgets, incident practices, and engineering standards.
  • Drive infrastructure architecture decisions and major platform evolutions as transaction volume and AI workloads scale.
  • Design and own durable event-driven messaging for inter-service communication, including migration from synchronous patterns.
  • Manage AWS infrastructure and infrastructure-as-code provisioning while operating Kubernetes and Docker workloads in production.
  • Build observability through monitoring, alerting, dashboards, and distributed tracing, and lead incident response, postmortems, and root-cause analysis.
  • Run chaos engineering and resilience experiments while mentoring engineers and raising reliability standards across teams.

Requirements

  • 7+ years of experience in site reliability, infrastructure, or closely related engineering roles.
  • Hands-on ownership of event-driven systems and messaging platforms such as Kafka, NATS, or RabbitMQ, including delivery semantics, consumer groups, dead letters, and backpressure.
  • Deep AWS and networking knowledge, including EC2, VPC, IAM, S3, and RDS, plus Terraform or Pulumi experience.
  • Production experience with Kubernetes, Docker, distributed-systems debugging, and automation using Go, Python, or a similar language.
  • Experience defining and operating SLOs, SLIs, error budgets, observability platforms, chaos experiments, and resilience testing.
  • Advanced written and spoken English proficiency required.

Nice to have

  • Production AI or MLOps infrastructure experience, including model serving, LLM inference, GPU/resource management, and agent observability.
  • Experience with multi-tenant container platforms, data pipelines, orchestration, or data warehouses.
  • Payments-industry experience and incident-management tooling such as PagerDuty, Opsgenie, or incident.io.
  • ECS, s6-overlay, or AI agent framework experience.
  • Spanish proficiency.

Culture & Benefits

  • Fully remote work with the ability to work from anywhere.
  • One-time home-office setup allowance and company-provided equipment.
  • Stock options and a health plan available wherever the role is based.
  • Flexible days off.
  • Language, professional, and personal development courses.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →