Staff Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Staff Site Reliability Engineer (AI): Building and evolving the reliability foundation for an AWS-based platform that provisions, deploys, and operates AI agents supporting global payment infrastructure, with an accent on event-driven architecture, observability, and scalable infrastructure. Focus on defining SLOs and error budgets, replacing synchronous communication with durable asynchronous messaging, and solving complex distributed-systems and production-reliability challenges.
Location: Remote globally; the role is listed across Europe and can be performed from anywhere.
Company
provides an AI-native global commerce operating system connecting enterprise merchants, banks, and wallets to payment, fraud prevention, KYC/KYB, and stablecoin infrastructure.
What you will do
- Define the platform-wide reliability strategy, including SLOs, SLIs, error budgets, incident practices, and engineering standards.
- Drive infrastructure architecture decisions and major platform evolutions as transaction volume and AI workloads scale.
- Design and own durable event-driven messaging for inter-service communication, including migration from synchronous patterns.
- Manage AWS infrastructure and infrastructure-as-code provisioning while operating Kubernetes and Docker workloads in production.
- Build observability through monitoring, alerting, dashboards, and distributed tracing, and lead incident response, postmortems, and root-cause analysis.
- Run chaos engineering and resilience experiments while mentoring engineers and raising reliability standards across teams.
Requirements
- 7+ years of experience in site reliability, infrastructure, or closely related engineering roles.
- Hands-on ownership of event-driven systems and messaging platforms such as Kafka, NATS, or RabbitMQ, including delivery semantics, consumer groups, dead letters, and backpressure.
- Deep AWS and networking knowledge, including EC2, VPC, IAM, S3, and RDS, plus Terraform or Pulumi experience.
- Production experience with Kubernetes, Docker, distributed-systems debugging, and automation using Go, Python, or a similar language.
- Experience defining and operating SLOs, SLIs, error budgets, observability platforms, chaos experiments, and resilience testing.
- Advanced written and spoken English proficiency required.
Nice to have
- Production AI or MLOps infrastructure experience, including model serving, LLM inference, GPU/resource management, and agent observability.
- Experience with multi-tenant container platforms, data pipelines, orchestration, or data warehouses.
- Payments-industry experience and incident-management tooling such as PagerDuty, Opsgenie, or incident.io.
- ECS, s6-overlay, or AI agent framework experience.
- Spanish proficiency.
Culture & Benefits
- Fully remote work with the ability to work from anywhere.
- One-time home-office setup allowance and company-provided equipment.
- Stock options and a health plan available wherever the role is based.
- Flexible days off.
- Language, professional, and personal development courses.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →