Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Site Reliability Engineer (AI): Ensuring the reliability, performance, and scalability of high-performance AI inference infrastructure with an accent on observability, incident management, and distributed systems. Focus on automating operational toil, optimizing GPU-backed workloads, and building resilient production-grade systems.
Company
builds high-performance infrastructure and serverless platforms to power fast, scalable AI inference for image, video, and emerging modalities.
What you will do
- Own and improve the reliability, availability, and performance of critical production services.
- Define and evolve reliability practices, including SLIs, SLOs, and production-readiness standards.
- Investigate complex production issues across distributed systems, APIs, networking, and GPU-backed workloads.
- Lead incident reviews and RCAs to turn failure modes into lasting engineering improvements.
- Reduce operational toil through automation, self-healing systems, and deployment safety.
- Collaborate with engineering teams on capacity planning, scaling, and architectural improvements.
Requirements
- Strong experience operating and troubleshooting production systems at scale in an SRE or Platform Engineering role.
- Deep understanding of distributed systems and debugging across applications, databases, queues, and infrastructure.
- Experience designing and operating observability systems using metrics, logs, and distributed tracing.
- Proficiency with Kubernetes, containers, IaC, and automated deployment practices.
- Ability to write software and automation using Python, Go, or PHP.
- Comfortable participating in an engineering on-call rotation and taking ownership of production problems.
Nice to have
- Experience with GPU environments or AI and ML workloads.
- Experience with RabbitMQ or other distributed messaging systems.
- Experience operating MySQL, Redis, or ClickHouse.
- Experience with global traffic management, load balancing, and CDN platforms.
Culture & Benefits
- Remote-first collective with twice-yearly in-person retreats.
- Flexible working hours outside of core collaboration blocks.
- Generous paid time off including vacation, sick days, and public holidays.
- Meaningful stock options to share in company growth.
- Paid family leave for maternity, paternity, and caregiving.
- Emphasis on recharging with downtime after intense release cycles.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →