Назад
Company hidden
2 дня назад

Site Reliability Engineer (AI)

Формат работы
remote
Тип работы
fulltime
Грейд
middle/senior
Английский
b2
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Site Reliability Engineer (AI): Ensuring the reliability, performance, and scalability of high-performance AI inference infrastructure with an accent on observability, incident management, and distributed systems. Focus on automating operational toil, optimizing GPU-backed workloads, and building resilient production-grade systems.

Company

hirify.global builds high-performance infrastructure and serverless platforms to power fast, scalable AI inference for image, video, and emerging modalities.

What you will do

  • Own and improve the reliability, availability, and performance of critical production services.
  • Define and evolve reliability practices, including SLIs, SLOs, and production-readiness standards.
  • Investigate complex production issues across distributed systems, APIs, networking, and GPU-backed workloads.
  • Lead incident reviews and RCAs to turn failure modes into lasting engineering improvements.
  • Reduce operational toil through automation, self-healing systems, and deployment safety.
  • Collaborate with engineering teams on capacity planning, scaling, and architectural improvements.

Requirements

  • Strong experience operating and troubleshooting production systems at scale in an SRE or Platform Engineering role.
  • Deep understanding of distributed systems and debugging across applications, databases, queues, and infrastructure.
  • Experience designing and operating observability systems using metrics, logs, and distributed tracing.
  • Proficiency with Kubernetes, containers, IaC, and automated deployment practices.
  • Ability to write software and automation using Python, Go, or PHP.
  • Comfortable participating in an engineering on-call rotation and taking ownership of production problems.

Nice to have

  • Experience with GPU environments or AI and ML workloads.
  • Experience with RabbitMQ or other distributed messaging systems.
  • Experience operating MySQL, Redis, or ClickHouse.
  • Experience with global traffic management, load balancing, and CDN platforms.

Culture & Benefits

  • Remote-first collective with twice-yearly in-person retreats.
  • Flexible working hours outside of core collaboration blocks.
  • Generous paid time off including vacation, sick days, and public holidays.
  • Meaningful stock options to share in company growth.
  • Paid family leave for maternity, paternity, and caregiving.
  • Emphasis on recharging with downtime after intense release cycles.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →