Назад
Company hidden
обновлено 3 дня назад

Lead Site Reliability Engineer

Формат работы
remote (только Romania)
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
Romania
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Lead Site Reliability Engineer (Kubernetes, AI): Building an automation-first reliability ecosystem for Avalara’s distributed SaaS platform across multi-cloud environments with an accent on observability, self-healing systems, and safe deployment practices. Focus on designing AI-driven operations, strengthening Kubernetes platform resilience, and leading incident response and lasting reliability improvements.

Location: Remote from Romania

Company

hirify.global develops a cloud compliance platform that connects tax and technology, processing customer API calls and tax returns at large scale.

What you will do

  • Own and evolve reliability strategy for distributed SaaS systems across multi-cloud platforms.
  • Design AI-driven operations for predictive monitoring, anomaly detection, automated root cause analysis, and incident resolution.
  • Build observability solutions with Prometheus, Grafana, and OpenTelemetry.
  • Create self-healing systems and automation frameworks that reduce manual operational work.
  • Improve feature-flagged deployments, progressive delivery, safe rollouts, CI/CD pipelines, and infrastructure-as-code environments.
  • Strengthen Kubernetes-based platform availability, scalability, and fault tolerance while mentoring engineers and leading post-incident improvements.

Requirements

  • 10+ years of experience in SaaS, distributed systems, or site reliability engineering.
  • Programming experience with Go, Java, or Python.
  • Deep hands-on experience with Prometheus, Grafana, and OpenTelemetry.
  • Experience with Kubernetes, containerization, and multi-cloud platforms such as AWS, GCP, Azure, or OCI.
  • Strong knowledge of Linux systems, networking, and cloud-native architectures.
  • Experience designing automation and applying AI or machine learning to operational workflows.

Culture & Benefits

  • AI-first workflows, decision-making, and products.
  • Paid time off and paid parental leave.
  • Eligible employees may receive bonuses.
  • Benefits may include private medical, life, and disability insurance, depending on location.
  • Inclusive culture supported by employee resource groups with senior leadership and executive sponsorship.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →