Назад
Company hidden
2 дня назад

Senior Site Reliability Engineer (Kubernetes)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/Australia/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (Kubernetes): Improving the reliability, resilience, security, and availability of production systems with an accent on Kubernetes, cloud infrastructure, automation, and observability. Focus on operating live infrastructure, responding to incidents, reducing toil, and developing resilient solutions across distributed teams.

Location: Remote, limited to Texas, the United States of America, Alberta, or British Columbia

Company

hirify.global provides Network as a Service that connects businesses to cloud providers, data centers, and each other.

What you will do

  • Improve production reliability, system resilience, security, and availability within an SRE team.
  • Operate live production infrastructure, handle alerts, participate in on-call rotations, and respond to incidents.
  • Develop automation, write code and effective runbooks, reduce operational toil, and prevent recurring problems.
  • Use observability systems for metrics, logs, and traces while maintaining a strong signal-to-noise ratio.
  • Collaborate with engineering teams and stakeholders on requirements, demonstrations, peer reviews, and technical solutions.
  • Conduct blameless post-incident reviews and promote DevOps, SRE, and industry best practices.

Requirements

  • 5+ years administering Linux systems and related infrastructure in production.
  • Strong Kubernetes and cloud infrastructure fundamentals; AWS experience is strongly preferred.
  • Experience with Bash and either Python or Go, infrastructure as code with Terraform, CI/CD, version control, and GitHub.
  • Experience with at least one of PostgreSQL, Cassandra, or ClickHouse and with production observability stacks.
  • Knowledge of SRE practices including SLIs, SLOs, SLAs, error budgets, blast radius, and blameless postmortems.
  • Self-directed, collaborative work style suited to an asynchronous, globally distributed team.

Nice to have

  • Bare-metal infrastructure experience.
  • AWS experience and Terraform expertise.

Culture & Benefits

  • Remote-first working environment with coworking options.
  • Four weeks of paid annual leave, parental leave, birthday leave, and purchased annual leave options.
  • Wellness allowance and employee wellbeing initiatives.
  • Study and training allowance plus five days of paid study leave.
  • Inclusive, collaborative environment with recognition programs and modern workspaces.

Hiring process

  • Candidates who meet the selection criteria are invited to an interview.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →