Назад
Company hidden
4 часа назад

Senior Site Reliability Engineer (Kubernetes)

128 500 - 190 000$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (Kubernetes): Designing, operating, and evolving highly available production platforms for a global SaaS environment with an accent on reliability, scalability, automation, and cloud-native infrastructure. Focus on building self-service platform capabilities, leading complex incident investigations, improving observability and deployment workflows, and applying AI-assisted engineering practices to reduce operational toil.

Location: McLean, Virginia, United States

Annual base salary: $128,500–$190,000

Company

hirify.global provides the hirify.global Experience Cloud, a SaaS platform for managing experiences, insights, and actions across customer, employee, patient, and resident journeys.

What you will do

  • Design, build, operate, and evolve highly available, scalable, and secure production platforms.
  • Partner with software engineering teams to improve application reliability, performance, scalability, and operational readiness.
  • Lead complex incident investigations, root cause analyses, and reliability improvement initiatives.
  • Build automation, self-service capabilities, infrastructure-as-code, deployment tooling, and platform solutions that reduce operational toil.
  • Support CI/CD and GitOps workflows and design observability strategies covering monitoring, logging, tracing, and alerting.
  • Drive SRE practices, platform standards, AI-assisted engineering workflows, and operational improvements across engineering teams; mentor junior engineers.

Requirements

  • 5+ years leading reliability, platform engineering, infrastructure, or cloud operations initiatives in production environments.
  • Experience operating large-scale production environments, Kubernetes, containerized workloads, and cloud platforms such as AWS, OCI, or GCP.
  • Strong Linux administration and troubleshooting skills, plus automation development with Python, Go, Bash, or similar languages.
  • Experience with Terraform or similar infrastructure-as-code technologies, CI/CD, GitOps, distributed systems, and production incident response.
  • Understanding of DNS, load balancing, TLS/SSL, routing, service networking, and operational reliability practices.
  • Ability to influence technical decisions across teams and participate in a rotating on-call schedule.

Nice to have

  • Experience with ArgoCD, multi-region or hybrid-cloud environments, and observability tools such as Prometheus, Grafana, Loki, or OpenTelemetry.
  • Experience with platform engineering, self-service infrastructure, high-scale SaaS, capacity planning, performance engineering, and resilience testing.
  • Knowledge of canary, blue/green, progressive delivery, feature flags, security, compliance, and regulatory requirements.
  • Experience using AI-assisted development, automation, or operational tooling to improve productivity and reliability.

Culture & Benefits

  • Work within a Site Reliability Engineering organization supporting a global SaaS platform.
  • Medical, dental, and vision insurance, 401(k), disability coverage, life and AD&D insurance.
  • Statutory leaves, paid parental leave, and paid holidays.
  • Equal opportunity workplace with accommodations available for applicants with disabilities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →