Назад
Company hidden
3 дня назад

Principal Site Reliability Engineer (Kubernetes)

190 000 - 220 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Site Reliability Engineer (Kubernetes): Building and maintaining reliable, scalable, and high-performance production infrastructure with an accent on Kubernetes operations, automation, observability, and incident response. Focus on defining SLOs and SLIs, optimizing capacity and performance, and solving complex reliability challenges through resilient services and operational best practices.

Location: Hybrid in Norwalk, Connecticut, USA

Salary: $190,000–$220,000 annually for positions in Connecticut and New York.

Company

hirify.global provides financial data, analytics, and software solutions to investment professionals and financial institutions worldwide.

What you will do

  • Monitor, maintain, and improve the reliability and availability of production systems and services.
  • Respond to incidents, participate in on-call support, and conduct blameless post-mortems.
  • Define and track Service Level Objectives and Service Level Indicators.
  • Collaborate with development and operations teams to build reliability into services.
  • Design automation that reduces operational toil and improves efficiency.
  • Contribute to capacity planning, performance optimization, system documentation, and runbooks.

Requirements

  • 8+ years of experience ensuring system and service reliability, scalability, and performance.
  • Hands-on Kubernetes experience required, including deployment, management, troubleshooting, cluster administration, networking, storage, and security.
  • Strong knowledge of Kubernetes concepts including Pods, Deployments, Services, ConfigMaps, and Ingress, plus experience with Helm.
  • Experience with cloud platforms, CI/CD tooling, monitoring and observability, infrastructure as code, configuration management, and programming or scripting.
  • Bachelor’s degree in computer science or a relevant field.
  • Strong analytical, communication, troubleshooting, automation, and incident-response skills.

Nice to have

  • Open-source contribution experience.
  • Familiarity with SRE principles from the Google SRE handbook.
  • Previous DevOps or Platform Engineering experience.

Culture & Benefits

  • Hybrid work environment.
  • Blameless culture focused on continuous learning and improvement.
  • Collaboration across technical and non-technical teams.
  • Employment with a company serving more than 200,000 investment professionals worldwide.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →