3 дня назад
Principal Site Reliability Engineer (Kubernetes)
190 000 - 220 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Site Reliability Engineer (Kubernetes): Building and maintaining reliable, scalable, and high-performance production infrastructure with an accent on Kubernetes operations, automation, observability, and incident response. Focus on defining SLOs and SLIs, optimizing capacity and performance, and solving complex reliability challenges through resilient services and operational best practices.
Location: Hybrid in Norwalk, Connecticut, USA
Salary: $190,000–$220,000 annually for positions in Connecticut and New York.
Company
provides financial data, analytics, and software solutions to investment professionals and financial institutions worldwide.
What you will do
- Monitor, maintain, and improve the reliability and availability of production systems and services.
- Respond to incidents, participate in on-call support, and conduct blameless post-mortems.
- Define and track Service Level Objectives and Service Level Indicators.
- Collaborate with development and operations teams to build reliability into services.
- Design automation that reduces operational toil and improves efficiency.
- Contribute to capacity planning, performance optimization, system documentation, and runbooks.
Requirements
- 8+ years of experience ensuring system and service reliability, scalability, and performance.
- Hands-on Kubernetes experience required, including deployment, management, troubleshooting, cluster administration, networking, storage, and security.
- Strong knowledge of Kubernetes concepts including Pods, Deployments, Services, ConfigMaps, and Ingress, plus experience with Helm.
- Experience with cloud platforms, CI/CD tooling, monitoring and observability, infrastructure as code, configuration management, and programming or scripting.
- Bachelor’s degree in computer science or a relevant field.
- Strong analytical, communication, troubleshooting, automation, and incident-response skills.
Nice to have
- Open-source contribution experience.
- Familiarity with SRE principles from the Google SRE handbook.
- Previous DevOps or Platform Engineering experience.
Culture & Benefits
- Hybrid work environment.
- Blameless culture focused on continuous learning and improvement.
- Collaboration across technical and non-technical teams.
- Employment with a company serving more than 200,000 investment professionals worldwide.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Replit
10 дней назад
Staff Site Reliability Engineer (Kubernetes/GCP)
250 000 - 325 000$
4 дня назад
Senior Site Reliability Engineer
140 000 - 180 000$
9 дней назад
Site Reliability Engineering Manager (AWS/Kubernetes)
205 000 - 255 000$
10 дней назад
Principal Site Reliability Engineer (AI)
165 000 - 185 000$
9 дней назад
Senior/Lead Site Reliability Engineer (Federal)
159 000 - 230 000$
Replit
10 дней назад
Site Reliability Engineer
210 000 - 275 000$