4 дня назад
Sr. Site Reliability Engineer
160 000 - 180 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Sr. Site Reliability Engineer (AWS/Kubernetes/Observability): Owning production reliability for CentralReach’s public and private cloud platforms with an accent on SLOs, error budgets, monitoring, incident response, and capacity planning. Focus on building multi-environment observability, automating toil reduction, improving cloud-native system performance, and implementing Datadog, Prometheus, and Grafana.
Location: Hybrid role associated with the Holmdel, New Jersey office
Base salary: $160,000–$180,000 USD per year
Company
provides software for autism and intellectual and developmental disability care, Applied Behavior Analysis, multidisciplinary therapy, and special education.
What you will do
- Own production reliability across availability, latency, performance, capacity, monitoring, emergency response, and uptime.
- Define and improve SLOs, SLIs, error budgets, dashboards, and observability practices.
- Analyze operational issues, lead incident response and root cause analysis, restore services, and maintain runbooks and standard operating procedures.
- Build automated observability and capacity-forecasting capabilities across multiple environments.
- Reduce operational toil through automation and continuous improvement.
- Collaborate with software engineering on releases, roadmap planning, operational readiness, and reliability practices.
Requirements
- Experience with monitoring, APM, and observability tools including Splunk, Prometheus, Datadog, and OpenTelemetry.
- Experience implementing logging, metrics, and tracing strategies.
- Strong understanding of CI/CD tools such as Jenkins, GitHub Actions, GitLab, Argo, and Kargo.
- Strong knowledge of AWS or other major cloud providers, cloud-native infrastructure, Kubernetes, and Helm.
- Programming experience in Java, Python, or Go, with familiarity with .NET application development.
- Strong understanding of Linux, Windows, software development, systems, networking, and cloud concepts.
Nice to have
- Experience using AI to improve productivity and amplify technical skills.
Culture & Benefits
- Hybrid work with collaborative offices in Holmdel, New Jersey; Fort Lauderdale, Florida; and Verona, Italy.
- Health benefits, generous paid time off, 401(k) matching, and paid parental leave.
- Career development support and wellness programs.
- Opportunities to participate in community engagement through CR Cares.
- Work environment focused on impact, inclusion, flexibility, innovation, and scale.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
8 дней назад
Staff Site Reliability Engineer (AI/ML)
112 500 - 187 500$
8 дней назад
Site Reliability Engineer (AWS)
120 000 - 185 000$
5 дней назад
Senior DevOps / Site Reliability Engineer (SRE) (Cybersecurity)
165 000 - 215 000$
6 дней назад
Site Reliability Engineer - Vice President (Kubernetes)
130 000 - 160 000$
10 дней назад
Senior Site Reliability Engineer (Temporal)
180 000 - 200 000$
5 дней назад
Site Reliability Engineer (Kubernetes)
180 000 - 220 000$