1 день назад
Site Reliability Engineer/L3 Support (AWS/Kubernetes)
110 000 - 130 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer/L3 Support (AWS/Kubernetes): Operating and improving a cloud-native SaaS platform in a FedRAMP High environment with an accent on reliability, observability, incident response, and operational automation. Focus on troubleshooting distributed Kubernetes applications, reducing MTTD and MTTR, strengthening resilience, and implementing secure infrastructure improvements.
Location: Remote from New York, US; U.S. Citizenship and eligibility to work on FedRAMP High systems are required.
Salary: $110,000–$130,000 annual base salary in New York, plus potential discretionary bonus and/or equity awards.
Company
SS&C is a financial services and healthcare technology company operating a cloud-native SaaS platform for financial services and healthcare organizations.
What you will do
- Own the operational health, reliability, availability, performance, and security of the FedRAMP High cloud platform.
- Monitor production services using telemetry, logs, metrics, distributed tracing, dashboards, and alerts.
- Investigate and resolve complex application and infrastructure incidents as the L3 escalation point.
- Lead incident response, root cause analysis, post-incident reviews, and corrective actions.
- Develop runbooks, standard operating procedures, observability improvements, and operational automation.
- Support deployments, infrastructure changes, disaster recovery exercises, resilience testing, and compliance activities.
Requirements
- U.S. Citizenship is required.
- 3–6 years of experience in SRE, production engineering, DevOps, platform engineering, or senior production support.
- Experience supporting mission-critical cloud production systems and troubleshooting distributed applications in Kubernetes.
- Strong knowledge of Linux, networking, infrastructure as code, configuration management, and scripting or programming.
- Experience with public cloud platforms, preferably AWS, and monitoring or observability tools such as Prometheus, Grafana, CloudWatch, Datadog, Splunk, or OpenTelemetry.
- Strong incident management, root cause analysis, analytical, troubleshooting, and communication skills.
Nice to have
- Experience with FedRAMP High, DoD IL5/IL6, or similar regulated environments.
- AWS services including EKS, RDS, IAM, CloudWatch, Route 53, VPC networking, and AWS Backup.
- CI/CD pipelines, deployment automation, Istio, PagerDuty, or Jira Service Management.
- AWS Associate or Professional certification.
Culture & Benefits
- Hybrid work model and flexible personal or vacation time off.
- Medical, dental, vision, employee assistance, parental leave, paid holidays, and sick leave.
- 401(k) matching and professional development reimbursement.
- Hands-on training through SS&C University and team-customized learning.
- Collaboration with software engineers, platform engineers, and security specialists.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Senior Site Reliability Engineer (AWS)
160 000 - 220 000$
1 день назад
Site Reliability Engineer (AWS)
2 дня назад
Principal Site Reliability Engineer (AWS/Kubernetes)
163 620 - 212 710$
4 дня назад
Senior Site Reliability Engineer - IT
103 200 - 273 700$
3 дня назад
Senior Site Reliability Engineer (AWS)
2 дня назад
Senior Site Reliability Engineer (AWS/Kubernetes)
160 000 - 180 000$