1 день назад
Site Reliability Engineer
90 000 - 110 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AWS/Cloud Infrastructure): Automating infrastructure, container platforms, observability pipelines, and production incident response for high-performance financial technology products with an accent on reliability engineering, cloud administration, and system operability. Focus on defining SLOs and SLIs, building Prometheus and Grafana monitoring, leading PagerDuty escalations, and improving infrastructure through blameless post-mortems.
Location: Hybrid role based in Boston or Waltham, Massachusetts, with an expectation of working from a Boston-area office six days per month.
Salary: USD 90,000–110,000 base salary per year
Company
SS&C provides AI-powered technology and services for financial services and healthcare organizations, including products used by investment firms to manage large-scale assets.
What you will do
- Automate infrastructure configuration with AWS CloudFormation, Chef, Packer, and Terraform.
- Automate operational tasks and ChatOps integrations using Lambda, Python, and Node.
- Maintain container build and deployment systems with Docker and Jenkins.
- Configure and maintain secrets management, service discovery, and container orchestration systems.
- Build observability pipelines and dashboards with Prometheus, Grafana, OpenTelemetry, and Splunk.
- Lead incident response, on-call escalations, SLO and SLI tracking, code reviews, production troubleshooting, and blameless post-mortems.
Requirements
- Bachelor’s or master’s degree in Computer Science or a similar discipline.
- 5+ years of experience working in a Linux environment.
- 5+ years of cloud computing administration and engineering experience with AWS, Azure, or Google Cloud.
- Extensive experience troubleshooting and debugging infrastructure issues.
- Ability to perform under pressure in a fast-paced environment and continuously improve owned systems.
- Availability to participate in a weekly on-call rotation.
Nice to have
- Experience with Python, Ruby, Node, C#, SQL, or Go.
- 3+ years of experience with containerized software deployment and orchestration tools such as Kubernetes or Rancher.
- 3+ years of experience with automation tools such as Chef, Puppet, Ansible, SaltStack, or Terraform.
- Experience using AI-assisted coding, AIOps, observability platforms, or LLM and AI service integrations.
Culture & Benefits
- Hybrid work model with business-casual dress.
- 401(k) matching, professional development reimbursement, and SS&C University training.
- Flexible personal and vacation time off, sick leave, and paid holidays.
- Medical, dental, vision, employee assistance, and parental leave benefits.
- Fitness club, travel, and other employee discounts.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Site Reliability Engineer (AWS)
7 дней назад
Sr. Site Reliability Engineer
160 000 - 180 000$
1 день назад
Cloud Site Reliability Engineer (AWS)
120 000 - 130 000$
5 дней назад
Principal Site Reliability Engineer (AI)
165 000 - 185 000$
3 дня назад
Site Reliability Engineering (SRE), The Core Engineering, Analyst, Dallas
4 дня назад
Senior Site Reliability Engineer (AI)
185 500 - 232 000$