8 дней назад
Site Reliability Engineer (AWS/Terraform)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AWS/Terraform): Building automation, observability tooling, and reliability standards across on-premises and AWS cloud infrastructure with an accent on infrastructure as code, SLOs, incident response, and production readiness. Focus on designing scalable deployment and monitoring systems, leading operational readiness reviews and major incidents, and using AI tools to automate troubleshooting and preventive reliability work.
Location: Remote in Australia
Company
develops data quality, data enrichment, and location intelligence software with an AI-first approach.
What you will do
- Define and maintain reliability standards across on-premises private cloud and AWS environments, including SLOs, SLIs, error budgets, alerting, logging, tracing, and MTTR improvements.
- Build infrastructure as code, deployment automation, monitoring, and observability tooling with Terraform, Ansible, Datadog, Python, and Bash.
- Guide CI/CD and deployment practices while embedding scalability, backup, recovery, and failure-mode planning into service design.
- Lead Operational Readiness Reviews, validate production and disaster recovery readiness, and maintain runbooks and reliability backlogs.
- Lead major incident response, perform root cause analyses, and implement preventive automation.
- Support security, compliance, vulnerability remediation, data protection, cross-team reliability reviews, and rotating on-call coverage.
Requirements
- Bachelor's degree in Computer Science, Information Systems, Engineering, or equivalent practical experience.
- At least 3 years of systems or infrastructure engineering experience in an enterprise production environment.
- Strong Linux experience, infrastructure-as-code experience with Terraform and/or Ansible, and AWS experience with EC2, ECS, S3, VPC, and IAM.
- Proficiency in Python or Bash, plus experience building monitoring and alerting systems; Datadog is preferred.
- Understanding of TCP/IP networking, DNS, load balancing, distributed systems, CI/CD, structured root cause analysis, SLOs, and operational readiness reviews.
- Active daily use of company-provided AI tools such as GitHub Copilot or Claude for infrastructure code, troubleshooting, testing, incident analysis, and documentation is required.
Nice to have
- Experience with Docker, ECS, Kubernetes, GitOps, Git, or GitLab.
- Knowledge of enterprise virtualization, hybrid cloud environments, ITIL, change management, or enterprise security tools.
- AWS Solutions Architect, SysOps Administrator, or DevOps Engineer certification.
Culture & Benefits
- Remote work arrangement in Australia.
- Participation in a rotating on-call schedule for critical escalations and changes.
- Collaboration with engineering teams across multiple infrastructure environments.
- Approximately 0% travel required.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Site Reliability Engineer (Kubernetes)
180 000 - 220 000$
11 часов назад
DevOps/SRE Engineer
10 дней назад
Site Reliability Engineer, Tech Lead (AI)
13 дней назад
Site Reliability Engineer (AWS/Terraform)
56 000 - 66 000€
Wheely
3 дня назад
Site Reliability Engineer (AWS)
5 000€
8 дней назад