Назад
обновлено 2 месяца назад

Site Reliability Engineer (AWS)

Формат работы
onsite
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
China
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (AWS): Building and operating scalable, observable, and resilient cloud infrastructure on AWS with an accent on Infrastructure as Code, automation, and reliability engineering. Focus on designing multi-region systems, implementing monitoring and alerting, managing incidents and deployments, and optimizing infrastructure cost and performance.

Location: Chengdu, China

Requirements

  • Bachelor’s degree in Computer Science, Software Engineering, Information Technology, or a related field.
  • At least 3 years of experience in SRE, DevOps, cloud infrastructure, or system administration.
  • Hands-on experience with AWS services including EC2, Lambda, ECS, EKS, RDS, ElastiCache, S3, SQS, SES, Auto Scaling, Load Balancers, VPC, Route 53, and CloudFront or similar edge networking.
  • Strong knowledge of Terraform or CloudFormation, distributed systems, cloud networking, firewalls, DNS, HTTP/TLS, and Linux, Windows, and container-based environments.
  • Proficiency in Python, Node.js, Bash, Ruby, or a related scripting language.
  • Experience with monitoring, alerting, logging, incident management, postmortems, zero-downtime deployments, blue/green or canary releases, and AWS cost optimization.

What you will do

  • Design, develop, and maintain AWS infrastructure using Terraform or AWS CloudFormation.
  • Operate reliable and scalable cloud environments across compute, containers, databases, storage, messaging, networking, and load balancing.
  • Lead and participate in architecture reviews covering reliability, scalability, security, performance, and infrastructure cost-efficiency.
  • Build monitoring, alerting, and logging solutions with tools such as CloudWatch, Prometheus, Grafana, and ELK, including AIOps capabilities for anomaly detection and predictive alerting.
  • Handle incidents, conduct root cause analysis and postmortems, improve CI/CD and deployment automation, and support machine learning model deployment lifecycles.
  • Maintain SLOs, SLAs, and error budgets while participating in on-call rotations and troubleshooting assigned incidents and tickets.

Culture & Benefits

  • Work with a global team across five continents.
  • Inclusive and respectful workplace with equal employment opportunity principles.
  • Reasonable accommodations are provided where needed for disability or religious practices.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →