2 дня назад
Site Reliability Engineer II (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer II, APAC (AWS/AI): Operating and improving production services in Australia with an accent on service-level objectives, observability, troubleshooting, and automation. Focus on investigating incidents across the application, data, infrastructure, and network layers, building AI-assisted tooling, and improving reliability through infrastructure as code and CI/CD.
Location: Remote in Australia. Applicants must be based in Australia and have unrestricted work rights in Australia. Periodic attendance at a office or collaborative work location may be required for certain events. Employer sponsorship and employer-administered work authorization are not available.
Company
provides Apple device management, deployment, and security solutions for workplace, education, and healthcare customers.
What you will do
- Implement and maintain service-level objectives, error budgets, and supporting service indicators with engineering teams.
- Investigate production issues across application, data, infrastructure, and network layers using logs, metrics, traces, profilers, and query plans.
- Build automation, AI-assisted tooling, and process improvements to eliminate operational toil and improve reliability.
- Create technical documentation, runbooks, postmortems, and proofs of concept for technical and non-technical audiences.
- Contribute to AI-agent guardrails, integrations, repository context, reusable skills, and prompt patterns.
- Support customer escalations, participate in an on-call rotation, shape reliability designs, and mentor less-experienced engineers.
Requirements
- At least 4 years of experience in software engineering, SRE, IT, or production operations.
- Production troubleshooting experience across the stack and experience operating services on AWS, including EC2, S3, EKS, RDS/Aurora, or CloudFront.
- Experience with observability tools such as Grafana, Prometheus, or LogicMonitor.
- Experience writing automation in Python, Go, Java, or a similar general-purpose language.
- Experience with Agile development and clear technical documentation.
- Must be based in Australia with unrestricted Australian work rights; visa sponsorship is not available.
Nice to have
- Experience with infrastructure as code and CI/CD tools such as Terraform, GitHub Actions, or Jenkins.
- Experience contributing to shared AI assets, repository context files, reusable skills, or prompt patterns.
Culture & Benefits
- Flexible, trust-based work environment with a focus on work-life balance.
- Small, empowered teams with a culture of ownership, respect, humility, and continuous improvement.
- Clear career development path with supportive leadership and management.
- Opportunity to improve products used by more than 75,000 global customers.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Site Reliability Engineer/L3 Support (AWS/Kubernetes)
110 000 - 130 000$
7 дней назад
Senior Engineer, Infrastructure Platform (AI)
5 дней назад
Site Reliability Engineer – AI-first Platform
7 дней назад
Staff Site Reliability Engineer (AI/ML)
241 000 - 270 000$
2 дня назад
AWS Cloud Engineer - Platform Operations
7 дней назад