11 дней назад
Senior Site Reliability Engineer (Hosted Infra)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Hosted Infra) (Cloud Infrastructure): Engineering and scaling multi-cloud infrastructure for Elastic Cloud across four cloud providers, 70+ regions, and tens of thousands of hosts with an accent on automation, infrastructure as code, host lifecycle management, and observability. Focus on building internal software, preventing incidents through alerting and monitoring, improving reliability, and participating in on-call response and postmortems.
Location: Australia
Company
develops the Search AI Platform, providing cloud-based search, security, and observability solutions for organizations.
What you will do
- Engineer internal software, tools, and services that automate large-scale infrastructure operations.
- Optimize the reliability and lifecycle of hosts across multiple cloud providers.
- Build alerting and monitoring systems that improve incident prevention and observability.
- Scale global infrastructure and evolve infrastructure management processes.
- Participate in code reviews, technical planning, documentation, mentoring, and knowledge sharing.
- Join the SRE on-call rotation, respond to incidents, improve runbooks, and contribute to postmortems and reliability improvements.
Requirements
- Production experience operating large-scale cloud compute environments with hundreds of hosts or more through automated workflows.
- Deep Linux systems experience, including operating-system-level debugging.
- Experience building software with Golang and reviewing maintainable code.
- Production experience with containerized workloads.
- Systems-thinking and customer-first approach to operational problems, with a focus on root-cause analysis.
- Ability to work across time zones, communicate clearly, and create documentation such as designs, runbooks, architecture decisions, and postmortems.
Nice to have
- Experience with Terraform or OpenTofu, Puppet or OpenVox, Ansible, Argo CD, Argo Workflows, CUE, Docker, Kubernetes, Ubuntu, or Ubuntu Live Patch.
- On-call incident response experience and familiarity with Stack, Graphite, Prometheus, or Influx.
- Hands-on experience engineering solutions with the Stack.
- Experience integrating AI tools into operational workflows to reduce toil without adding unnecessary complexity.
Culture & Benefits
- Distributed working environment with flexible locations and schedules for many roles.
- Health coverage for employees and families in many locations.
- Competitive pay based on the work performed rather than previous salary.
- Generous vacation allowance and at least 16 weeks of parental leave.
- Up to $2,000 in matched financial donations and up to 40 hours of paid volunteer time annually.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →