11 дней назад
Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI/Cloud): Protecting the availability and performance of customer-facing applications through incident response, observability, deployments, and operational readiness with an accent on cloud infrastructure, Kubernetes, and responsible AI use. Focus on troubleshooting distributed environments, automating operational work, identifying supportability gaps, and improving reliability across production systems.
Location: Remote, Romania
Company
provides an AI-first cloud compliance platform that connects tax and technology for businesses.
What you will do
- Monitor customer-facing and revenue-critical applications and respond to production events.
- Coordinate incidents, escalations, communications, and resolution activities.
- Use observability and data analysis to identify operational risks before they affect customers.
- Deploy applications and review application readiness, supportability, observability, and release maturity.
- Troubleshoot application, infrastructure, platform, cloud, container, database, and networking issues across GCP, OCI, AWS, and Azure.
- Build automation and operational tools, recommend reliability improvements, and apply AI responsibly to troubleshooting and operational insight.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field.
- 3+ years of experience supporting large-scale cloud solutions in enterprise production environments.
- 2+ years of experience with production deployments, monitoring, incident response, and operational communications.
- Experience with Linux and Windows, observability, reliability practices, containerized applications or Kubernetes, SQL, and database services.
- Scripting experience with PowerShell, Python, Bash, or Perl, plus automation and performance diagnostics experience.
- Applied experience using AI tools such as Glean, ChatGPT, Claude, or similar platforms to improve analysis, quality, or automation.
Culture & Benefits
- 24x7x365 engineering operations support for customer-facing applications.
- Paid time off and paid parental leave.
- Eligibility for bonuses and a broader compensation package.
- Location-dependent private medical, life, and disability insurance.
- Inclusive culture supported by employee resource groups and senior leadership sponsorship.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
13 дней назад
Site Reliability Engineer (AI)
200 000 - 400 000$
12 дней назад
Staff Observability Engineer (AI)
290 000 - 375 000PLN
Valletta.Software | AI-Care
3 часа назад
Senior DevOps / SRE Support Engineer (LATAM)
5 000 - 5 500$
13 дней назад
Senior Site Reliability Developer (SaaS)
123 250 - 166 750CAD
5 часов назад
DevOps/SRE Engineer
Vallettasoftware Software Development
13 дней назад
Senior DevOps / SRE Support Engineer (Kubernetes)
5 000 - 5 500$