19 часов назад
Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI) (Python/Cloud/LLM Evaluation): Building resilient production infrastructure, automated test suites, and observability frameworks for secure enterprise AI applications with an accent on cloud reliability, banking-scale performance, and model evaluation. Focus on designing chaos-resilient environments, validating RAG and LLM outputs, automating CI/CD and IaC checks, and enforcing security and data privacy standards.
Location: Abu Dhabi, United Arab Emirates — on-site
Company
is an AI and data consultancy delivering bespoke intelligent systems for enterprise clients, with a focus on financial services and banking across the UAE and GCC.
What you will do
- Build load-testing, fault-injection, and self-healing automation for web and mobile backends and AI microservices.
- Develop Python-based, CI/CD-integrated regression suites for application logic, data pipelines, and Infrastructure as Code states.
- Define SLIs and SLOs and implement telemetry, dashboards, and proactive alerting.
- Integrate AI evaluation harnesses such as DeepEval and Ragas to benchmark model accuracy, hallucinations, and RAG retrieval performance.
- Automate security scans and privacy checks while monitoring cloud efficiency and spend.
- Collaborate with architects, data engineers, AI specialists, and senior enterprise stakeholders to communicate and deliver secure AI solutions.
Requirements
- High proficiency in Python and Bash, with experience in PyTest, Selenium, or Robot Framework.
- Hands-on experience with Azure or AWS, including cloud networking, auto-scaling, containers, and serverless reliability.
- Experience with GitHub Actions, Azure DevOps, Terraform, K6 or JMeter, and observability tools such as Azure Monitor, CloudWatch, Grafana, or Prometheus.
- Understanding of prompt engineering, RAG architectures, and model evaluation metrics including ROUGE, BLEU, and LLM benchmarks.
- Experience in financial services, banking, or highly regulated enterprise environments with strict uptime and data privacy requirements.
Nice to have
- Experience managing production Kubernetes clusters, service meshes, and container security scanners.
- Familiarity with UAE Central Bank guidelines, NESA security compliance, and data residency mandates.
- AWS SysOps Administrator, Azure Administrator Associate, CKA, or SRE certification.
Culture & Benefits
- Work on secure, high-impact AI systems for enterprise clients.
- Collaborate with architects, data engineers, and AI specialists in a culture of accountability and practical problem-solving.
- Visa sponsorship is available for the successful candidate.
- Personal health insurance, professional development, certification support, and role-related subscription reimbursement.
- Monthly employee incentives and career advancement opportunities.
Hiring process
- The application and interview process is designed to be accessible, predictable, and fair.
- Reasonable adjustments can be requested during the application and interview process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
12 часов назад
Senior DevOps Engineer (AI)
7 дней назад
Senior DevOps Engineer (Data Store)
10 часов назад
Senior Site Reliability Engineer (Cloud Infrastructure)
5 часов назад
Senior / Staff Platform Engineer (AI)
3 дня назад
Platform Reliability Engineer (Cloud)
NDA
1 день назад
Senior DevOps Engineer
8 000 - 12 000$