18 дней назад
Senior Engineer - Site Reliability (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Engineer - Site Reliability (AI): Enhancing the reliability and performance of AI platforms in Abu Dhabi with an accent on Kubernetes operations, observability, automation, and production incident response. Focus on designing preventive solutions through root cause analysis, improving SLO/SLI adoption, deploying releases with minimal disruption, and maintaining GPU-enabled infrastructure.
Location: Abu Dhabi, United Arab Emirates; onsite role
Company
develops and commercializes industrial artificial intelligence products and applications for the oil and gas industry as part of the G42 and ADNOC ecosystem.
What you will do
- Maintain and evolve monitoring, alerting, observability, and incident response systems.
- Lead reliability projects, root cause analyses, and preventive improvements for complex production issues.
- Analyze service performance, identify bottlenecks, and implement measurable reliability improvements.
- Automate infrastructure and deployment workflows, including CI/CD pipelines and application releases.
- Drive SLO/SLI adoption and collaborate with engineering teams, project managers, and solution architects.
- Mentor junior engineers, evaluate new technologies, and maintain operational knowledge bases.
Requirements
- Bachelor’s degree in Business Analytics, Data Science, Computer Science, Engineering, or a related field; a master’s degree is preferred.
- At least 5 years of experience in SRE, DevOps, systems administration, or platform engineering, including Kubernetes cluster management.
- Strong Linux/Unix administration skills and hands-on experience deploying and operating Kubernetes or OpenShift clusters.
- Experience with monitoring and observability platforms such as Prometheus, Grafana, ELK, or Sentry, plus CI/CD pipelines and automation.
- Proficiency in Python or Bash, familiarity with databases, and experience with backups, restoration, network security, and cloud platforms.
- Ability to troubleshoot high-pressure production incidents, perform root cause analysis, restore services, and maintain GPU-compatible container images.
Nice to have
- Experience with infrastructure-as-code tools such as Ansible or Terraform.
- Knowledge of Redis, Memcached, RabbitMQ, Kafka, Postgres, Elasticsearch, or ClickHouse.
- Working knowledge of OAuth 2.0, OpenID Connect, SAML 2.0, Kerberos, or LDAP.
Culture & Benefits
- Fast-paced environment focused on initiative and transformational industrial technology projects.
- Collaboration with professionals from around the world.
- Healthcare, education support for dependents, leave benefits, and other rewards.
- Career support and recognition for individual contributions.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →