2 часа назад
Senior DevOps Engineer, AI Platform
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior DevOps Engineer, AI Platform (Azure/OCI/Kubernetes): Building and operating production infrastructure for AI platforms, agent runtimes, web applications, backend services, and APIs with an accent on Kubernetes, cloud networking, CI/CD, and observability. Focus on translating technical designs into secure, scalable infrastructure, supporting asynchronous AI workloads, and solving production reliability, troubleshooting, and cost optimization challenges.
Location: Remote in Canada
Company
helps businesses establish and protect their online presence through domain, brand, website, security, and digital infrastructure services.
What you will do
- Translate technical designs into reliable, scalable, secure, and observable production cloud infrastructure across Microsoft Azure and Oracle Cloud Infrastructure.
- Design, provision, operate, and troubleshoot Kubernetes environments, primarily Azure Kubernetes Service and Oracle Kubernetes Engine.
- Support AI workloads including LLM gateways, Python agent runtimes, RAG workers, MCP services, background workers, and asynchronous processing pipelines.
- Manage networking and platform dependencies including ingress and egress, load balancers, DNS, TLS, private connectivity, routing, firewalls, databases, caches, queues, and scheduled jobs.
- Build CI/CD pipelines with Jenkins and Bitbucket using Docker, Helm, Kubernetes, ArgoCD, and container registries.
- Own observability, production readiness, incident troubleshooting, root cause analysis, scalability, reliability, and infrastructure cost optimization.
Requirements
- 7+ years of experience in DevOps, SRE, platform engineering, cloud infrastructure, or a related role.
- Strong production Kubernetes experience, including networking, scheduling, storage, autoscaling, security, and troubleshooting.
- Strong Microsoft Azure experience with AKS, networking, identity, storage, and monitoring; OCI experience is preferred.
- Experience with Jenkins, Bitbucket, Docker, Terraform, Helm, Kubernetes, and infrastructure as code.
- Experience supporting production web applications and backend services, including REST APIs, microservices, background workers, databases, caching, and messaging systems.
- Strong Linux, systems, production troubleshooting, observability, and application infrastructure knowledge, including Python and backend frameworks such as FastAPI.
Nice to have
- Experience with AI or machine learning platforms, LLM gateways, agent runtimes, RAG pipelines, or MCP services.
- Experience with Cloudflare, Envoy, ArgoCD, GitOps, and OpenTelemetry.
- Experience building reusable infrastructure platforms for high-scale SaaS or customer-facing applications.
Culture & Benefits
- Hands-on ownership of infrastructure delivery from design through UAT and production.
- Close collaboration with AI engineers, application engineers, architects, and platform teams.
- Focus on reusable infrastructure patterns that help engineering teams launch services consistently.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →