9 часов назад
Tech Lead DevOps (AI/MLOps)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Tech Lead DevOps (AI/MLOps): Building and operating scalable cloud, CI/CD, and AI/ML infrastructure with an accent on automation, reliability, security, and MLOps. Focus on designing Kubernetes and multi-cloud platforms, supporting GPU workloads, implementing observability and AIOps, and leading platform engineering practices.
Location: Hyderabad, India
Company
operates a global online vehicle auction platform connecting vehicle sellers with buyers across more than 190 countries.
What you will do
- Lead and mentor the DevOps and Platform Engineering team and define cloud, DevOps, and AI infrastructure strategy.
- Design and manage scalable CI/CD pipelines, automated releases, Infrastructure as Code, and GitOps workflows.
- Build and support MLOps infrastructure for model training, deployment, monitoring, drift detection, and retraining.
- Architect and operate AWS, Azure, or GCP infrastructure with Kubernetes, Docker, and OpenShift.
- Implement monitoring, logging, tracing, alerting, SRE metrics, incident response, and AIOps capabilities.
- Drive DevSecOps, IAM, secrets management, vulnerability management, compliance, and infrastructure security.
Requirements
- 8+ years of experience in DevOps, Cloud Engineering, or Platform Engineering, including leadership experience.
- Hands-on experience with AWS, Azure, or GCP; CI/CD tools; Terraform or CloudFormation; Docker; Kubernetes; and configuration management tools.
- Strong scripting or programming skills in Python, Bash, Go, or a similar language.
- Experience with observability tools such as Prometheus, Grafana, ELK, Datadog, or Splunk.
- Experience with AI/ML infrastructure, MLOps tools, model deployment, lifecycle management, and data pipelines.
- Knowledge of cloud-native security, AI-assisted operations, distributed systems, and high-availability environments.
Nice to have
- Experience with GPU infrastructure and optimization of AI/ML workloads.
- Experience scaling high-traffic or AI-driven applications.
- Cloud, Kubernetes, DevOps, or AI/ML certifications.
- Experience with Agile/Scrum, SRE practices, and platform engineering concepts.
Culture & Benefits
- Emphasis on diversity, inclusion, collaboration, and continuous improvement.
- Cross-functional collaboration with Engineering, QA, Security, Product, Data Engineering, and AI/ML teams.
- Work in fast-paced, high-availability environments focused on reliability and operational efficiency.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
1 день назад
ML Ops Support
3 часа назад
Principal Engineer, Software Engineering - Development Operations (DevOps)
7 дней назад
Cloud Administrator
3 часа назад
Cloud Engineering Specialist (AWS)
10 часов назад
Staff Site Reliability Engineer (Linux/Network Troubleshooting/Scripting) (AI)
4 часа назад