4 дня назад
AI Cloud Senior DevOps Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Cloud Senior DevOps Engineer (Kubernetes/Docker/MLOps): Building and operating cloud-native infrastructure, CI/CD pipelines, and GPU clusters for AI products and model inference with an accent on high availability, infrastructure as code, and observability. Focus on automating multi-cloud deployments, designing disaster recovery and self-healing systems, and resolving complex production incidents.
Location: Remote within San Jose, CA or Austin, TX
Company
provides Bitcoin mining infrastructure, data center operations, and AI cloud computing capabilities.
What you will do
- Design and maintain end-to-end CI/CD and MLOps pipelines for software applications and machine learning models.
- Build, optimize, and scale Kubernetes- and Docker-based cloud infrastructure, including GPU clusters for AI workloads and model inference.
- Design high-availability production environments with disaster recovery, self-healing, capacity planning, and performance tuning.
- Automate reproducible infrastructure provisioning across cloud environments using Terraform, Ansible, and Helm.
- Develop monitoring, logging, and alerting systems with tools such as Prometheus, Grafana, and ELK/EFK.
- Lead incident resolution, root cause analysis, security governance, release workflows, and collaboration with R&D, Data Science, Security, and Business teams.
Requirements
- 5+ years of hands-on experience in DevOps, SRE, or cloud infrastructure roles.
- Bachelor's degree or higher in Computer Science, Engineering, or a related technical field.
- Expert knowledge of Linux and networking principles, including TCP/IP, DNS, HTTP, load balancing, and VPCs.
- Deep production-level experience with Docker and Kubernetes.
- Experience designing and managing AWS, GCP, Azure, Alibaba Cloud, or comparable public and hybrid cloud infrastructure.
- Strong coding or scripting skills in Go, Python, Shell, or another major language, plus knowledge of CI/CD, IaC, observability, and SRE practices.
Nice to have
- MLOps, model serving, model inference, or GPU cluster experience, including vLLM, TGI, or Triton Inference Server.
- Experience with large-scale distributed or high-concurrency systems and Internal Developer Platforms.
- Knowledge of Zero Trust, DevSecOps, SOC2, or ISO27001.
- Technical leadership, mentoring, or DevOps team management experience.
Culture & Benefits
- Full-time role within the AI Cloud team.
- Cross-functional collaboration across engineering, data science, security, and business functions.
- Work focused on automation, developer productivity, platform engineering, system reliability, and security.
- Equal employment opportunity commitment in accordance with applicable laws.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →