7 дней назад
Cloud Senior DevOps Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Cloud Senior DevOps Engineer (AI): Building the CI/CD, infrastructure-as-code, MLOps, and Internal Developer Platform foundations for an AI-operated GPU cloud with an accent on Kubernetes, multi-cloud infrastructure, high availability, and observability. Focus on designing automated deployment paths for software and machine-learning models, provisioning GPU clusters, enforcing security and compliance, and leading incident remediation.
Location: Remote within San Jose, California, or Austin, Texas
Company
develops Bitcoin mining infrastructure, AI computational infrastructure, and cloud capabilities for high-demand artificial intelligence workloads.
What you will do
- Design, implement, and maintain end-to-end GitLab-based CI/CD and MLOps pipelines for software applications and machine-learning models.
- Build and scale Kubernetes- and Docker-based cloud-native infrastructure, including specialized GPU clusters for AI workloads and model inference.
- Own high-availability architecture, disaster recovery, self-healing, capacity planning, and production performance tuning.
- Automate reproducible infrastructure provisioning across cloud environments with Terraform, Ansible, and Helm.
- Build observability systems using monitoring, logging, alerting, and AI model telemetry.
- Develop the Internal Developer Platform, improve deployment workflows, enforce security and compliance standards, and lead major incident resolution.
Requirements
- Bachelor’s degree or higher in Computer Science, Engineering, or a related technical field, plus 5+ years of hands-on DevOps, SRE, or cloud infrastructure experience.
- Expert knowledge of Linux and networking fundamentals, including TCP/IP, DNS, HTTP, load balancing, and VPCs.
- Deep production-level expertise with Docker and Kubernetes.
- Experience designing and managing infrastructure on public or hybrid cloud platforms such as AWS, GCP, Azure, or Alibaba Cloud, including multi-cloud strategies.
- Strong programming or scripting skills in at least one major language, such as Go, Python, or Shell.
- Practical understanding of CI/CD, infrastructure as code, observability, and SRE principles, with strong problem-solving and cross-team communication skills.
Nice to have
- Experience with MLOps, model serving or inference frameworks such as vLLM, TGI, or Triton Inference Server, and GPU cluster management.
- Experience with large-scale distributed systems, Internal Developer Platforms, Zero Trust, DevSecOps, SOC 2, or ISO 27001.
- Technical leadership, mentoring, or DevOps team management experience.
- Experience applying LLM-driven automation to code, configuration, pipelines, or incident response.
Culture & Benefits
- Work on infrastructure connecting AI research, data science, product development, and production operations.
- Emphasis on paved roads, golden paths, automation, and developer autonomy.
- Focus on turning incidents, tickets, and manual runbooks into durable automation.
- Collaboration across R&D, Data Science, Security, and Business teams.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
ComboTech
3 часа назад
Senior DevOps Engineer (igaming)
3 дня назад
Senior DevOps Engineer (Cloud)
155 000 - 175 000$
3 дня назад
Platform Engineer (AWS/Kubernetes)
150 000 - 250 000$
5 дней назад
Platform Engineer (AWS)
170 000 - 210 000$
4 дня назад
Senior Release and Deployment Engineer (DevOps)
5 дней назад