Назад
Company hidden
4 дня назад

AI Cloud Senior DevOps Engineer

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/US/Norway +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Cloud Senior DevOps Engineer (Kubernetes/Docker/MLOps): Building and operating cloud-native infrastructure, CI/CD pipelines, and GPU clusters for AI products and model inference with an accent on high availability, infrastructure as code, and observability. Focus on automating multi-cloud deployments, designing disaster recovery and self-healing systems, and resolving complex production incidents.

Location: Remote within San Jose, CA or Austin, TX

Company

hirify.global provides Bitcoin mining infrastructure, data center operations, and AI cloud computing capabilities.

What you will do

  • Design and maintain end-to-end CI/CD and MLOps pipelines for software applications and machine learning models.
  • Build, optimize, and scale Kubernetes- and Docker-based cloud infrastructure, including GPU clusters for AI workloads and model inference.
  • Design high-availability production environments with disaster recovery, self-healing, capacity planning, and performance tuning.
  • Automate reproducible infrastructure provisioning across cloud environments using Terraform, Ansible, and Helm.
  • Develop monitoring, logging, and alerting systems with tools such as Prometheus, Grafana, and ELK/EFK.
  • Lead incident resolution, root cause analysis, security governance, release workflows, and collaboration with R&D, Data Science, Security, and Business teams.

Requirements

  • 5+ years of hands-on experience in DevOps, SRE, or cloud infrastructure roles.
  • Bachelor's degree or higher in Computer Science, Engineering, or a related technical field.
  • Expert knowledge of Linux and networking principles, including TCP/IP, DNS, HTTP, load balancing, and VPCs.
  • Deep production-level experience with Docker and Kubernetes.
  • Experience designing and managing AWS, GCP, Azure, Alibaba Cloud, or comparable public and hybrid cloud infrastructure.
  • Strong coding or scripting skills in Go, Python, Shell, or another major language, plus knowledge of CI/CD, IaC, observability, and SRE practices.

Nice to have

  • MLOps, model serving, model inference, or GPU cluster experience, including vLLM, TGI, or Triton Inference Server.
  • Experience with large-scale distributed or high-concurrency systems and Internal Developer Platforms.
  • Knowledge of Zero Trust, DevSecOps, SOC2, or ISO27001.
  • Technical leadership, mentoring, or DevOps team management experience.

Culture & Benefits

  • Full-time role within the AI Cloud team.
  • Cross-functional collaboration across engineering, data science, security, and business functions.
  • Work focused on automation, developer productivity, platform engineering, system reliability, and security.
  • Equal employment opportunity commitment in accordance with applicable laws.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →