Назад
Company hidden
7 дней назад

Cloud Senior DevOps Engineer (AI)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Cloud Senior DevOps Engineer (AI): Building the CI/CD, infrastructure-as-code, MLOps, and Internal Developer Platform foundations for an AI-operated GPU cloud with an accent on Kubernetes, multi-cloud infrastructure, high availability, and observability. Focus on designing automated deployment paths for software and machine-learning models, provisioning GPU clusters, enforcing security and compliance, and leading incident remediation.

Location: Remote within San Jose, California, or Austin, Texas

Company

hirify.global develops Bitcoin mining infrastructure, AI computational infrastructure, and cloud capabilities for high-demand artificial intelligence workloads.

What you will do

  • Design, implement, and maintain end-to-end GitLab-based CI/CD and MLOps pipelines for software applications and machine-learning models.
  • Build and scale Kubernetes- and Docker-based cloud-native infrastructure, including specialized GPU clusters for AI workloads and model inference.
  • Own high-availability architecture, disaster recovery, self-healing, capacity planning, and production performance tuning.
  • Automate reproducible infrastructure provisioning across cloud environments with Terraform, Ansible, and Helm.
  • Build observability systems using monitoring, logging, alerting, and AI model telemetry.
  • Develop the Internal Developer Platform, improve deployment workflows, enforce security and compliance standards, and lead major incident resolution.

Requirements

  • Bachelor’s degree or higher in Computer Science, Engineering, or a related technical field, plus 5+ years of hands-on DevOps, SRE, or cloud infrastructure experience.
  • Expert knowledge of Linux and networking fundamentals, including TCP/IP, DNS, HTTP, load balancing, and VPCs.
  • Deep production-level expertise with Docker and Kubernetes.
  • Experience designing and managing infrastructure on public or hybrid cloud platforms such as AWS, GCP, Azure, or Alibaba Cloud, including multi-cloud strategies.
  • Strong programming or scripting skills in at least one major language, such as Go, Python, or Shell.
  • Practical understanding of CI/CD, infrastructure as code, observability, and SRE principles, with strong problem-solving and cross-team communication skills.

Nice to have

  • Experience with MLOps, model serving or inference frameworks such as vLLM, TGI, or Triton Inference Server, and GPU cluster management.
  • Experience with large-scale distributed systems, Internal Developer Platforms, Zero Trust, DevSecOps, SOC 2, or ISO 27001.
  • Technical leadership, mentoring, or DevOps team management experience.
  • Experience applying LLM-driven automation to code, configuration, pipelines, or incident response.

Culture & Benefits

  • Work on infrastructure connecting AI research, data science, product development, and production operations.
  • Emphasis on paved roads, golden paths, automation, and developer autonomy.
  • Focus on turning incidents, tickets, and manual runbooks into durable automation.
  • Collaboration across R&D, Data Science, Security, and Business teams.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →