2 дня назад
Cloud Service Security Platform DevOps & Maintenance (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Cloud Service Security Platform DevOps & Maintenance (Kubernetes/Terraform): Building the CI/CD, infrastructure-as-code, MLOps pipelines, and internal developer platform for an AI-operated GPU cloud with an accent on cloud-native infrastructure, high availability, and security governance. Focus on managing GPU clusters, designing observability and disaster recovery, and automating incident remediation across multi-cloud environments.
Location: Remote within San Jose, California or Austin, Texas
Company
Technologies Group builds Bitcoin mining infrastructure, AI computational infrastructure, data centers, and cloud capabilities for high-demand artificial intelligence workloads.
What you will do
- Design, implement, and maintain CI/CD pipelines for software applications and machine learning models, including automated testing, deployment, rollback, and release workflows.
- Build and scale cloud-native infrastructure with Kubernetes and Docker, including specialized GPU clusters for AI workloads and model inference.
- Provision reproducible multi-cloud infrastructure with Terraform, Ansible, Helm, and other infrastructure-as-code tools.
- Design high-availability architecture, disaster recovery, self-healing mechanisms, capacity planning, and performance tuning for production systems.
- Build monitoring, logging, and alerting systems with Prometheus, Grafana, and ELK/EFK, including telemetry for AI model metrics.
- Collaborate with R&D, data science, security, and business teams; lead incident response, root-cause analysis, and preventative automation.
Requirements
- Bachelor's degree or above in Computer Science, Engineering, or a related technical field, plus 5+ years of DevOps, SRE, or cloud infrastructure experience.
- Expert knowledge of Linux and networking principles, including TCP/IP, DNS, HTTP, load balancing, and VPCs.
- Deep production experience with Docker and Kubernetes orchestration.
- Experience designing and managing infrastructure on public or hybrid cloud platforms such as AWS, GCP, Azure, or Alibaba Cloud.
- Strong programming or scripting skills in Go, Python, Shell, or another major language, with an automation-focused engineering mindset.
- Practical knowledge of CI/CD, infrastructure as code, observability, SRE, security standards, Zero Trust access controls, secrets management, and compliance frameworks such as SOC 2 or ISO 27001.
Nice to have
- Experience with MLOps, model serving or inference frameworks such as vLLM, TGI, or Triton Inference Server, and GPU cluster management.
- Experience with large-scale distributed systems, high-concurrency environments, or AI platforms.
- Internal Developer Platform design, DevSecOps, automated security testing, or Zero Trust architecture experience.
- Technical leadership, mentoring, or DevOps team management experience.
- Experience integrating LLM-driven code or configuration assistance into engineering pipelines.
Culture & Benefits
- Full-time position within an AI Cloud team building infrastructure for AI products.
- Work focused on developer autonomy through paved roads, golden paths, and automated tooling.
- Cross-functional collaboration across engineering, data science, security, and business teams.
- Equal employment opportunity in accordance with applicable country, state, and local laws.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →