3 дня назад
Cloud Senior DevOps Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Cloud Senior DevOps Engineer (Kubernetes/AI Infrastructure): Building and operating cloud-native infrastructure, CI/CD and MLOps pipelines, and GPU clusters for AI workloads with an accent on high availability, observability, security, and infrastructure automation. Focus on designing multi-cloud platforms, resolving complex incidents, implementing disaster recovery, and supporting production model inferencing.
Location: Singapore, Penang, Malaysia, or Taiwan
Company
is a technology company providing Bitcoin mining solutions, ASIC mining hardware, datacenter operations, and AI cloud capabilities.
What you will do
- Design, implement, and maintain CI/CD pipelines for software applications and machine learning models, including automated testing, deployment, and rollback.
- Lead troubleshooting and root cause analysis during complex system anomalies and major incidents.
- Build monitoring, logging, and alerting systems using tools such as Prometheus, Grafana, and ELK/EFK.
- Automate reproducible infrastructure provisioning across cloud environments with Terraform, Ansible, and Helm.
- Design highly available production systems, disaster recovery strategies, self-healing mechanisms, capacity plans, and performance improvements.
- Build and scale Kubernetes- and Docker-based cloud infrastructure, including GPU clusters for AI workloads and model inferencing.
Requirements
- 5+ years of hands-on experience in DevOps, SRE, or cloud infrastructure roles.
- Bachelor's degree or above in Computer Science, Engineering, or a related technical field.
- Strong coding and scripting skills in at least one major language, such as Go, Python, or Shell.
- Experience designing and managing public or hybrid cloud infrastructure, including multi-cloud strategies.
- Deep production-level knowledge of Docker and Kubernetes orchestration.
- Expert knowledge of Linux and networking fundamentals, including TCP/IP, DNS, HTTP, load balancing, and VPCs.
Nice to have
- Experience with MLOps, model serving frameworks such as vLLM, TGI, or Triton Inference Server, and GPU cluster management.
- Experience with large-scale distributed systems, high-concurrency environments, or AI platforms.
- Experience building Internal Developer Platforms and applying Zero Trust, DevSecOps, SOC2, or ISO27001 practices.
- Technical leadership, mentoring, or DevOps team management experience.
Culture & Benefits
- Inclusive environment valuing diverse backgrounds and perspectives.
- Startup-style environment in a fast-growing digital asset technology company.
- Autonomy, personal accountability, and opportunities to contribute to new projects.
- Training, mentoring, welfare benefits, and professional development opportunities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →