Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff/Senior Staff Engineer, Kubernetes (AWS/Alibaba Cloud): Operating and scaling large-scale production Kubernetes clusters across Alibaba Cloud and AWS with an accent on multi-cloud governance, high availability, security, and observability. Focus on automating releases and infrastructure, resolving complex cluster failures, optimizing cloud-native workloads, and leading incident response and post-mortem improvements.
Location: Singapore, Singapore. Applicants must have the current right to work in Singapore and must not require visa sponsorship.
Company
OKX is a crypto exchange and developer of OKX Wallet and other blockchain products serving retail users and large institutions.
What you will do
- Own the lifecycle, scaling, upgrades, operations, fault diagnosis, and performance tuning of large-scale production Kubernetes clusters.
- Operate and optimize Alibaba Cloud and AWS environments, including container services, networking, storage, IAM, load balancing, databases, object storage, cost management, and disaster recovery.
- Lead containerization and microservices operations while improving Pod scheduling, resource quotas, network policies, image management, and logging.
- Build monitoring, alerting, logging, and distributed tracing systems; define runbooks, change processes, incident response plans, and security controls.
- Develop Shell and Python automation and integrate Jenkins, GitLab CI, ArgoCD, Terraform, Ansible, and Helm for automated releases, inspections, backups, and infrastructure management.
- Lead incident response, root-cause analysis, post-mortems, technical documentation, and knowledge sharing for multi-cloud Kubernetes operations.
Requirements
- Bachelor’s degree or above in computer science or a related field.
- 4+ years of hands-on experience operating production Kubernetes clusters and independently resolving complex failures and performance bottlenecks.
- Strong knowledge of Kubernetes components, scheduling, networking, storage, cluster management, containerd/Docker, Calico/Flannel, CSI, and multi-cluster operations.
- Production experience operating Alibaba Cloud and AWS, including ACK, ECS, EKS, EC2, VPC, IAM, load balancing, databases, object storage, monitoring, and disaster recovery.
- Proficiency in Linux administration, Shell and Python automation, CI/CD, Infrastructure as Code, service mesh, and observability tools.
- Current right to work in Singapore is required; visa sponsorship is not available.
Nice to have
- Experience with 100+ node public cloud environments, multi-cloud cost optimization, and Kubernetes security hardening.
- CKA or CKS certification.
- Experience scheduling AI/LLM workloads, including GPU scheduling and distributed training.
Culture & Benefits
- Learning and development programs with education subsidies.
- Team-building programs and company events.
- Wellness and meal allowances.
- Comprehensive healthcare coverage for employees and dependants.
- Competitive total compensation package.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
1 день назад
DevSecOps Engineer (Kubernetes)
7 дней назад
Cloud DevOps Engineer (AI)
9 дней назад
Senior DevOps Engineer (Kubernetes)
PrimeHunters
2 часа назад
Senior DevOps Engineer (Fintech)
4 часа назад
Senior DevOps Engineer (Fintech)
13 дней назад
Staff Infrastructure & DevOps Engineer (AWS)
211 000 - 235 000$