2 месяца назад
Infrastructure Operations Engineer (APAC) (AI)
165 000 - 205 000SGD
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Infrastructure Operations Engineer (APAC) (AI): Operating and scaling large-scale GPU infrastructure across Linux, bare metal systems, provisioning workflows, and reliability operations with an accent on automation, observability, and incident response. Focus on building infrastructure platforms, troubleshooting complex systems, and reducing manual toil through Kubernetes, Terraform, Ansible, and software automation.
Location: Remote for candidates residing in Singapore; Monday–Friday, 8:00 AM–5:00 PM local time (UTC+8), with regular on-call participation.
Annual base salary: SGD 165,000–205,000.
Company
develops an end-to-end platform for developing, training, and deploying AI systems, combining developer-focused software with large-scale GPU compute infrastructure.
What you will do
- Design, build, and roll out infrastructure platforms and operational patterns that reduce incidents and support customer-facing and internal features.
- Deploy updates and improvements for internal and end-customer use cases.
- Operate large-scale GPU environments, Linux systems, bare-metal infrastructure, and provisioning workflows.
- Improve platform reliability, observability, and operational efficiency through automation.
- Collaborate with Infrastructure Engineering, Network Operations, Customer Success, and Software Platform teams.
- Participate in a distributed primary/secondary on-call rotation.
Requirements
- 8+ years of experience working with Linux as a server or hosting platform.
- 5+ years of experience with AWS.
- At least 2 years of experience with Kubernetes, container fundamentals, Terraform, and Ansible.
- At least 2 years managing network-attached storage using NFS, Ceph, or similar protocols.
- Experience with Prometheus, the ELK stack, GitOps workflows, and automation using Python, Go, Bash, or other languages.
- Strong networking fundamentals, experience building complex systems, and effective written and verbal communication.
Nice to have
- Experience troubleshooting and provisioning bare-metal hardware, including Dell systems.
- Experience with GPU servers in bare-metal or virtualized environments.
- Experience with SONiC switches, Palo Alto firewalls, Juniper Networks, 400Gb Ethernet, or InfiniBand.
- Experience with VAST storage systems.
Culture & Benefits
- Medical, dental, and vision coverage for employees and eligible dependents.
- Equity through RSUs, performance-based discretionary bonus, and location-dependent retirement benefits.
- Unlimited PTO, company holidays, floating holidays, and a two-week winter company closure.
- Paid parental and family leave, wellness and work-from-home stipends, and an annual learning allowance.
- Four weeks of paid sabbatical leave after four years of service.
- Flexible schedules and hybrid work options for office-based teams; benefits may vary by location, team, and role.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Platform Automation Engineer
100 000 - 150 000$
10 дней назад
AI Platform Engineer
130 000 - 180 000$
4 дня назад
Systems Automation Engineer (Terraform)
120 000 - 135 000$
3 дня назад
Senior Azure Cloud Engineer (DevOps)
130 000 - 170 000$
6 дней назад
DevOps Automation Engineer (Terraform)
75 000 - 95 000$