1 день назад
Infrastructure Operations Engineer (APAC) (AI)
165 000 - 205 000SGD
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Infrastructure Operations Engineer (APAC) (AI): Operating and scaling large-scale GPU infrastructure across Linux, bare metal systems, provisioning workflows, and reliability operations with an accent on automation, observability, and incident response. Focus on building infrastructure platforms, troubleshooting complex systems, and reducing manual toil through Kubernetes, Terraform, Ansible, and software automation.
Location: Remote for candidates residing in Singapore; Monday–Friday, 8:00 AM–5:00 PM local time (UTC+8), with regular on-call participation.
Annual base salary: SGD 165,000–205,000.
Company
develops an end-to-end platform for developing, training, and deploying AI systems, combining developer-focused software with large-scale GPU compute infrastructure.
What you will do
- Design, build, and roll out infrastructure platforms and operational patterns that reduce incidents and support customer-facing and internal features.
- Deploy updates and improvements for internal and end-customer use cases.
- Operate large-scale GPU environments, Linux systems, bare-metal infrastructure, and provisioning workflows.
- Improve platform reliability, observability, and operational efficiency through automation.
- Collaborate with Infrastructure Engineering, Network Operations, Customer Success, and Software Platform teams.
- Participate in a distributed primary/secondary on-call rotation.
Requirements
- 8+ years of experience working with Linux as a server or hosting platform.
- 5+ years of experience with AWS.
- At least 2 years of experience with Kubernetes, container fundamentals, Terraform, and Ansible.
- At least 2 years managing network-attached storage using NFS, Ceph, or similar protocols.
- Experience with Prometheus, the ELK stack, GitOps workflows, and automation using Python, Go, Bash, or other languages.
- Strong networking fundamentals, experience building complex systems, and effective written and verbal communication.
Nice to have
- Experience troubleshooting and provisioning bare-metal hardware, including Dell systems.
- Experience with GPU servers in bare-metal or virtualized environments.
- Experience with SONiC switches, Palo Alto firewalls, Juniper Networks, 400Gb Ethernet, or InfiniBand.
- Experience with VAST storage systems.
Culture & Benefits
- Medical, dental, and vision coverage for employees and eligible dependents.
- Equity through RSUs, performance-based discretionary bonus, and location-dependent retirement benefits.
- Unlimited PTO, company holidays, floating holidays, and a two-week winter company closure.
- Paid parental and family leave, wellness and work-from-home stipends, and an annual learning allowance.
- Four weeks of paid sabbatical leave after four years of service.
- Flexible schedules and hybrid work options for office-based teams; benefits may vary by location, team, and role.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Nscale
1 день назад
Infrastructure Software Engineer (AI)
150 000 - 215 000$
6 часов назад
Senior SRE (GPU Infrastructure)
168 000 - 252 000$
SandboxAQ
2 дня назад
Staff Platform Engineer (AI)
121 600 - 228 000$
2 дня назад
Security Platform Engineer (AI)
160 000 - 180 000$
18 часов назад
Infrastructure Engineer
3 дня назад
Senior Systems Engineer (Cloud Infrastructure)
150 000 - 250 000$