Назад
Company hidden
1 день назад

Infrastructure Operations Engineer (APAC) (AI)

165 000 - 205 000SGD
Формат работы
remote (только Singapore)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Infrastructure Operations Engineer (APAC) (AI): Operating and scaling large-scale GPU infrastructure across Linux, bare metal systems, provisioning workflows, and reliability operations with an accent on automation, observability, and incident response. Focus on building infrastructure platforms, troubleshooting complex systems, and reducing manual toil through Kubernetes, Terraform, Ansible, and software automation.

Location: Remote for candidates residing in Singapore; Monday–Friday, 8:00 AM–5:00 PM local time (UTC+8), with regular on-call participation.

Annual base salary: SGD 165,000–205,000.

Company

hirify.global develops an end-to-end platform for developing, training, and deploying AI systems, combining developer-focused software with large-scale GPU compute infrastructure.

What you will do

  • Design, build, and roll out infrastructure platforms and operational patterns that reduce incidents and support customer-facing and internal features.
  • Deploy updates and improvements for internal and end-customer use cases.
  • Operate large-scale GPU environments, Linux systems, bare-metal infrastructure, and provisioning workflows.
  • Improve platform reliability, observability, and operational efficiency through automation.
  • Collaborate with Infrastructure Engineering, Network Operations, Customer Success, and Software Platform teams.
  • Participate in a distributed primary/secondary on-call rotation.

Requirements

  • 8+ years of experience working with Linux as a server or hosting platform.
  • 5+ years of experience with AWS.
  • At least 2 years of experience with Kubernetes, container fundamentals, Terraform, and Ansible.
  • At least 2 years managing network-attached storage using NFS, Ceph, or similar protocols.
  • Experience with Prometheus, the ELK stack, GitOps workflows, and automation using Python, Go, Bash, or other languages.
  • Strong networking fundamentals, experience building complex systems, and effective written and verbal communication.

Nice to have

  • Experience troubleshooting and provisioning bare-metal hardware, including Dell systems.
  • Experience with GPU servers in bare-metal or virtualized environments.
  • Experience with SONiC switches, Palo Alto firewalls, Juniper Networks, 400Gb Ethernet, or InfiniBand.
  • Experience with VAST storage systems.

Culture & Benefits

  • Medical, dental, and vision coverage for employees and eligible dependents.
  • Equity through RSUs, performance-based discretionary bonus, and location-dependent retirement benefits.
  • Unlimited PTO, company holidays, floating holidays, and a two-week winter company closure.
  • Paid parental and family leave, wellness and work-from-home stipends, and an annual learning allowance.
  • Four weeks of paid sabbatical leave after four years of service.
  • Flexible schedules and hybrid work options for office-based teams; benefits may vary by location, team, and role.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →