3 дня назад
Systems Engineer (HPC)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Systems Engineer (HPC): Designing, operating, and improving high-performance compute and storage infrastructure across on-premises and cloud environments with an accent on Linux systems, containerized workloads, and observability. Focus on troubleshooting complex infrastructure layers, automating repetitive workflows, and optimizing CPU/GPU cluster utilization across GCP, AWS, and Azure.
Location: Singapore, Singapore; hybrid working opportunities are available.
Company
is a quantitative trading firm developing high-performance electronic trading infrastructure and business support platforms.
What you will do
- Design, support, and operate HPC compute and storage infrastructure across on-premises and cloud environments.
- Maintain and improve large-scale Linux systems covering compute, storage, networking, automation, and monitoring.
- Troubleshoot issues across operating system, storage, networking, and cluster scheduling layers.
- Manage batch and containerized CPU and GPU workloads across diverse compute resources.
- Develop cloud infrastructure on GCP, AWS, and Azure, along with HPC management tools, access modules, and internal libraries.
- Build observability pipelines, improve cluster utilization, manage infrastructure deployments and upgrades, and automate repetitive workflows.
Requirements
- Bachelor’s degree or higher in computer science, engineering, or a related field.
- 1–5 years of relevant experience in Linux systems, DevOps, HPC, or infrastructure engineering.
- Strong knowledge of Linux internals, networking fundamentals, and hardware and server components.
- Experience with distributed or networked storage and at least one major cloud provider: GCP, AWS, or Azure.
- Familiarity with Infrastructure-as-Code and configuration management tools such as Ansible, Terraform, or Salt.
- Strong scripting or programming skills in Python, Go, or Bash, plus hands-on experience with Docker or Podman and Kubernetes.
Nice to have
- Experience with Slurm or HTCondor batch schedulers and GPU-based compute platforms.
- Exposure to CI/CD systems, release automation, and hybrid on-premises and cloud HPC environments.
- Interest in performance tuning, scalability, and systems optimization.
Culture & Benefits
- Collaborative, welcoming workplace with a diverse team and minimal hierarchy.
- Generous paid time off and regional savings and financial wellness plans.
- Free breakfast, lunch, and snacks daily.
- In-office wellness experiences, wellness expense reimbursement, sports teams, and fitness events.
- Volunteer opportunities, charitable giving, social events, and celebrations.
- Workshops and continuous learning opportunities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →