3 дня назад
Systems Engineer (HPC/Cloud)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Systems Engineer (HPC/Cloud): Designing, operating, and improving high-performance compute and storage infrastructure across on-premises and cloud environments with an accent on Linux systems, cloud platforms, containerized workloads, and observability. Focus on automating infrastructure operations, troubleshooting complex OS, storage, networking, and scheduling issues, and optimizing CPU/GPU cluster performance and utilization.
Location: Hong Kong, Hong Kong; hybrid working opportunities available
Company
is a global quantitative trading firm that develops electronic trading infrastructure and operates independent systematic trading teams.
What you will do
- Design, support, and operate HPC compute and storage infrastructure across on-premises and cloud environments.
- Maintain and improve large-scale Linux systems covering compute, storage, networking, automation, and monitoring.
- Troubleshoot issues across operating system, storage, networking, and cluster scheduling layers.
- Manage batch and containerized CPU and GPU workloads across diverse compute resources.
- Develop cloud infrastructure on GCP, AWS, and Azure, together with HPC management tools, access modules, and internal libraries.
- Build observability pipelines, analyze system performance, automate repetitive workflows, and manage deployments and infrastructure lifecycle processes.
Requirements
- Bachelor’s degree or higher in computer science, engineering, or a related field.
- 1–5 years of relevant experience in Linux systems, DevOps, HPC, or infrastructure engineering.
- Strong knowledge of Linux internals, including process scheduling, virtual memory, filesystems, and networking.
- Experience with distributed or networked storage, at least one major cloud provider, and Infrastructure-as-Code or configuration management tools.
- Strong scripting or programming skills in Python, Go, or Bash, plus hands-on experience with Docker or Podman and Kubernetes.
- Understanding of TCP/IP, Ethernet, hardware and server components, troubleshooting, automation, and end-user experience.
Nice to have
- Experience managing GPU-based compute platforms and hybrid on-premises/cloud HPC environments.
- Exposure to CI/CD systems, release automation, performance tuning, scalability, and systems optimization.
- Experience with batch schedulers such as Slurm or HTCondor.
Culture & Benefits
- Collaborative, welcoming, and diverse workplace with limited hierarchy and an emphasis on respectful teamwork.
- Generous paid time off policies and regional savings plans and financial wellness tools.
- Hybrid working opportunities.
- Free breakfast, lunch, and snacks, plus in-office wellness experiences and selected wellness reimbursements.
- Company-sponsored sports teams, fitness events, volunteer opportunities, charitable giving, social events, and continuous learning workshops.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Infrastructure Engineer (AI Hardware)
150 000 - 250 000$
7 часов назад
Associate Linux/Windows Engineer - New Grad
6 часов назад
Associate Linux/Window Engineer (Platform Services)
2 часа назад
HPC Systems Engineer (High-Performance Computing)
Lambda
1 день назад
Site Reliability Engineer (AI Infrastructure)
240 000 - 356 000$
2 дня назад
IT Operations Engineer (Cloud Infrastructure)
100 000 - 155 000$