обновлено 1 месяц назад
Systems Engineer (HPC/Cloud)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Systems Engineer (HPC/Cloud): Designing, operating, and improving high-performance computing and storage infrastructure across on-premises and cloud environments with an accent on Linux systems, networking, automation, and observability. Focus on managing CPU and GPU workloads, developing infrastructure across GCP, AWS, and Azure, and solving complex performance and reliability challenges across distributed systems.
Location: Singapore, Singapore
Company
Quantitative trading firm developing electronic trading infrastructure, market access, data, compute, research, and risk management platforms.
What you will do
- Design, support, and operate HPC compute and storage infrastructure across on-premises and cloud environments.
- Maintain and improve large-scale Linux systems covering compute, storage, networking, automation, and monitoring.
- Troubleshoot complex issues across operating systems, storage, networking, and cluster scheduling layers.
- Manage and optimize batch and containerized CPU and GPU workloads across diverse compute resources.
- Develop cloud infrastructure across GCP, AWS, and Azure, including management tools, access modules, and internal libraries.
- Build observability pipelines, improve cluster utilization, manage infrastructure lifecycle changes, and automate repetitive workflows.
Requirements
- Bachelor’s degree or higher in computer science, engineering, or a related field.
- 1–5 years of experience in Linux systems, DevOps, HPC, or infrastructure engineering.
- Strong understanding of Linux internals, including process scheduling, virtual memory, filesystems, and networking.
- Experience with distributed or networked storage, at least one major cloud provider, Infrastructure-as-Code or configuration management tools, and container technologies.
- Strong scripting or programming skills in Python, Go, or Bash, together with knowledge of TCP/IP, Ethernet, and server hardware.
- Strong troubleshooting, communication, automation, and operational excellence skills with a focus on end-user experience.
Nice to have
- Experience with Slurm or HTCondor batch schedulers and GPU-based compute platforms.
- Exposure to CI/CD systems, release automation, and hybrid on-premises and cloud HPC environments.
- Interest in performance tuning, scalability, and systems optimization.
Culture & Benefits
- Hybrid working opportunities in a collaborative and welcoming workplace.
- Generous paid time off and regional savings plans and financial wellness tools.
- Free breakfast, lunch, and snacks daily.
- In-office wellness experiences, wellness expense reimbursement, sports teams, and fitness events.
- Volunteer opportunities, charitable giving, social events, and continuous learning workshops.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →