8 часов назад
Staff Platform Engineer - High Performance Computing Platform Management (HPC)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Platform Engineer - High Performance Computing Platform Management (HPC): Designing, implementing, and managing resilient HPC infrastructure across compute, storage, networking, and job scheduling systems with an accent on high availability, scalability, security, and performance. Focus on optimizing clusters, implementing InfiniBand and Ethernet solutions, managing capacity and resource allocation, and supporting data scientists and developers.
Location: Singapore, Singapore; on-site. Only Singapore Citizens will be considered.
Company
is an agency under Singapore's Ministry of Defence, working on cloud infrastructure and big data platforms.
What you will do
- Lead the delivery of a resilient, scalable, and secure HPC platform covering compute nodes, storage, networks, and job scheduling.
- Design, implement, and manage HPC infrastructure to support organisational requirements.
- Build storage and high-performance networking solutions using SAN, NAS, object storage, InfiniBand, and Ethernet.
- Plan capacity, forecast resource needs, and manage procurement and deployment of hardware and software.
- Optimize, monitor, and troubleshoot HPC clusters, job scheduling, and resource allocation.
- Collaborate with data scientists and developers to optimize application performance and provide platform support.
Requirements
- Degree in Computer Science, Computer Engineering, or a related field.
- 8+ years of experience managing HPC systems and strong knowledge of HPC architectures, including clusters, grids, and clouds.
- Experience with Linux or Unix, HPC schedulers such as Slurm, Torque, or LSF, and storage systems including SAN, NAS, and object storage.
- Experience with high-performance networking, cloud platforms such as AWS, Azure, or Google Cloud, and scripting with Python, Perl, or Bash.
- Experience with Docker, Kubernetes, Knative, Run:AI, Grafana, Prometheus, Kyverno, ArgoCD, Rancher, NVIDIA BCM, and NVIDIA SuperPOD architecture.
- Experience leading engineering teams and managing security and compliance for infrastructure platforms. Singapore citizenship is required.
Nice to have
- Certifications in NVIDIA AI Infrastructure and Operations or Certified Kubernetes Administrator.
- Experience with machine learning or deep learning frameworks such as TensorFlow or PyTorch.
- Familiarity with agile development methodologies and Git.
Culture & Benefits
- Purposeful work focused on engineering and operational excellence.
- Modern technologies and technical stacks.
- Strong engineering culture and work-life balance.
- Environment that supports innovation and professional growth.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
8 часов назад
Staff Slurm Cluster & HPC Engineer
15 часов назад
Systems Engineer - R&D (Linux/HPC)
150 000 - 300 000$
8 часов назад
Network Reliability Engineer
9 часов назад
Network Engineer (AI Infrastructure)
5 дней назад
Systems Engineer (Weekend Warrior)
8 часов назад