3 дня назад
Network Engineer, AI Cluster Commissioning (AI Networking)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Network Engineer, AI Cluster Commissioning (AI Networking): Deploying, commissioning, and optimizing high-performance Ethernet and RDMA networks for large-scale AI clusters with an accent on low-latency interconnects, RoCE fabrics, network automation, and security. Focus on diagnosing layer 1–4 issues, validating multi-node performance, automating network changes, and coordinating on-site deployments across Australia and Southeast Asia.
Location: Based in Australia or Singapore, with regular visits to project sites in Australia and Southeast Asia
Company
develops and operates energy-efficient AI infrastructure, including large-scale GPU cloud platforms and AI Factories, across the Asia Pacific region.
What you will do
- Deploy, configure, commission, test, and benchmark high-throughput Ethernet and RDMA networks for multi-node AI clusters.
- Diagnose and remediate Layer 1–4 network issues, congestion, link errors, and performance bottlenecks across HPC, storage, and AI environments.
- Develop monitoring tools and automate provisioning, backups, compliance checks, and network configuration using Ansible, NetBox, Bash, and Python.
- Implement CI/CD workflows for network changes and maintain configuration consistency, auditability, and topology visibility.
- Coordinate commissioning activities with engineering, operations, infrastructure, security, vendors, and technology partners.
- Support network security through segmentation, access control, zero-trust architecture, vulnerability remediation, SIEM integration, and incident response.
Requirements
- Bachelor’s degree in network engineering, computer science, or a related technical field.
- 5+ years of network engineering experience, including Linux host networking and complex technical projects.
- Experience with high-performance Ethernet, IPv4, IPv6, BGP, and RoCE.
- Knowledge of network automation and Infrastructure as Code tools, including Ansible, NetBox, Bash, and Python.
- Strong analytical, problem-solving, communication, collaboration, and project management skills.
- Willingness to travel internationally and domestically for on-site deployments and commissioning.
Nice to have
- Hands-on experience with the NVIDIA Spectrum Ethernet Platform and RoCE.
- NVIDIA InfiniBand, subnet manager configuration, UFM, and multi-tenancy networking.
- Network security, firewall management, segmentation, secure network design, and zero-trust architecture.
- Experience with Cumulus Linux, SONiC, NVIDIA Air, DPU/SmartNICs, and network performance tools.
Culture & Benefits
- Work alongside founders and specialists in AI infrastructure, energy systems, and next-generation computing.
- Founder-led environment with accessible leadership, fast decisions, and limited bureaucracy.
- Early ownership and opportunities to grow into new technical domains.
- Exposure to large-scale AI infrastructure projects across the Asia Pacific region.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Network Deployment Engineer (Networking)
8 дней назад
Professional Services Engineer (High-Performance Storage)
8 дней назад
Cloud Network Test Engineer
2 дня назад
Senior Technical Specialist (Network Operations)
2 дня назад
Principal Specialist Internet Optimisation
7 дней назад