обновлено 1 месяц назад
Senior Network Engineer (AI/HPC)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Network Engineer (AI/HPC): Designing, maintaining, and scaling InfiniBand and Ethernet networks for Linux-based high-performance computing and artificial intelligence environments with an accent on low-latency networking, data center infrastructure, and operational reliability. Focus on building network architectures, optimizing RDMA/RoCE and routing environments, automating operations, and resolving complex issues across cloud-scale systems.
Location: Hybrid in South Korea, with 3 days in the office. Conditions: Seongnam-si, Gyeonggi-do, South Korea
Company
is a global technology company that designs, builds, deploys, and manages high-performance AI, HPC, memory, compute, and infrastructure solutions.
What you will do
- Develop network configurations and architectures for projects and technology roadmaps.
- Design, enhance, maintain, and upgrade InfiniBand and Ethernet networks.
- Validate network infrastructure with solutions architects and field engineers against customer requirements.
- Diagnose and resolve network issues at scale, coordinate escalations, and work with vendors and engineering teams.
- Support rack lifecycle processes and cloud-scale network and storage environments.
- Document configurations, procedures, and best practices while developing automation and operational tools.
Requirements
- 8+ years of hands-on experience with enterprise-scale networks.
- In-depth knowledge of data center environments, servers, and network equipment.
- Proven experience with NVIDIA InfiniBand and Ethernet networks using Cumulus, Arista, or Juniper.
- Experience installing, monitoring, maintaining, and documenting data center network equipment and processes.
- Ability to act as a technical escalation point, resolve problems, and communicate clearly with team members and clients.
- Participation in a weekly on-call rotation and willingness to respond to network and server errors after hours.
Nice to have
- Experience managing RDMA or RoCE environments.
- Experience with low-latency, high-bandwidth networking optimization.
- Knowledge of VXLAN, EVPN, BGP, and OSPF.
- Knowledge of HPC libraries and GPU networking technologies, including NCCL, UCX, MPI, NVLink, and NVSwitch.
- Familiarity with UFM or OpenSM.
Culture & Benefits
- Work in a fast-paced, complex managed services environment supporting HPC and AI workloads.
- Collaborate across internal teams, customers, vendors, and engineering organizations.
- Participate in continuous learning, technology research, knowledge-sharing, and team activities.
- Work in a flexible, outcome-focused environment that values ownership, innovation, and servant leadership.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Senior Systems Engineer (AI)
250 000 - 400 000$
12 дней назад
Principal Operations Engineer, Network (AI)
225 000 - 270 000$
4 дня назад
SME Network Engineer
130 000 - 160 000$
4 дня назад
Cloud Network Security Engineer (Azure/AWS)
175 000 - 200 000$
14 дней назад
Senior Network Engineer
12 дней назад
Systems Engineer (Enterprise Networking)
98 000 - 138 000$