2 дня назад
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure) (GPU cloud infrastructure): Installing, commissioning, operating, and maintaining regional and core GPU cloud deployments across datacenter environments with an accent on hardware bring-up, networking, platform installation, and operational readiness. Focus on validating high-density AI infrastructure, troubleshooting GPU and network fabrics, integrating Kubernetes and storage platforms, and producing reliable handover documentation.
Location: Europe
Company
Radian Arc provides an infrastructure-as-a-service platform for cloud gaming, artificial intelligence, and machine learning applications inside telecommunications carrier networks.
What you will do
- Coordinate physical deployment, rack-and-stack activities, cabling validation, asset tracking, and datacenter handover for GPU, compute, storage, and network infrastructure.
- Bring up servers and validate BIOS, BMC, firmware, NIC, DPU, GPU, NVMe, RAID, PCIe topology, thermals, and hardware health.
- Support OOB, north-south, storage, and east-west network deployment, including BGP, ECMP, VLAN/VRF, EVPN/VXLAN, OVS/OVN, and RoCE/RDMA validation.
- Install and validate Linux hosts, NVIDIA drivers, CUDA, containers, KVM/QEMU, CloudStack, Kubernetes, KubeVirt, GPU Operator, and storage integrations.
- Perform maintenance, incident response, root-cause analysis, acceptance testing, performance baselining, and operational readiness reporting.
- Maintain telemetry, dashboards, runbooks, as-built records, and deployment procedures while coordinating with engineering and external infrastructure vendors.
Requirements
- Strong hands-on experience deploying and maintaining datacenter infrastructure from bare metal through production readiness.
- Experience with GPU, HPC, AI cloud, private cloud, or high-density compute environments; familiarity with HGX, NVL72-style architectures, NVLink/NVSwitch, power density, and cooling requirements.
- Practical expertise across Linux troubleshooting, GPU servers, NVIDIA drivers, firmware, PCIe topology, networking, storage, and OOB management.
- Knowledge of VLANs, VRFs, BGP, ECMP, OVS/OVN, routing, cabling, optics, transceivers, and datacenter handover.
- Strong documentation skills, including accurate as-built records, asset records, validation reports, and operational runbooks.
- Ability to troubleshoot across hardware, network, host, and platform layers and communicate field issues clearly to engineering and leadership.
Nice to have
- Experience with liquid-cooled GPU deployments and next-generation high-density AI infrastructure.
Culture & Benefits
- Attractive compensation package based on expertise and experience.
- Friendly, internationally diverse, flexible, and hybrid-friendly work environment.
- Opportunity to join a fast-growing scale-up focused on GPU-based edge computing and AI infrastructure.
- Career development opportunities as the platform evolves toward core AI infrastructure.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →