7 часов назад
Senior GPU Systems & Fabric Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior GPU Systems & Fabric Engineer (AI): Building the high-performance hardware foundation for AI cloud computing by integrating GPUs, Kubernetes, and low-latency networking with an accent on Linux kernel internals, GPU architectures, and high-speed interconnects. Focus on optimizing RDMA and InfiniBand fabrics, automating GPU and NIC remediation, and enabling topology-aware multi-tenant workloads.
Location: Remote within San Jose, US or Austin, TX
Company
provides Bitcoin mining solutions and AI computational infrastructure, including data center, bare-metal, and cloud capabilities.
What you will do
- Architect and maintain NVIDIA and AMD GPU device plugin and Kubernetes Operator integrations.
- Configure and optimize RDMA, SR-IOV, RoCEv2, and InfiniBand networking for distributed AI training.
- Build automated DCGM-based remediation pipelines to detect, isolate, and reset degraded GPU and NIC components.
- Implement GPU slicing with MIG and vGPU for multi-tenant inference workloads.
- Tune kernel parameters, device drivers, CUDA, and NCCL to optimize containerized AI workloads.
- Collaborate with scheduling and storage teams on topology-aware placement, data movement, hardware standards, and complex performance investigations.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
- 5+ years of systems engineering experience with strong proficiency in Linux kernel internals, C, or Go.
- Hands-on experience with NVIDIA H100/A100 GPU architectures, CUDA runtimes, RDMA, and InfiniBand.
- Deep understanding of containerized environments and Kubernetes device plugin architecture.
- Experience operating, debugging, and scaling bare-metal systems in large-scale production or HPC environments.
- Familiarity with Terraform, Ansible, and CI/CD infrastructure automation.
Nice to have
- Experience in high-velocity, high-growth engineering environments.
Culture & Benefits
- Full-time employment.
- Collaborative work across infrastructure, scheduling, storage, and reliability engineering teams.
- Opportunities to mentor team members and establish documentation standards for an evolving AI hardware stack.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
16 часов назад
HPC Engineer (AI)
18 часов назад
Senior Infrastructure Engineer (OpenStack)
6 дней назад
Senior Network Engineer (InfiniBand / UFM)
170 000 - 210 000$
14 часов назад
Senior HPC Cluster Engineer
145 920 - 209 241$
6 дней назад
Infrastructure Engineer (GPU & Compute)
180 000 - 220 000$
Lambda
5 дней назад
Site Reliability Engineer (AI Infrastructure)
240 000 - 356 000$