23 часа назад
GPU Systems Engineer
200 000 - 300 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
GPU Systems Engineer (GPU infrastructure/HPC): Design, deploy, and operate distributed GPU clusters spanning thousands of nodes, with an accent on Linux systems, GPU workload performance, and infrastructure automation. Focus on profiling AI workloads, troubleshooting GPUDirect RDMA across hardware and network layers, and building self-healing tooling for large-scale compute fleets.
Location: New York, hybrid
Salary: $200,000–$300,000 annual base salary, plus eligible discretionary bonus.
Company
is a quantitative trading firm that develops high-performance electronic trading infrastructure and operates a global engineering organization.
What you will do
- Design, deploy, scale, and operate distributed GPU clusters, including hardware selection and network topology.
- Identify performance bottlenecks across compute, storage, networking, and their integration points.
- Profile and benchmark GPU workloads with researchers and convert findings into measurable performance improvements.
- Build provisioning, monitoring, diagnostics, and self-healing automation for fleets of thousands of nodes.
- Own infrastructure projects from architecture and implementation through long-term support.
- Qualify new hardware and software generations and work with vendors to resolve complex issues.
Requirements
- 5+ years of experience engineering large-scale Linux systems in HPC, AI, or distributed-infrastructure environments.
- Deep Linux knowledge, including installation, performance tuning, debugging, and kernel-level investigation.
- Hands-on experience troubleshooting distributed GPU workloads and understanding GPU performance.
- Working experience with GPUDirect RDMA and data movement between GPUs and networks.
- Python for automation and tooling, plus CUDA or C/C++ experience for reading, profiling, and debugging GPU code.
- Experience with configuration management tools such as Salt, Ansible, Puppet, or Chef, and the ability to diagnose issues across hardware, operating-system, and network layers.
Nice to have
- Experience with additional parts of the NVIDIA stack, including NCCL and NVLink.
Culture & Benefits
- Hybrid working opportunities in a collaborative, low-hierarchy environment.
- Generous paid time off policies.
- Savings plans and financial wellness tools available in each region.
- Free breakfast, lunch, and snacks in the office.
- Wellness reimbursements, sports teams, fitness events, volunteer opportunities, and charitable giving.
- Workshops and continuous learning opportunities, along with regular social events.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
1 день назад
Sr. Software Engineer (Systems Engineering, C/C++)
196 350 - 292 600$
1 день назад
Systems Software Engineer (AI Cloud)
120 000 - 180 000$
1 день назад
Infrastructure Engineer (AI Hardware)
150 000 - 250 000$
1 день назад
Senior Infrastructure Engineer (GPU Infrastructure)
180 000 - 300 000$
2 дня назад
Software Engineer – Linux Infrastructure and Distributed Computing (C++/Linux)
136 300 - 231 700$
1 день назад
Systems Operations Support Engineer (AI Infrastructure)
90 000 - 160 000$