13 часов назад
GPU Systems Engineer (AI)
150 000 - 300 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
GPU Systems Engineer (AI) (HPC/GPU infrastructure): Designing, building, and optimizing large-scale distributed GPU compute clusters for trading and research with an accent on performance tuning, workload profiling, and infrastructure automation. Focus on diagnosing bottlenecks across compute, storage, and networking layers, deploying systems across thousands of nodes, and resolving complex hardware, operating system, and network issues.
Location: Austin, TX, United States; Chicago, Illinois, United States; London, United Kingdom; New York, NY, United States
Base salary: 150,000–300,000 USD per year, plus discretionary performance-based bonuses and benefits.
Company
applies a scientific approach to financial trading and operates a large-scale computing environment for algorithmic trading research and development.
What you will do
- Design, build, and optimize distributed GPU compute clusters for HPC and AI workloads.
- Identify and resolve performance bottlenecks across compute, storage, and networking layers.
- Profile, benchmark, and fine-tune GPU-based workloads with research and development teams.
- Automate system deployment, monitoring, and troubleshooting across thousands of nodes.
- Own critical infrastructure projects from concept through implementation and ongoing support.
- Test and deploy hardware and software, partnering with vendors to resolve complex issues.
Requirements
- 5+ years of large-scale Linux systems engineering experience in HPC, AI, or distributed infrastructure.
- Extensive experience with Linux installation, performance tuning, and troubleshooting.
- Expertise in distributed GPU workload troubleshooting and GPU optimization.
- Proficiency in Python scripting and automation frameworks.
- Experience with NVIDIA technologies including NCCL, GPUDirect RDMA, and NVLink.
- Familiarity with configuration management tools such as Salt, Ansible, Puppet, or Chef, and the ability to diagnose hardware, operating system, and network issues.
Nice to have
- CUDA or C/C++ experience.
Culture & Benefits
- Collaborative environment spanning research, engineering, and infrastructure teams.
- Work on high-impact automation and computing challenges in algorithmic trading.
- Culture focused on openness, transparency, diverse expertise, and togetherness.
- Competitive benefits package and eligibility for discretionary performance-based bonuses.
- AI tools are prohibited during interviews or assessments unless explicitly authorized.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
13 часов назад
Senior Performance Engineer (AI)
135 000 - 170 000$
2 дня назад
Systems Administrator (HPC)
150 000 - 175 000$
6 дней назад
Infrastructure Engineer (GPU & Compute)
180 000 - 220 000$
Lambda
5 дней назад
Site Reliability Engineer (AI Infrastructure)
240 000 - 356 000$
6 дней назад
Infrastructure Engineer (Storage)
180 000 - 220 000$
6 часов назад