обновлено 2 дня назад
HPC Solutions Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
HPC Solutions Engineer (GPU/HPC): Configuring and maintaining large-scale GPU clusters for premium clients with an accent on distributed computing, high-speed networking, machine learning environments, and infrastructure automation. Focus on orchestrating heterogeneous compute resources, tuning system performance, integrating ML workloads into production, and building reusable infrastructure implementations.
Location: Fully remote
Company
A startup building a marketplace that connects independent data centers and compute providers with users seeking diverse, high-performance computing resources.
What you will do
- Lead technical discovery with customers, define requirements and deliverables, and help them use distributed GPU computing resources effectively.
- Recommend tools and develop reusable boilerplate and reference implementations for future customers.
- Manage NVIDIA GPU clusters and coordinate with IT on efficient operation, including InfiniBand networking.
- Deploy and maintain machine learning environments using virtual storage, distributed computing tools, and SLURM.
- Automate infrastructure provisioning and management with Ansible and Terraform.
- Collaborate with data scientists and engineers on production ML integration, performance optimization, documentation, and operational procedures.
Requirements
- 7+ years of experience in high-performance computing, distributed machine learning, GPU computing, and/or system architecture.
- Proficiency managing NVIDIA GPU environments and familiarity with GPU computing frameworks and libraries.
- Strong experience with InfiniBand and other high-speed networking technologies.
- Experience with HPC job schedulers, preferably SLURM.
- Expertise in Ansible and Terraform, plus experience deploying and managing virtual storage solutions.
- Strong Python coding skills and familiarity with machine learning libraries and frameworks, alongside excellent communication and teamwork skills.
Culture & Benefits
- Fully remote work with a high-accountability, high-agency culture.
- Exposure to diverse hardware configurations, distributed compute locations, and cutting-edge GPU use cases.
- Competitive salary, equity, and benefits.
- Flexible paid time off.
- People-centric, mission-driven, experimental working environment focused on individual ownership and results.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Staff Slurm Cluster & HPC Engineer
5 дней назад
HPC Engineer (AI)
5 дней назад
Senior HPC Cluster Engineer
145 920 - 209 241$
3 дня назад
Cluster Administration Engineer (GPU/HPC)
200 000 - 400 000$
3 дня назад
Engineer, Storage and Data Protection (AI/HPC)
4 дня назад
Senior Storage Infrastructure Engineer (AI)
100 000 - 300 000$