Назад
Company hidden
обновлено 2 дня назад

HPC Solutions Engineer

Формат работы
remote (Global)
Тип работы
fulltime
Грейд
senior
Английский
b2
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
HPC Solutions Engineer (GPU/HPC): Configuring and maintaining large-scale GPU clusters for premium clients with an accent on distributed computing, high-speed networking, machine learning environments, and infrastructure automation. Focus on orchestrating heterogeneous compute resources, tuning system performance, integrating ML workloads into production, and building reusable infrastructure implementations.

Location: Fully remote

Company

A startup building a marketplace that connects independent data centers and compute providers with users seeking diverse, high-performance computing resources.

What you will do

  • Lead technical discovery with customers, define requirements and deliverables, and help them use distributed GPU computing resources effectively.
  • Recommend tools and develop reusable boilerplate and reference implementations for future customers.
  • Manage NVIDIA GPU clusters and coordinate with IT on efficient operation, including InfiniBand networking.
  • Deploy and maintain machine learning environments using virtual storage, distributed computing tools, and SLURM.
  • Automate infrastructure provisioning and management with Ansible and Terraform.
  • Collaborate with data scientists and engineers on production ML integration, performance optimization, documentation, and operational procedures.

Requirements

  • 7+ years of experience in high-performance computing, distributed machine learning, GPU computing, and/or system architecture.
  • Proficiency managing NVIDIA GPU environments and familiarity with GPU computing frameworks and libraries.
  • Strong experience with InfiniBand and other high-speed networking technologies.
  • Experience with HPC job schedulers, preferably SLURM.
  • Expertise in Ansible and Terraform, plus experience deploying and managing virtual storage solutions.
  • Strong Python coding skills and familiarity with machine learning libraries and frameworks, alongside excellent communication and teamwork skills.

Culture & Benefits

  • Fully remote work with a high-accountability, high-agency culture.
  • Exposure to diverse hardware configurations, distributed compute locations, and cutting-edge GPU use cases.
  • Competitive salary, equity, and benefits.
  • Flexible paid time off.
  • People-centric, mission-driven, experimental working environment focused on individual ownership and results.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →