Назад
Company hidden
9 часов назад

Automated Testing Engineer (GPU Infrastructure)

172 500 - 210 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Automated Testing Engineer (GPU Infrastructure): Validating large-scale, multi-node GPU clusters and building automated integration testing for high-performance AI and HPC workloads with an accent on CI/CD automation, distributed scaling, and interconnect fabric testing. Focus on developing Python or Go cluster orchestration, benchmarking NCCL/RCCL collective communication, and analyzing performance bottlenecks across Linux, hypervisor, and physical network layers.

Location: San Francisco or Sunnyvale, California, US; onsite

Salary: $172,500–$210,000 per year, plus Restricted Stock Units

Company

hirify.global is an AI infrastructure company building vertically integrated energy, data center, and cloud systems for demanding AI workloads.

What you will do

  • Build CI/CD platforms and automated integration testing for low-level infrastructure and distributed control planes.
  • Design and execute validation tests across multi-node virtualized GPU clusters to verify scaling, stability, and multi-tenant isolation.
  • Develop Python or Go automation frameworks to provision, configure, and stress-test cluster environments.
  • Validate NVLink, Infinity Fabric, InfiniBand, and RoCE interconnects for low-latency, high-bandwidth communication.
  • Run NCCL/RCCL collective communication benchmarks, including AllReduce and AllGather.
  • Analyze CPU and multi-node communication regressions across guest operating systems, hypervisors, and physical fabrics.

Requirements

  • 5+ years of relevant experience and a bachelor's or master's degree in Computer Science, Electrical Engineering, or a related technical field.
  • Experience building and deploying automated integration tests for AI cloud environments, from low-level Linux systems to distributed control planes.
  • Advanced Python and/or Bash scripting skills and strong knowledge of CI/CD pipelines and GitLab tooling.
  • Working knowledge of Kubernetes, Docker, Terraform, and Postgres.
  • Familiarity with NVIDIA CUDA/NCCL and/or AMD ROCm/RCCL in multi-node environments.
  • Strong understanding of RDMA, RoCE, InfiniBand, Linux kernel internals, PCIe topology, VFIO, HugePages, and IOMMU.

Nice to have

  • Experience with MNNVL or specialized AI fabric architectures.
  • Familiarity with NVIDIA Nsight, AMD Omniperf, and hardware-level debugging or performance profiling tools.
  • Knowledge of Kubernetes GPU orchestration and specialized device plugins.

Culture & Benefits

  • Competitive compensation, equity, and Restricted Stock Units.
  • Paid time off, holidays, leave programs, and parental leave.
  • Health, dental, vision, HSA contributions, life insurance, and disability coverage.
  • Professional development, tuition reimbursement, and mental health support.
  • 401(k) plan with company match of up to 4% of salary.
  • Commuter benefits, cell phone stipend, daily meal allowance, and global travel insurance.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →