9 часов назад
Automated Testing Engineer (GPU Infrastructure)
172 500 - 210 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Automated Testing Engineer (GPU Infrastructure): Validating large-scale, multi-node GPU clusters and building automated integration testing for high-performance AI and HPC workloads with an accent on CI/CD automation, distributed scaling, and interconnect fabric testing. Focus on developing Python or Go cluster orchestration, benchmarking NCCL/RCCL collective communication, and analyzing performance bottlenecks across Linux, hypervisor, and physical network layers.
Location: San Francisco or Sunnyvale, California, US; onsite
Salary: $172,500–$210,000 per year, plus Restricted Stock Units
Company
is an AI infrastructure company building vertically integrated energy, data center, and cloud systems for demanding AI workloads.
What you will do
- Build CI/CD platforms and automated integration testing for low-level infrastructure and distributed control planes.
- Design and execute validation tests across multi-node virtualized GPU clusters to verify scaling, stability, and multi-tenant isolation.
- Develop Python or Go automation frameworks to provision, configure, and stress-test cluster environments.
- Validate NVLink, Infinity Fabric, InfiniBand, and RoCE interconnects for low-latency, high-bandwidth communication.
- Run NCCL/RCCL collective communication benchmarks, including AllReduce and AllGather.
- Analyze CPU and multi-node communication regressions across guest operating systems, hypervisors, and physical fabrics.
Requirements
- 5+ years of relevant experience and a bachelor's or master's degree in Computer Science, Electrical Engineering, or a related technical field.
- Experience building and deploying automated integration tests for AI cloud environments, from low-level Linux systems to distributed control planes.
- Advanced Python and/or Bash scripting skills and strong knowledge of CI/CD pipelines and GitLab tooling.
- Working knowledge of Kubernetes, Docker, Terraform, and Postgres.
- Familiarity with NVIDIA CUDA/NCCL and/or AMD ROCm/RCCL in multi-node environments.
- Strong understanding of RDMA, RoCE, InfiniBand, Linux kernel internals, PCIe topology, VFIO, HugePages, and IOMMU.
Nice to have
- Experience with MNNVL or specialized AI fabric architectures.
- Familiarity with NVIDIA Nsight, AMD Omniperf, and hardware-level debugging or performance profiling tools.
- Knowledge of Kubernetes GPU orchestration and specialized device plugins.
Culture & Benefits
- Competitive compensation, equity, and Restricted Stock Units.
- Paid time off, holidays, leave programs, and parental leave.
- Health, dental, vision, HSA contributions, life insurance, and disability coverage.
- Professional development, tuition reimbursement, and mental health support.
- 401(k) plan with company match of up to 4% of salary.
- Commuter benefits, cell phone stipend, daily meal allowance, and global travel insurance.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Baseten
1 день назад
Software Engineer (Testing Infrastructure)
165 000 - 330 000$
2 дня назад
Staff QA Engineer (Cinema Imaging)
152 900 - 210 200$
5 часов назад
Senior Quality Assurance Engineer (AI)
220 000 - 240 000$
8 часов назад
Test Development Engineer (AI Hardware)
170 000 - 210 000$
8 часов назад
Test Engineer / Senior Test Engineer (Robotics)
148 000 - 181 000$
10 часов назад
Software Engineer (Cloud Security)
170 000 - 196 000$