Назад
Company hidden
4 часа назад

Member of Technical Staff — Inference-Multi-Hardware (AI Infrastructure)

200 000 - 400 000$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff — Inference-Multi-Hardware (AI Infrastructure): Bringing up and optimizing AI inference and training systems across NVIDIA and AMD GPUs, Google TPUs, server CPUs, and emerging accelerators with an accent on portable kernels, hardware abstractions, and distributed execution. Focus on porting runtimes to immature platforms, profiling memory and compiler behavior, debugging cross-platform correctness, and delivering production performance with vendor engineering teams.

Location: Palo Alto, California, United States

Salary: $200,000–$400,000 annual salary + equity

Company

hirify.global is an infrastructure-first AI company building open systems for large-scale inference and training, including SGLang and Miles.

What you will do

  • Bring up inference and training systems on new accelerator platforms and drive them to competitive performance.
  • Design hardware abstractions that keep a single codebase performant across vendors.
  • Port and optimize high-performance kernels across programming models and memory architectures.
  • Build cross-platform benchmarking, profiling, and regression detection systems.
  • Debug numerical divergence and correctness gaps between platforms.
  • Collaborate with vendor engineering, kernel, runtime, distributed systems, and product teams on performance improvements.

Requirements

  • 4+ years of experience in systems, performance, or ML infrastructure engineering.
  • Deep expertise in at least one accelerator programming model, such as CUDA, ROCm/HIP, Pallas/XLA, Triton, or a vendor SDK.
  • Strong understanding of accelerator architecture, including memory hierarchy, bandwidth limits, and occupancy.
  • Experience writing or optimizing high-performance ML kernels and working with distributed execution or communication libraries such as NCCL, RCCL, or MPI.
  • Proficiency in C++ and Python.
  • Strong system-level debugging and profiling skills, including on platforms with incomplete tooling.
  • Production track record of shipped performance work.

Nice to have

  • Experience bringing up ML workloads on new silicon or working across multiple vendor ecosystems.
  • Experience with compiler stacks, hardware abstraction layers, quantization, mixed precision, or portable kernel interfaces.
  • Experience with distributed inference or training frameworks, CPU inference optimization, or large-scale collective communication.
  • Open-source contributions to kernels, compilers, or ML systems, and collaboration with silicon vendors or cloud partners.

Culture & Benefits

  • Work on open inference and training infrastructure used for frontier-level AI systems.
  • Collaborate directly with NVIDIA, Google, frontier AI labs, and other technical partners.
  • Contribute hardware-specific optimizations, benchmarks, and portability improvements to open-source SGLang and Miles.
  • Equity is included in the compensation package.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →