Назад
Company hidden
обновлено 7 часов назад

Senior Inference Engineer (AI)

185 000 - 300 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Senior Inference Engineer (AI): Building and optimizing inference pipelines for AI-driven video generation products with an accent on GPU parallelism and inference acceleration. Focus on implementing CUDA kernels, optimizing tensor/sequence parallelism, and reducing model latency for scalable deployment.

Location: On-site in Palo Alto, CA (3–5 days a week)

Salary: $185,000 – $300,000

Company

hirify.global is an AI startup creating seamless and intuitive video creation tools to empower creativity by breaking down technical barriers.

What you will do

  • Lead and implement advanced inference acceleration techniques, including attention optimization and quantization.
  • Optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency.
  • Develop high-performance computing kernels and distributed workloads using CUDA and NCCL.
  • Collaborate with research and engineering teams to bring videogen and large language models (LLMs) into production.
  • Drive rigorous code reviews and mentor fellow engineers on best practices in GPU programming.

Requirements

  • 5+ years of engineering experience with a track record in inference acceleration and model deployment at scale.
  • Expertise in quantization, attention acceleration, and deep learning compiler stacks.
  • Deep knowledge of GPU programming (CUDA, NCCL) and distributed inference parallelism.
  • Familiarity with video generation (videogen) models and large language models (LLMs).
  • Must be based in or able to work on-site in Palo Alto, CA.

Nice to have

  • Experience with high-throughput video or real-time streaming model deployment.
  • Familiarity with distributed training and optimization toolkits.
  • Contributions to open source projects in AI infrastructure or deep learning compilers.
  • Startup or rapid prototyping experience.

Culture & Benefits

  • Competitive AI industry salary and equity in a fast-growing startup.
  • Comprehensive health benefits and monthly stipends.
  • Regular company retreats.
  • Collaborative and supportive office culture based in Palo Alto.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →