Назад
Company hidden
5 часов назад

Software Engineer (AI)

175 000 - 220 000$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer (AI): Optimizing GPU kernels, distributed systems, and model execution for high-throughput training and inference across LLM, VLM, and video workloads with an accent on CUDA, Triton, mixed precision, quantization, and hardware efficiency. Focus on building low-latency sampling, distributed routing, model sharding, performance benchmarks, and communication optimizations across multi-GPU and multi-node environments.

Location: San Mateo, United States

Salary: $175,000–$220,000 per year, plus equity

Company

hirify.global provides an AI platform for building, training, and serving specialized models across text, image, embedding, audio, and multimodal workloads.

What you will do

  • Optimize system and GPU performance for high-throughput AI training and inference workloads.
  • Profile and resolve GPU-, kernel-, latency-, throughput-, memory-, and compute-efficiency bottlenecks.
  • Implement low-level optimizations with CUDA, Triton, PyTorch, and related performance tooling.
  • Scale LLM, VLM, and video-model systems across multi-GPU and multi-node environments.
  • Build benchmarking, monitoring, and performance-configuration infrastructure.
  • Collaborate with ML researchers and infrastructure teams on hardware-efficient model architectures, runtimes, and distributed systems.

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience.
  • 5+ years of experience in performance optimization or high-performance computing systems.
  • Proficiency in CUDA or ROCm and experience with GPU profiling tools such as Nsight, nvprof, or CUPTI.
  • Familiarity with PyTorch and performance-critical model execution.
  • Experience debugging and optimizing distributed systems in multi-GPU environments.
  • Deep understanding of GPU architecture, parallel programming models, and compute kernels.

Nice to have

  • Master’s or PhD in Computer Science, Electrical Engineering, or a related field.
  • Experience optimizing LLM, VLM, or video models for training and inference.
  • Knowledge of ML compiler stacks such as torch.compile, Triton, or XLA.
  • Open-source contributions to ML or HPC infrastructure.
  • Experience with cloud-scale AI infrastructure, Kubernetes, or hardware-aware model design.

Culture & Benefits

  • Work on low-latency inference, scalable model serving, and production AI infrastructure.
  • Collaborate with engineers and AI researchers on advanced generative AI systems.
  • High ownership and direct impact on system speed, scalability, and cost efficiency.
  • Equity is included in the compensation package.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →