Назад
Company hidden
5 дней назад

Machine Learning System Engineer (AI)

147 906 - 232 650$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning System Engineer (AI): Designing and optimizing large-scale model-serving systems with an accent on distributed infrastructure, GPU inference, and production reliability. Focus on building high-concurrency serving systems, accelerating inference through kernels and quantization, and deploying open-source LLMs with robust CI/CD.

Location: Remote in the United States, with listed locations in Seattle and San Francisco.

Base salary: $147,906–$193,100 for Zone C, $160,380–$209,385 for Zone B, or $178,200–$232,650 for Zone A.

Company

hirify.global develops software products that help teams collaborate and manage different types of work.

What you will do

  • Design and implement scalable distributed infrastructure for model serving, including load balancing, auto-scaling, batch scheduling, and global KV caching.
  • Optimize model inference latency and throughput under production workloads.
  • Build reliable, high-concurrency serving systems handling billions of requests.
  • Benchmark, fine-tune, and accelerate inference engines using low-level and algorithmic optimizations.
  • Create CI/CD infrastructure for model deployment and inference engine updates.
  • Partner with senior ML engineers to fine-tune and deploy open-source LLMs.

Requirements

  • 3+ years of software engineering experience.
  • Deep low-level systems programming experience with C, C++, or Rust.
  • Experience with large-scale, high-concurrency production serving.
  • Experience with GPU inference engines such as vLLM, SGLang, Triton, or TensorRT-LLM.

Nice to have

  • 1+ years of system performance optimization experience.
  • Experience with GPU kernels, quantization, speculative decoding, or distillation.
  • Experience testing, benchmarking, and improving inference service reliability.
  • Experience designing CI/CD infrastructure for inference.
  • Strong background in batching, caching, load balancing, and parallelism.

Culture & Benefits

  • Choice of working remotely or from an office in the United States.
  • Health and wellbeing resources.
  • Paid volunteer days.
  • Potential eligibility for benefits, bonuses, commissions, and equity.
  • Accommodations and adjustments are available during the recruitment process.

Hiring process

  • Identity verification may be required as a condition of employment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →