Назад
11 месяцев назад

Model Performance Engineer (AI)

180 000 - 360 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Model Performance Engineer (AI): Building and operating Model APIs for high-performance hosted inference across distributed systems, model serving, and developer tooling with an accent on low-latency backend services, CUDA optimization, and multi-GPU performance. Focus on productionizing inference runtime improvements, designing benchmarking frameworks, and implementing observability, quotas, authentication, and usage metering.

Location: Hybrid in San Francisco, Montreal, New York, Seattle, or Toronto

Salary: $180K–$360K annually, plus equity

Company

Baseten provides mission-critical inference infrastructure, applied AI research, and developer tooling for companies deploying AI models in production.

What you will do

  • Design, build, and operate Model APIs for structured outputs, tool and function calling, and multimodal serving.
  • Profile and optimize TensorRT-LLM and CUDA kernels, implement custom CUDA operators, and improve memory allocation and multi-GPU communication.
  • Productionize inference runtime improvements including speculative decoding, guided generation, quantization, batching, KV-cache reuse, scheduling, and routing.
  • Build benchmarking frameworks covering model architectures, batch sizes, sequence lengths, and hardware configurations.
  • Develop observability through metrics, traces, and logs while measuring speed, reliability, and quality.
  • Implement API versioning, validation, usage metering, quotas, authentication, and developer-friendly serving experiences.

Requirements

  • 3+ years of experience building and operating distributed systems or large-scale APIs.
  • Experience owning low-latency, reliable backend services, including rate limiting, authentication, quotas, metering, and migrations.
  • Strong performance and infrastructure skills, including profiling, tracing, capacity planning, and SLO management.
  • Ability to debug performance and reliability issues across application behavior, runtimes, and infrastructure internals.
  • Strong written communication and ability to create clear design documents and collaborate across functions.

Nice to have

  • Experience with LLM runtimes such as vLLM, SGLang, or TensorRT-LLM, or contributions to open-source inference engines.
  • Knowledge of Kubernetes, service meshes, API gateways, or distributed scheduling.
  • Background in developer-facing infrastructure or open-source APIs.

Culture & Benefits

  • Competitive compensation with meaningful equity.
  • Flexible PTO and a company-wide Winter Break.
  • Paid parental leave and a fertility and family-building stipend.
  • U.S. employees and dependents receive 100% coverage of medical, dental, and vision insurance.
  • U.S. employees have access to a company-facilitated 401(k).
  • Exposure to a variety of ML startups and AI infrastructure work.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →