Назад
5 дней назад

Software Engineer (AI Inference)

165 000 - 330 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
middle/senior/lead
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer (AI Inference) (HPC/LLM): Building benchmarking, profiling, observability, and development tools for high-performance AI inference infrastructure with an accent on GPU systems, model evaluation, and runtime performance. Focus on automating performance testing, identifying latency-cost-quality trade-offs, and optimizing model serving across distributed compute and networking stacks.

Location: Hybrid in San Francisco, Montreal, New York, Seattle, or Toronto

Salary: $165,000–$330,000 per year, plus equity

Company

Baseten provides AI inference infrastructure, applied AI research, and developer tooling for companies bringing machine-learning models into production.

What you will do

  • Evaluate and automate LLM quality benchmarks and custom performance suites for long-context, KV-cache reuse, and disaggregated serving workloads.
  • Develop GPU-enabled internal development environments for high-performance model experimentation.
  • Build and contribute to open-source benchmarking and model-evaluation tools, including InferenceMAX and genai-bench.
  • Profile systems with PyTorch Profiler, NVIDIA Nsight Systems, and py-spy to identify compute and networking bottlenecks.
  • Create real-time dashboards, alerts, CI/CD performance tests, and release automation for model runtimes.
  • Develop optimization tools to identify the best latency, cost, and quality configuration for each model and workload.

Requirements

  • Mid-to-senior engineering experience with strong technical depth and communication skills.
  • Understanding of GPU memory subsystems, InfiniBand, and data movement across clusters.
  • Experience or strong interest in scripting, stress testing, and fuzz testing to identify system limits.
  • Curiosity about Transformer mathematics, FLOPs, and memory requirements.
  • Familiarity with Python and willingness to master the NVIDIA software stack.
  • Ability to drive cross-team initiatives, work through ambiguous requirements, and mentor engineers.

Nice to have

  • Familiarity with C++.

Culture & Benefits

  • High ownership and autonomy while leading a small team and building tools from scratch.
  • Opportunities to contribute to open-source projects and develop expertise in GPU orchestration and LLM inference.
  • Meaningful equity and competitive compensation.
  • Flexible PTO, including a company-wide winter break, and paid parental leave.
  • U.S. employees receive full medical, dental, and vision coverage, plus access to a company-facilitated 401(k).
  • Fertility and family-building support through Carrot.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →