Назад
56 минут назад

Distributed Systems Engineer (AI Inference)

180 000 - 360 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Distributed Systems Engineer (AI Inference): Building and operating the distributed runtime and Model APIs that power large-scale LLM inference with an accent on routing, autoscaling, Kubernetes orchestration, and advanced model-serving capabilities. Focus on designing low-latency production systems, debugging GPU and infrastructure reliability issues, and balancing performance, scalability, operational simplicity, and developer experience.

Location: Hybrid in San Francisco, Montreal, New York, Seattle, or Toronto

Salary: $180K–$360K annually, plus equity

Company

Baseten provides infrastructure, applied AI research, and developer tooling for deploying and operating production AI models and large-scale inference.

What you will do

  • Build infrastructure and orchestration systems for distributed LLM inference, including routing, autoscaling, scheduling, and runtime management.
  • Design, build, and operate Model APIs supporting structured outputs, tool and function calling, and multimodal serving.
  • Implement API versioning, validation, usage metering, quotas, and authentication.
  • Develop observability through metrics, traces, logs, benchmarks, testing, and release automation.
  • Debug and harden production systems across Kubernetes, distributed runtimes, networking, and GPU workloads.
  • Own projects end to end and collaborate with inference performance engineers to deliver optimizations to customers.

Requirements

  • Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 3+ years of experience building and operating distributed systems, backend infrastructure, or large-scale APIs.
  • Experience owning low-latency, reliable backend services with rate limiting, authentication, quotas, metering, and migrations.
  • Knowledge of profiling, tracing, capacity planning, SLO management, and production reliability practices.
  • Ability to debug performance and reliability issues across application, runtime, and infrastructure layers.
  • Strong developer-experience, written communication, and cross-functional collaboration skills.

Nice to have

  • Experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or Dynamo.
  • Deep Kubernetes experience, including operators and custom resources, plus service meshes or API gateways.
  • Experience with distributed scheduling, autoscaling, service orchestration, or production GPU workloads.
  • Background in developer-facing infrastructure, APIs, open-source infrastructure, or ML systems.
  • Familiarity with observability tooling, CI/CD systems, or release automation.

Culture & Benefits

  • Competitive compensation with meaningful equity.
  • Flexible PTO and a company-wide Winter Break.
  • Paid parental leave and a fertility and family-building stipend.
  • U.S.-only: medical, dental, and vision insurance fully covered for employees and dependents.
  • U.S.-only: company-facilitated 401(k).
  • Exposure to ML startups and opportunities for learning and networking.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →