Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Model Performance Engineer (AI): Building and operating Model APIs for high-performance hosted inference across distributed systems, model serving, and developer tooling with an accent on low-latency backend services, CUDA optimization, and multi-GPU performance. Focus on productionizing inference runtime improvements, designing benchmarking frameworks, and implementing observability, quotas, authentication, and usage metering.
Location: Hybrid in San Francisco, Montreal, New York, Seattle, or Toronto
Salary: $180K–$360K annually, plus equity
Company
Baseten provides mission-critical inference infrastructure, applied AI research, and developer tooling for companies deploying AI models in production.
What you will do
- Design, build, and operate Model APIs for structured outputs, tool and function calling, and multimodal serving.
- Profile and optimize TensorRT-LLM and CUDA kernels, implement custom CUDA operators, and improve memory allocation and multi-GPU communication.
- Productionize inference runtime improvements including speculative decoding, guided generation, quantization, batching, KV-cache reuse, scheduling, and routing.
- Build benchmarking frameworks covering model architectures, batch sizes, sequence lengths, and hardware configurations.
- Develop observability through metrics, traces, and logs while measuring speed, reliability, and quality.
- Implement API versioning, validation, usage metering, quotas, authentication, and developer-friendly serving experiences.
Requirements
- 3+ years of experience building and operating distributed systems or large-scale APIs.
- Experience owning low-latency, reliable backend services, including rate limiting, authentication, quotas, metering, and migrations.
- Strong performance and infrastructure skills, including profiling, tracing, capacity planning, and SLO management.
- Ability to debug performance and reliability issues across application behavior, runtimes, and infrastructure internals.
- Strong written communication and ability to create clear design documents and collaborate across functions.
Nice to have
- Experience with LLM runtimes such as vLLM, SGLang, or TensorRT-LLM, or contributions to open-source inference engines.
- Knowledge of Kubernetes, service meshes, API gateways, or distributed scheduling.
- Background in developer-facing infrastructure or open-source APIs.
Culture & Benefits
- Competitive compensation with meaningful equity.
- Flexible PTO and a company-wide Winter Break.
- Paid parental leave and a fertility and family-building stipend.
- U.S. employees and dependents receive 100% coverage of medical, dental, and vision insurance.
- U.S. employees have access to a company-facilitated 401(k).
- Exposure to a variety of ML startups and AI infrastructure work.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Writer
9 дней назад
Software Engineer (AI)
146 000 - 240 000$
7 дней назад
Software Engineer (AI)
210 000 - 265 000$
TypeSafe AI
4 дня назад
Member of Technical Staff Model Capabilities (AI)
150 000 - 250 000$
TypeSafe AI
4 дня назад
Member of Technical Staff Backend Platform (AI)
150 000 - 250 000$
Anthropic
4 дня назад
Staff+ Software Engineer, Inference Velocity (AI)
405 000 - 485 000$
7 дней назад
Head of Software Engineering (AI)
237 000 - 340 000$