3 часа назад
Backend Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Backend Engineer (AI): Building and scaling Ollama’s cloud inference platform with an accent on high-throughput, low-latency distributed systems. Focus on GPU fleet management, multi-tenant infrastructure, and ensuring reliability for large-scale model serving.
Location: On-site in Palo Alto
Company
is the largest developer network in the open-model ecosystem, providing a local-first runtime for developers to access open models, backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.
What you will do
- Build and scale the inference platform serving requests for .com.
- Design routing and capacity layers to optimize workload placement across GPUs and regions.
- Manage multi-tenant infrastructure including isolation, quotas, metering, and billing.
- Develop reliability, observability, and cost control systems for the platform.
Requirements
- Deep experience with high-throughput, low-latency distributed systems.
- Proven ability to own production services end-to-end with a focus on cost/performance tradeoffs.
- Experience with Kubernetes, GPU scheduling, or inference infrastructure.
- Strong understanding of reliability, SLOs, and capacity planning.
Nice to have
- Experience building inference platforms or GPU fleet management.
- Background in billing or metering for AI services.
Culture & Benefits
- Small, talent-dense team with a flat, low-ego, and fast-moving culture.
- Focus on truth-seeking, passion, and design-driven development.
- Opportunity to work on infrastructure processing trillions of tokens.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →