2 часа назад
Staff Engineer (AI Inference Platform)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Engineer (AI Inference Platform): Building and evolving the orchestration layer for a globally distributed inference platform with an accent on Kubernetes, high availability, security, and production reliability. Focus on designing active-active systems, reducing latency and tail latency, resolving complex production bottlenecks, and leading architecture across platform and infrastructure teams.
Location: On-site in Sunnyvale, United States, or Toronto, Canada
Company
builds large-scale AI hardware and infrastructure for high-speed model training and inference.
What you will do
- Design, develop, test, and maintain production software for the Inference Platform orchestration layer.
- Shape the technical direction for Kubernetes operators, custom resource definitions, failure domains, service boundaries, and platform evolution.
- Architect active-active systems with rapid failover, graceful degradation, clear SLOs, and improved latency, throughput, capacity efficiency, and resilience.
- Write and review production code, make high-consequence architectural decisions, and set engineering standards through design and code reviews.
- Lead complex production issues, observability, incident response, capacity planning, and post-incident improvements.
- Partner with ML, Product, Infrastructure, and Cloud teams on scalable system designs and cross-functional technical decisions.
Requirements
- 8+ years of software engineering experience, including substantial individual-contributor work on large-scale distributed systems or cloud infrastructure.
- Deep expertise in distributed systems architecture, ideally with Kubernetes.
- Experience designing highly available, latency-sensitive systems and optimizing latency, throughput, and efficiency in high-QPS environments.
- Experience with certificates, TLS, and mTLS.
- Strong proficiency in Go or C++ and the ability to contribute production code directly.
- Experience with metrics, logging, tracing, alerting, incident response, and SLO-driven operations, plus the ability to influence senior engineers and cross-functional partners.
Nice to have
- Experience with ML inference infrastructure, model serving systems, or GPU-accelerated workloads.
- Experience with TTFT and tail-latency reduction.
Culture & Benefits
- Opportunity to build an AI platform beyond the constraints of GPUs.
- Access to cutting-edge AI research, including publishing and open-source work.
- Work on a high-performance AI supercomputer platform.
- Combination of startup vitality and business stability.
- Non-corporate culture focused on individual beliefs, learning, growth, and inclusion.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 часов назад
Principal Software Engineer (Go)
225 000 - 270 000$
4 часа назад
Sr Staff Software Engineer (AI Infrastructure)
245 000 - 295 000$
7 дней назад
Software Engineer, Strategic Projects (AI)
18 202 - 25 237CAD
Amazon AI
23 часа назад
Senior Software Development Engineer (AI)
168 100 - 227 400$
7 часов назад
Senior Software Engineer (C++)
215 000 - 265 000$
5 дней назад
Sr Software Engineer (C++, AI/ML, OOP, Linux)
11 358 - 19 308$