58 минут назад
Principal Engineer (AI Inference)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Engineer (AI Inference): Building and evolving the Inference Cloud Platform for high-availability, low-latency, multi-region AI inference with an accent on distributed systems architecture, active-active failover, and high-QPS performance. Focus on designing graceful degradation, solving production bottlenecks, and driving reliability, observability, and capacity improvements across teams.
Location: On-site at the Sunnyvale office, United States
Company
builds specialized AI hardware and cloud inference infrastructure designed to deliver high-speed AI model training and inference.
What you will do
- Define and prioritize high-leverage technical problems for the Inference Cloud Platform.
- Set long-term direction for multi-region topology, failure domains, service boundaries, and platform evolution.
- Architect active-active systems with rapid failover, circuit breaking, backpressure, load shedding, and clear SLOs.
- Improve latency, throughput, capacity efficiency, and resilience under unpredictable AI workload demand.
- Write and review production code and designs for critical paths, including build-versus-buy decisions.
- Lead production incidents, observability, capacity planning, post-incident improvements, cross-team technical strategy, and mentorship.
Requirements
- 10+ years of software engineering experience, including substantial individual contributor experience with large-scale distributed systems or cloud infrastructure.
- Deep expertise in distributed systems architecture, networking, compute orchestration, container platforms, and multi-region production services.
- Demonstrated experience building highly available, latency-sensitive systems at scale and optimizing high-QPS performance.
- Strong proficiency in Go, C++, or Python, with the ability to contribute production code directly.
- Experience with metrics, logging, tracing, alerting, incident response, and SLI/SLO/SLA-driven operations.
- Ability to influence senior engineers and cross-functional partners through technical judgment, communication, and credibility.
Nice to have
- Experience with TTFT and tail-latency reduction.
- Experience with ML inference infrastructure, model serving systems, or GPU-accelerated workloads.
Culture & Benefits
- Build AI platforms beyond the constraints of traditional GPUs.
- Opportunities to publish and open-source AI research.
- Work with a high-performance AI supercomputer platform.
- Startup vitality combined with job stability.
- Non-corporate culture focused on individual respect, continuous learning, and inclusion.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 часа назад
Software Engineer (Robotics/Cloud Infrastructure)
200 000 - 280 000$
3 часа назад
Software Engineer (AI Infrastructure)
170 000 - 205 000$
Databricks
6 дней назад
Senior Backend Engineer (AI)
165 300 - 219 675$
Baseten
2 дня назад
Software Engineer (Observability)
165 000 - 330 000$
6 часов назад
Member of Technical Staff (AI Infrastructure)
150 000 - 350 000$
5 часов назад
Senior Software Engineer (AI)
190 000 - 210 000$