2 часа назад
Reliability Tech Lead (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Reliability Tech Lead (AI): Building and operating reliability systems for a massive-scale, low-latency AI inference service with an accent on SLOs, fault tolerance, incident response, and observability. Focus on designing failover and graceful-degradation mechanisms, developing chaos-testing and fault-injection tooling, and improving reliability across multi-region cloud deployments and specialized data centers.
Location: US and Canada offices
Company
builds wafer-scale AI hardware and an inference platform designed to deliver high-speed model training and inference.
What you will do
- Define reliability strategy, SLOs, and incident-response frameworks for the inference service.
- Design and implement fault detection, graceful degradation, failover, throttling, and recovery mechanisms across regions and data centers.
- Lead incident management, postmortems, root-cause analysis, and prevention initiatives.
- Architect for redundancy, durability, observability, and debuggability across the inference stack.
- Develop tooling and frameworks for chaos testing, load simulation, and distributed fault injection.
- Partner with software, infrastructure, and hardware teams while monitoring reliability metrics and mentoring engineers.
Requirements
- Bachelor’s or master’s degree in computer science or a related field.
- 7+ years of experience in backend, infrastructure, or reliability engineering for large-scale distributed systems.
- Strong programming skills in at least one backend language, such as Python, C++, Go, or Rust.
- Deep experience with SLO, SLI, and SLA design, incident response, and postmortem practices.
- Excellent communication and cross-functional leadership skills.
Nice to have
- Experience building large-scale AI infrastructure systems.
Culture & Benefits
- Work on a breakthrough AI platform and one of the world’s fastest AI supercomputers.
- Opportunities to publish and open-source AI research.
- Job stability combined with startup vitality.
- Simple, non-corporate work culture that respects individual beliefs.
- Continuous learning, growth, and support in an inclusive work environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 часа назад
Senior Production Engineer (AI Infrastructure)
170 000 - 205 000$
6 часов назад
Senior Site Reliability Engineer (AI Infrastructure)
215 000 - 275 000$
5 часов назад
Production Engineer (AI Infrastructure)
Windsurf
6 дней назад
Site Reliability Engineer (AI)
2 дня назад
Senior Software / Site Reliability Lead Engineer (AI)
142 696 - 158 303$
6 часов назад