10 дней назад
Software Engineer (Kernel Reliability)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineer (Kernel Reliability) (AI hardware and systems): Improving the reliability of advanced compute clusters and inference, training, and internal production services with an accent on kernel debugging, failure analysis, and scalable reliability tooling. Focus on diagnosing complex system failures, enhancing debug tools, collaborating on hardware-software architectures, and driving incident-response and root-cause analysis improvements.
Location: United States and Canada
Company
Builds large-scale AI compute hardware and software infrastructure designed to accelerate model training and inference beyond GPU-based systems.
What you will do
- Contribute to the technical roadmap for kernel-centric reliability across internal and customer-facing systems.
- Develop tooling and provide hands-on debugging support to reduce downtime after system and service failures.
- Enhance diagnostic and debug tools to accelerate failure analysis.
- Collaborate with software, ASIC, and hardware architecture teams on reliable, debuggable systems and next-generation architectures.
- Participate in incident response, root-cause analysis, and post-mortems, driving follow-up improvements.
Requirements
- Strong programming skills in C/C++ and Python.
- Solid foundations in operating systems, computer architecture, and systems programming.
- Ability to debug complex issues using logs, traces, and standard debugging workflows.
- Interest in root-cause analysis; relevant skills may be demonstrated through projects, internships, or coursework.
- New college graduates are welcome.
Nice to have
- Exposure to parallel or distributed programming, including message passing, multicore, GPU, or embedded systems.
- Experience with debuggers, core dumps, tracing, sanitizers, profilers, or other diagnostic tools.
- Familiarity with deadlocks, livelocks, race conditions, and debugging parallel applications.
- Knowledge of instruction pipelining, multithreading, networking, and memory systems.
- Familiarity with monitoring, incident response, and post-mortem practices.
Culture & Benefits
- Work on AI hardware and software infrastructure designed to overcome GPU limitations.
- Opportunity to contribute to open-source AI research and work with high-performance AI systems.
- Startup vitality combined with job stability.
- Non-corporate culture that respects individual beliefs and supports continuous learning and growth.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
CoreWeave
11 дней назад
Principal Engineer, Distributed Systems (Security)
227 000 - 303 000$
TCOM
13 дней назад
Software Engineer (C/C++)
14 дней назад
Senior Software Engineer (Distributed Systems)
Relevance AI
13 дней назад
Staff Engineer (AI)
11 дней назад
Senior Software Engineer (AI/Code Security)
124 000 - 329 200$
11 дней назад
Software Engineer, C/C++
150 825 - 251 375$