Назад
Company hidden
10 дней назад

Software Engineer (Kernel Reliability)

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer (Kernel Reliability) (AI hardware and systems): Improving the reliability of advanced compute clusters and inference, training, and internal production services with an accent on kernel debugging, failure analysis, and scalable reliability tooling. Focus on diagnosing complex system failures, enhancing debug tools, collaborating on hardware-software architectures, and driving incident-response and root-cause analysis improvements.

Location: United States and Canada

Company

Builds large-scale AI compute hardware and software infrastructure designed to accelerate model training and inference beyond GPU-based systems.

What you will do

  • Contribute to the technical roadmap for kernel-centric reliability across internal and customer-facing systems.
  • Develop tooling and provide hands-on debugging support to reduce downtime after system and service failures.
  • Enhance diagnostic and debug tools to accelerate failure analysis.
  • Collaborate with software, ASIC, and hardware architecture teams on reliable, debuggable systems and next-generation architectures.
  • Participate in incident response, root-cause analysis, and post-mortems, driving follow-up improvements.

Requirements

  • Strong programming skills in C/C++ and Python.
  • Solid foundations in operating systems, computer architecture, and systems programming.
  • Ability to debug complex issues using logs, traces, and standard debugging workflows.
  • Interest in root-cause analysis; relevant skills may be demonstrated through projects, internships, or coursework.
  • New college graduates are welcome.

Nice to have

  • Exposure to parallel or distributed programming, including message passing, multicore, GPU, or embedded systems.
  • Experience with debuggers, core dumps, tracing, sanitizers, profilers, or other diagnostic tools.
  • Familiarity with deadlocks, livelocks, race conditions, and debugging parallel applications.
  • Knowledge of instruction pipelining, multithreading, networking, and memory systems.
  • Familiarity with monitoring, incident response, and post-mortem practices.

Culture & Benefits

  • Work on AI hardware and software infrastructure designed to overcome GPU limitations.
  • Opportunity to contribute to open-source AI research and work with high-performance AI systems.
  • Startup vitality combined with job stability.
  • Non-corporate culture that respects individual beliefs and supports continuous learning and growth.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →