Назад
1 день назад

Software Engineer - Platform Infrastructure (Rust, C++)

180 000 - 440 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Software Engineer - Platform Infrastructure (Rust, C++): Designing and implementing a large-scale distributed system for one of the world's largest supercomputing clusters with an accent on low-level optimization and system reliability. Focus on profiling GPUs, Linux kernel, and networking to achieve peak efficiency in AI training.

Location: Palo Alto, CA

Salary: $180,000 - $440,000 USD

Company

xAI is dedicated to creating AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge.

What you will do

  • Design and implement a large-scale distributed system powering a global supercomputing cluster.
  • Optimize performance across the low-level stack, including GPUs, Linux kernel, networking, and filesystems.
  • Collaborate on hardware, software, and algorithm co-design to advance AI training.
  • Maintain and innovate the codebase to ensure maximum scalability and reliability.
  • Develop internal tools to enhance team productivity and streamline engineering workflows.

Requirements

  • Systems programming experience in C, C++, or Rust.
  • Deep understanding of computer systems fundamentals from transistors to high-level applications.
  • Hands-on expertise with Kubernetes (K8s), including cluster architecture, CNI, CSI, and service mesh.
  • Must be located in or able to work from Palo Alto, CA.

Nice to have

  • Full-stack debugging skills from the kernel/OS level up through container orchestration.
  • Deep knowledge of OS internals, including process scheduling, memory management, and synchronization.
  • Proficiency in low-level optimization and performance analysis techniques.
  • Experience with Linux kernel debugging tools such as perf, gdb, strace, and Wireshark.
  • Expertise in GitOps workflows, Helm, Operators, and container technologies (Docker, containerd, crio).
  • Experience with distributed system observability tools like Prometheus, Grafana, and OpenTelemetry.

Culture & Benefits

  • Flat organizational structure where initiative and excellence are rewarded.
  • High-performance environment focused on engineering excellence and curiosity.
  • Competitive compensation package including equity.
  • Comprehensive medical, vision, and dental coverage.
  • Access to a 401(k) retirement plan and life/disability insurance.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →