Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Systems Engineer (AI Cloud): Developing and evolving a Kubernetes-native test framework to qualify host images and infrastructure on physical GPU hardware with an accent on low-level system validation and HPC integration. Focus on hardening the framework core in Rust, extending coverage to Slurm-on-Kubernetes, and implementing AI-native triage tools.
Location: Must be based in the US (Livingston, NJ; New York, NY; Sunnyvale, CA; or Bellevue, WA) and meet U.S. Government export control regulations (U.S. person/permanent resident)
Salary: $153,000 – $204,000
Company
CoreWeave is a specialized cloud provider delivering high-performance GPU infrastructure designed to accelerate AI breakthroughs for labs and global enterprises.
What you will do
- Own and evolve the Kubernetes-native test framework for qualifying host images on real hardware.
- Extend validation coverage to HPC, fabric verification, and Slurm-on-Kubernetes (SUNK).
- Optimize the CI pipeline for speed and reliability, focusing on boot consistency and eliminating flaky tests.
- Build end-to-end reporting paths, including structured results, storage, and triage dashboards.
- Implement AI-native testing using LLMs for log triage, failure classification, and regression detection.
- Collaborate with firmware, kernel, and imaging teams to embed testing into the release process.
Requirements
- 3+ years of experience building test infrastructure, systems software, or platform tooling at scale.
- Proficiency in Python and a strong systems background in Rust, Go, C, or C++.
- Experience operating in Kubernetes environments and managing containerized deployments.
- Deep Linux systems knowledge, including the boot chain, kernel, drivers, and low-level debugging.
- Must be a U.S. person or eligible for export control authorization (citizen, green card holder, etc.).
Nice to have
- Experience with Rust and Kubernetes-native workflow orchestration (e.g., Argo Workflows).
- HPC expertise including InfiniBand/RoCE and GPU/accelerator validation.
- Knowledge of Slurm or Slurm-on-Kubernetes (SUNK).
- Experience applying LLMs to test workflows for triage and failure detection.
- Contributions to open-source systems projects or Rust crates.
Culture & Benefits
- 100% company-paid medical, dental, and vision insurance.
- 401(k) with a generous employer match and Employee Stock Purchase Program (ESPP).
- Flexible PTO and comprehensive disability/life insurance.
- Catered daily lunch at office and data center locations.
- Support for family-forming and mental wellness through dedicated benefit programs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
8 дней назад
Senior Systems Engineer (Infrastructure)
CrowdStrike
6 дней назад
Engineer II - Virtualization (Remote)
100 000 - 145 000$
9 дней назад
HPC Windows Server Infrastructure Engineer (AI)
136 300 - 199 900$
13 дней назад
VP of Systems Engineering (AI Infrastructure) - West Coast
12 дней назад
Senior Systems Engineer (OpenStack)
100 000 - 215 000$
9 дней назад