4 дня назад
Staff AI Infrastructure Engineer
235 000 - 353 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff AI Infrastructure Engineer (GPU Infrastructure/AI): Architecting and operating Luma's 10k+ GPU fleet, including scheduling, resource management, reliability, and inference systems with an accent on Linux, distributed systems, Kubernetes, and low-level infrastructure behavior. Focus on debugging failures across hardware, kernels, runtimes, and orchestration, improving utilization under extreme demand, and building the systems required for new model capabilities.
Location: Redwood City, California, United States; hybrid
Base salary: $235,000–$353,000 annually, plus equity
Company
builds unified general intelligence that can generate, understand, and operate in the physical world, with a focus on multimodal intelligence and vision.
What you will do
- Architect and operate a heterogeneous fleet of more than 10,000 GPUs under extreme demand.
- Improve infrastructure utilization, performance, scheduling, placement, and resource management as cluster size and concurrency grow.
- Debug failures across hardware, operating systems, kernels, runtimes, containers, networking, storage, and orchestration.
- Partner with research and product teams to build infrastructure for new model capabilities and scale inference while maintaining reliability and latency.
- Set company-wide reliability standards, shape research and product architecture, and eliminate recurring classes of instability.
- Hire, develop, and mentor systems and reliability engineers.
Requirements
- Deep expertise in Linux and distributed systems.
- Production experience operating GPU or accelerator clusters.
- Strong proficiency with Kubernetes and modern open-source infrastructure.
- Ability to debug across hardware, kernels, runtimes, and orchestration, including behavior under contention and at scale.
- Ability to write code, build automation, and reason about bottlenecks, failure modes, and trade-offs.
- Strong technical judgment and production ownership in high-severity failure situations.
Nice to have
- Experience raising reliability standards across a company and influencing product and research architecture early.
- Ability to build strong engineering partnerships and attract and develop experienced engineers.
- Curiosity about how models use infrastructure and how systems improvements expand model capabilities.
Culture & Benefits
- Hybrid work in Redwood City, California.
- Equity offered in addition to base salary.
- Equal opportunity employment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →