2 дня назад
Principal Systems Software Engineer (AI Infrastructure)
260 000 - 340 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Systems Software Engineer (AI Infrastructure): Designing and productionizing a unified AI infrastructure fabric across BMaaS, IaaS, and CaaS with an accent on GPU throughput, virtualization, high-performance networking, and kernel-level systems. Focus on architecting RDMA and SR-IOV paths, optimizing memory and compute for large-scale AI workloads, and leading R&D from hyperscale prototypes to production systems.
Location: San Francisco, CA, US; on-site
Salary: $260,000–$340,000 annually, plus significant equity and bonus
Company
builds vertically integrated energy and AI infrastructure, operating systems from energy and data center layers through cloud services.
What you will do
- Architect a unified infrastructure fabric spanning Bare-Metal-as-a-Service, Intelligent IaaS, and Elastic CaaS for large-scale AI workloads.
- Design low-latency GPU infrastructure using InfiniBand, RDMA, SR-IOV, KVM, custom micro-VMs, Kubernetes, and Slurm.
- Lead R&D workstreams covering memory, networking, compute, virtualization, and GPU scheduling.
- Guide the technical roadmap and advise executive leadership on hardware/software co-design decisions.
- Debug complex I/O-path race conditions and optimize kernel-level memory pinning for GPU clusters.
- Produce technical white papers and RFCs and represent in open-source and industry communities.
Requirements
- 12+ years designing and shipping core infrastructure at a major hyperscaler or specialized HPC cloud.
- Deep expertise in the Linux kernel, KVM, QEMU, Firecracker, RoCE v2, and InfiniBand.
- Experience designing software that maximizes NVIDIA or AMD GPU and high-speed NIC performance.
- Experience leading cross-functional teams through ambiguous, high-impact infrastructure projects.
- Significant industry contributions through patents, open-source work, research, RFCs, or technical white papers.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related analytical field, or equivalent professional experience.
Nice to have
- Patents related to network virtualization, GPU scheduling, or distributed file systems.
- Maintainer status or major contributions to the Linux Kernel, Kubernetes, or specialized HPC projects.
- Experience optimizing infrastructure for large language model training and inference at scale.
- Peer-reviewed publications at systems venues such as OSDI, SOSP, NSDI, or SC.
Culture & Benefits
- Competitive compensation with significant equity and bonus opportunities.
- Restricted stock units and a 401(k) plan with company matching up to 4% of salary.
- Paid time off, paid holidays, parental leave, and volunteer time off.
- Health, dental, vision, HSA contributions, life insurance, and short- and long-term disability coverage.
- Professional development, tuition reimbursement, mental health support, commuter benefits, and a cell phone stipend.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Principal AI Systems Architect
195 000 - 285 000$
2 часа назад
Principal Engineer - Software Systems
150 000 - 200 000$
7 дней назад
Principal Engineer (Kubernetes)
174 051 - 227 233$
2 дня назад
Solutions Architect (AI Infrastructure)
200 000 - 280 000$
2 дня назад
Software Architect (AI)
204 000 - 245 000$
2 дня назад
Staff Software Engineer (Kubernetes)
215 000 - 265 000$