5 часов назад
Supercomputing Platform & Infrastructure Engineer (AI)
200 000 - 550 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Supercomputing Platform & Infrastructure Engineer (AI): Designing and operating large-scale GPU clusters for model training and inference with an accent on Terraform-based infrastructure-as-code, Kubernetes orchestration, and high-throughput networking and storage. Focus on automating fault detection and recovery, debugging cross-layer failures, and improving observability and reliability across distributed infrastructure.
Location: On-site in San Francisco, with relocation support to SF where possible
Salary: $200,000–$550,000 annual base salary, depending on experience; equity is also included in total compensation.
Company
is building safe artificial general intelligence through research automation, code generation, frontier-scale pre-training, domain-specific reinforcement learning, long-context models, and inference-time compute.
What you will do
- Design, operate, and scale GPU clusters supporting model training and inference workloads.
- Build and maintain Terraform-driven infrastructure across cloud and hybrid environments.
- Deploy, operate, and optimize Kubernetes clusters for AI workload scheduling.
- Develop scalable infrastructure-as-code patterns for compute, networking, and storage provisioning.
- Optimize high-throughput networking and storage, deployment reproducibility, and environment consistency.
- Automate fault detection and recovery while improving observability and platform reliability.
Requirements
- Strong software engineering skills and experience building production infrastructure systems.
- Deep hands-on Terraform experience, including module design, state management, environment isolation, and large-scale deployments.
- Experience operating production GPU infrastructure or high-performance distributed systems.
- Strong understanding of networking and storage systems.
- Experience with major cloud platforms such as GCP, AWS, Azure, or OCI.
- Experience owning production-critical infrastructure end to end and debugging issues across hardware, drivers, networking, storage, operating systems, and cloud layers.
Culture & Benefits
- Small, fast-paced, highly focused team working toward safe AGI.
- Equity as a significant part of total compensation.
- 401(k) plan with 6% salary matching.
- Health, dental, and vision insurance for employees and dependents.
- Unlimited paid time off.
- Visa sponsorship and relocation stipend to San Francisco where possible.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Infrastructure Engineer (AI)
200 000 - 350 000$
7 часов назад
Site Reliability Engineer (AI)
100 000 - 300 000$
4 часа назад
Infrastructure Engineer (AI)
148 000 - 230 000$
6 часов назад
Infrastructure Engineer (AI)
200 000 - 400 000$
7 часов назад
Infrastructure Engineer (AI)
160 000 - 245 000$
4 часа назад
Infrastructure Platform Engineer (AI)
150 000 - 350 000$