6 дней назад
Senior Infrastructure Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Infrastructure Engineer (AI): Building and operating large-scale distributed systems for AI cluster management with an accent on Kubernetes operators, bare-metal automation, and observability. Focus on designing high-performance control planes, ensuring system reliability, and scaling infrastructure for wafer-scale AI hardware.
Location: Sunnyvale, CA (Hybrid)
Company
Systems builds breakthrough AI hardware, including the world's largest AI chip, to deliver industry-leading training and inference speeds for global enterprises and AI-native startups.
What you will do
- Develop declarative, CRD-driven automation for bare-metal networking, OS, and application software across large-scale clusters.
- Build and maintain Kubernetes operators to schedule complex inference workloads with priority queues and health-aware placement.
- Design gRPC control-plane services, including authorization, admission webhooks, and multi-tenant quota policies.
- Implement robust metrics and log pipelines using Prometheus and Grafana for wafer-scale systems and network fabric.
- Ensure system reliability through failure detection, HA control planes, and automated recovery mechanisms.
Requirements
- 5+ years of experience building and operating production distributed systems or infrastructure software.
- Proficiency in Go and Python for production-quality code.
- Deep expertise in Kubernetes, including writing controllers, operators, CRDs, and admission webhooks.
- Strong debugging skills across distributed systems, Linux, and networking.
- Practical experience with Prometheus and Grafana, including PromQL and exporter design.
- Demonstrated adoption of AI tools in engineering workflows with a focus on verification and rigor.
Nice to have
- Experience with bare-metal or HPC fleet operations.
- Knowledge of RDMA/RoCE, eBPF, Ceph, NVMe-oF, or etcd.
- Familiarity with scheduler internals and inference serving stacks.
Culture & Benefits
- Opportunity to work on one of the fastest AI supercomputers in the world.
- Non-corporate work culture that respects individual beliefs.
- Balance of startup vitality with long-term job stability.
- Support for continuous learning, growth, and open-source research contributions.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Writer
9 дней назад
Infrastructure Engineer (AI)
155 400 - 273 700$
9 дней назад
Member of Technical Staff, Infrastructure Engineer (AI)
175 000 - 240 000$
Perplexity
9 дней назад
Infrastructure Engineer (AI)
220 000 - 405 000$
11 дней назад
Senior DevOps Engineer (AI)
7 дней назад
Senior Systems Engineer (Cloud Infrastructure)
150 000 - 250 000$
10 дней назад