23 часа назад
Infrastructure Engineer (AI Hardware)
150 000 - 250 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Infrastructure Engineer (AI Hardware): Lead development of next-generation infrastructure tooling and hybrid HPC clusters to support AI ASIC simulation, synthesis, and CI workflows with an accent on scalable orchestration, observability, and workload migration. Focus on designing programmable infrastructure control planes, real-time telemetry systems, and fault injection frameworks to ensure reliability and performance at scale.
Location
Location: On-site in San Jose, United States
Relocation support available for candidates moving to San Jose.
Company
builds hardware and software infrastructure for frontier AI inference, focusing on high-performance ASICs and platform engineering.
What you will do
- Design and build orchestration layers for hybrid high-performance compute clusters supporting AI ASIC workflows.
- Develop programmable infrastructure control planes ensuring reproducibility and auditability.
- Create tools to enable engineers to leverage massive parallelism without complexity.
- Prototype workload migration strategies between on-premise and cloud environments balancing performance and cost.
- Implement real-time telemetry, tracing, and observability stacks with alerting and synthetic testing frameworks.
- Build fault injection and synthetic testing systems to validate infrastructure reliability under stress.
Requirements
- Must be located on-site in San Jose or willing to relocate.
- 8+ years experience in infrastructure engineering, systems programming, or backend development with performance and hardware interaction focus.
- Strong programming skills in Python, Go, Rust, and C++.
- Expertise in Linux, virtualization, containerization, and CI/CD pipelines.
- Experience with Infrastructure as Code tools like Terraform, Ansible, Puppet.
- Knowledge of observability tools such as Prometheus, Grafana, VictoriaMetrics, and distributed tracing.
Nice to have
- Experience with Bazel build system.
- Familiarity with ASIC development flows and EDA tools like Synopsys, Cadence, Verilator.
- Hands-on experience with AWS, GCP, Azure hybrid cloud deployments.
- Experience managing bare-metal servers, network hardware, and high-performance storage.
- Profiling and optimizing compute environments for latency and throughput.
Culture & Benefits
- Medical, dental, and vision insurance with generous coverage.
- $500 monthly credit for waiving medical benefits.
- $2,000 monthly housing subsidy for living near the office.
- Relocation support to San Jose.
- Wellness benefits including fitness and mental health.
- Daily lunch and dinner provided at the office.
- Unlimited compute budget with ROI justification.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Lambda
3 дня назад
Storage Engineer
267 000 - 356 000$
24 часа назад
Senior Cloud Support Engineer (AI)
125 000 - 145 000$
Lambda
6 дней назад
Software Engineer (Cloud Infrastructure)
206 000 - 275 000$
22 часа назад
Cloud Support Engineer
145 000 - 175 000$
1 день назад
Senior Infrastructure Engineer (GPU Infrastructure)
180 000 - 300 000$
19 часов назад
HPC Infrastructure Engineer
147 750 - 221 625$