4 часа назад
Software Engineer (AI Infrastructure)
170 000 - 205 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineer (AI Infrastructure): Developing diagnostic, observability, automation, and repair tooling for high-performance GPU compute clusters with an accent on distributed systems, reliability, and hardware operations. Focus on building AI agents for component-level diagnosis and remediation, validating GPU systems with PyTorch and NVIDIA NCCL, and maintaining fleet availability across data center environments.
Location: San Francisco, CA, United States; on-site
Salary: $170,000–$205,000 per year plus bonus
Company
builds vertically integrated energy and AI infrastructure, operating systems from power generation through cloud computing to support large-scale AI workloads.
What you will do
- Develop deep-level diagnostics and troubleshooting tools for hardware faults in GPU racks and high-density compute systems.
- Build automation and troubleshooting tooling for NVIDIA A100, H200, GB200, B200, and AMD 350X/355X GPU platforms.
- Develop AI agents for component-level diagnosis, remediation, critical-environment management, and hardware repair workflows.
- Create post-repair validation and testing tools using burn-in testing, PyTorch, and NVIDIA NCCL.
- Own deployment, monitoring, and operational support for tooling that improves GPU fleet availability and performance.
- Develop automation for facility power management and direct liquid-cooling hardware systems.
Requirements
- 4–6 years of software engineering experience.
- Strong programming skills in at least one of Go, Python, Java, or Rust.
- Expertise in distributed systems, reliability, and cloud platforms such as Kubernetes, infrastructure as code, and GCP.
- Ability to identify problems, develop scalable solutions, set technical direction, and deliver independently.
- Strong analytical, problem-solving, communication, and collaboration skills.
- Ability to work on-site in San Francisco, United States.
Nice to have
- Experience with Temporal and Kubernetes.
- Experience working directly with hardware vendors.
- Background in large-scale GPU fleet operations or hyperscale data center environments.
Culture & Benefits
- Industry-competitive compensation with restricted stock units and bonus eligibility.
- Health, vision, dental, HSA, life insurance, and disability coverage.
- 401(k) with a 100% employer match up to 4% of salary.
- Paid parental leave, generous paid time off, and holiday schedule.
- Tuition reimbursement, cell phone reimbursement, Calm subscription, and legal services.
- Company-paid commuter benefit of $50 per pay period.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Snowflake
23 часа назад
Backend Infrastructure Engineer (AI)
200 000 - 270 000$
Snowflake
1 час назад
Staff Software Engineer (AI-first)
160 000 - 230 000$
Lambda
6 дней назад
Staff Engineer (Kubernetes/AI)
314 000 - 465 000$
6 часов назад
Senior Software Engineer (AI)
190 000 - 210 000$
6 дней назад
API Platform Engineer (AI)
16 667 - 29 167$
Baseten
2 дня назад
Software Engineer (Observability)
165 000 - 330 000$