Назад
Company hidden
6 часов назад

Staff Hardware Systems Engineer (AI Infrastructure)

215 000 - 260 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Hardware Systems Engineer (AI Infrastructure): Driving the lifecycle of GPU- and CPU-based compute platforms from prototype bring-up through production, with an accent on validation, performance characterization, distributed AI workloads, and system reliability. Focus on analyzing hardware/software bottlenecks, debugging compute and infrastructure systems, tuning cluster configurations, and building automation for large-scale AI and HPC environments.

Location: On-site in San Francisco or Sunnyvale, California, United States

Salary: $215,000–$260,000 per year, plus Restricted Stock Units

Company

hirify.global builds vertically integrated, energy-first AI infrastructure spanning energy, data centers, cloud services, and compute systems.

What you will do

  • Drive the lifecycle of next-generation compute platforms from evaluation and prototype bring-up through validation, deployment, and production readiness.
  • Define and execute performance characterization and validation strategies for CPU, GPU, and accelerated computing platforms.
  • Analyze training and inference workloads, including dense, MoE, long-context, and multimodal models, to understand compute, memory, communication, and I/O behavior.
  • Translate workload and platform insights into cluster tuning recommendations covering topology, parallelism, scheduling, power, and software stack configuration.
  • Lead system-level debugging across compute, memory, storage, networking, accelerators, and platform firmware.
  • Partner with hardware, firmware, software, infrastructure, reliability, operations, and vendor engineering teams on prototyping, qualification, NPI, and production support.

Requirements

  • 8+ years of experience in hardware systems, platform, performance, ML systems, infrastructure engineering, or a related field.
  • Hands-on experience with large-scale GPU or accelerated computing infrastructure and distributed training or inference workloads.
  • Experience with benchmarking, profiling, performance optimization, system bring-up, validation, and root-cause analysis.
  • Strong understanding of server and accelerator architectures, including CPU, GPU, memory, storage, networking, PCIe, InfiniBand, or NVLink.
  • Experience developing automation, testing, diagnostics, or data-analysis frameworks using Python, Shell, or similar languages.
  • Bachelor’s or Master’s degree in Electrical Engineering, Computer Engineering, Computer Science, or equivalent experience.

Nice to have

  • Experience with RDMA, RoCE, CXL, NVLink, fabric-level performance analysis, inference serving, training frameworks, or ML compiler/runtime stacks.
  • Familiarity with x86 and ARM server platforms and fleet-level observability, diagnostics, performance, or reliability systems.
  • Experience introducing compute technologies into production cloud or large-scale data center environments.
  • Knowledge of infrastructure efficiency, power, cooling, performance-per-dollar, total cost of ownership, or sustainable hardware design.

Culture & Benefits

  • Health, vision, and dental insurance options, including HDHP and PPO plans.
  • Employer HSA contributions, paid parental leave, life insurance, disability coverage, and Teladoc.
  • 401(k) with a 100% employer match up to 4% of salary.
  • Generous paid time off, holidays, commuter benefits, cell phone reimbursement, and tuition reimbursement.
  • Restricted Stock Units, Calm subscription, and MetLife Legal benefits.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →