Назад
Company hidden
8 часов назад

Senior Hardware Systems Engineer (AI Infrastructure)

170 000 - 205 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Hardware Systems Engineer (AI Infrastructure): Driving the lifecycle of CPU-, GPU-, and accelerated-computing platforms from prototype bring-up through production with an accent on validation, workload characterization, performance tuning, and reliability. Focus on analyzing training and inference workloads, solving complex hardware/software bottlenecks, and optimizing cluster configurations across compute, memory, networking, storage, and accelerators.

Location: Sunnyvale, California, United States; on-site

Salary: $170,000–$205,000 per year, plus Restricted Stock Units

Company

hirify.global is a vertically integrated AI infrastructure company building and operating energy-first compute systems for large-scale AI workloads.

What you will do

  • Drive the lifecycle of next-generation compute platforms from evaluation and prototype bring-up through validation, deployment, and production readiness.
  • Define and execute performance characterization and validation strategies for CPU, GPU, and accelerated-computing platforms.
  • Characterize training and inference workloads, including dense, MoE, long-context, and multimodal models, across compute, memory, communication, and I/O behavior.
  • Translate workload and platform findings into cluster-level recommendations covering topology, parallelism, scheduling, power, and software-stack configuration.
  • Build performance profiles and reference configurations for deploying, tuning, and scaling clusters across model families and workload classes.
  • Lead system-level debugging and root-cause analysis across compute, memory, storage, networking, accelerators, firmware, and software.

Requirements

  • 5–6+ years of experience in hardware systems, platform, performance, ML systems, infrastructure engineering, or a related field.
  • Hands-on experience with large-scale GPU or accelerated-computing infrastructure and distributed training or inference workloads.
  • Experience with benchmarking, performance profiling, system optimization, bring-up, validation, and complex hardware/software issue resolution.
  • Understanding of server and accelerator architectures, including CPU, GPU, memory, storage, networking, PCIe, InfiniBand, and NVLink.
  • Experience developing automation, testing, diagnostics, or data-analysis frameworks using Python, Shell, or similar languages.
  • Bachelor’s or Master’s degree in Electrical Engineering, Computer Engineering, Computer Science, or equivalent experience.

Nice to have

  • Experience with RDMA, RoCE, CXL, NVLink, fabric-level performance analysis, inference serving, training frameworks, or ML compiler/runtime stacks.
  • Familiarity with x86 and ARM server platforms and fleet-level observability, diagnostics, performance, or reliability systems.
  • Experience introducing compute technologies into production cloud or large-scale data-center environments.
  • Knowledge of infrastructure efficiency, power, cooling, performance-per-dollar, total cost of ownership, or sustainable hardware design.

Culture & Benefits

  • Collaboration across hardware, firmware, software, networking, infrastructure, reliability, operations, vendor, and customer teams.
  • Health, vision, and dental insurance options, including HDHP and PPO plans, plus employer HSA contributions.
  • 401(k) with a 100% employer match up to 4% of salary.
  • Paid parental leave, paid time off, holidays, life insurance, and short- and long-term disability coverage.
  • Tuition reimbursement, cell phone reimbursement, Teladoc, Calm subscription, legal services, and a company-paid commuter benefit of $300 per month.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →