Назад
Company hidden
2 дня назад

Compute / Server Platform Architect (AI)

Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Compute / Server Platform Architect (AI): Owning the server-side platform architecture for Cerebras CS3-based AI training and inference clusters with an accent on CPU, memory, IO, PCIe, networking, and predictable system performance. Focus on building capacity and scaling models, validating hardware configurations through benchmarking, and driving qualification, vendor collaboration, and cross-stack debugging.

Location: US and Canada offices

Company

hirify.global builds AI accelerator hardware and CS3-based systems for high-speed model training and inference.

What you will do

  • Own the architecture, configurations, server roles, and lifecycle strategy for hirify.global AI clusters.
  • Define server formulas, capacity plans, scaling ratios, and headroom policies for different cluster sizes and workloads.
  • Specify CPU, memory, PCIe, NIC, NVMe, OS, BIOS, firmware, and driver configurations.
  • Translate software and runtime behavior into measurable hardware requirements and communicate technical guardrails to software teams.
  • Build performance and scaling models, run benchmarks and workload experiments, and drive cross-stack bottleneck resolution.
  • Lead vendor evaluations, platform qualification, production adoption, and root-cause analysis for hardware and software regressions.

Requirements

  • PhD in Computer Science or Electrical/Computer Engineering with 8+ years of industry experience, or a bachelor's/master's degree in CS or EE with 10+ years of industry experience.
  • 5+ years of experience in server platform architecture, systems performance engineering, or large-scale infrastructure design for AI/ML, HPC, or performance-sensitive distributed systems.
  • Deep knowledge of x86 server architecture, CPU microarchitecture, cache hierarchies, NUMA, memory controllers, and memory bandwidth and latency tradeoffs.
  • Strong Linux systems knowledge, including profiling, performance analysis, scheduling, syscall overhead, memory management, and tuning.
  • Experience with high-performance IO, NIC behavior, RDMA/RoCE, NVMe, capacity modeling, and rigorous benchmarking.
  • Ability to work with vendors and cross-functional teams, document tradeoffs, and drive technical decisions; familiarity with C, C++, and Python.

Culture & Benefits

  • Work on an AI platform designed to extend beyond GPU limitations.
  • Opportunity to contribute to AI research, model releases, and open-source projects.
  • Work with one of the fastest AI supercomputers in the world.
  • Combination of startup vitality and job stability.
  • Non-corporate culture focused on individual beliefs, learning, growth, and inclusion.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →