Назад
1 час назад

Principal Systems Architect (AI)

314 000 - 465 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Systems Architect (AI): Architecting and defining scalable compute platforms for AI/ML and high-throughput workloads with an accent on CPU/GPU architecture, memory hierarchies, and accelerator topologies. Focus on designing systems around high-bandwidth fabrics and optimizing for density, power, and performance.

Location: Must be based in or able to commute to San Jose, San Francisco, or Bellevue offices 4 days per week

Salary: $314,000 – $465,000 per year (varies by location)

Company

Lambda is a leader in AI cloud infrastructure, providing high-performance compute solutions to researchers and enterprises to make superintelligence ubiquitous.

What you will do

  • Architect and define scalable compute platforms optimized for AI/ML, simulation, and high-throughput workloads.
  • Develop compute system standards and design patterns to ensure consistency, performance, and maintainability.
  • Evaluate emerging CPU, GPU, and accelerator technologies to make architectural tradeoff decisions regarding density, power, and cost.
  • Collaborate with product and engineering teams to map workload requirements to platform capabilities.
  • Define compute platform roadmaps and architectural reference designs for hardware and cluster deployment.
  • Mentor systems engineers on performance tuning, sizing, and architectural strategy.

Requirements

  • 7+ years of experience architecting large-scale 10k-100k+ GPU HPC or cloud compute platforms.
  • Deep knowledge of CPU/GPU architectures, memory hierarchies, and accelerator topologies.
  • Experience designing systems around high-bandwidth, low-latency fabrics (NVLink, InfiniBand, and RoCE).
  • Strong understanding of system performance tuning, resource scheduling, and thermal/power optimization.
  • Ability to work across hardware and software boundaries, including OS behavior and orchestration layers.
  • Must be able to work onsite 4 days per week in San Jose, San Francisco, or Bellevue.

Nice to have

  • Hands-on experience with AI/ML workloads and performance characteristics.
  • Familiarity with HPC orchestration tools like Slurm or Kubernetes.
  • Experience with GPU virtualization technologies.
  • Background in compute telemetry and real-time performance profiling.

Culture & Benefits

  • Generous cash and equity compensation.
  • Comprehensive health, dental, and vision coverage for employees and dependents.
  • 401k Plan with 2% company match.
  • Flexible paid time off policy.
  • Wellness and commuter stipends for select roles.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →