Назад
Company hidden
2 дня назад

Inference Architecture Interns (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
trainee
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Inference Architecture Interns (AI) (AI accelerators and inference): Developing and optimizing compute architectures, runtimes, and model integrations for next-generation inference hardware with an accent on performance modeling, distributed execution, and hardware–software co-design. Focus on porting transformer models, optimizing routing and communication layers, profiling bottlenecks, and implementing high-performance Model Toolkit components.

Location: On-site at the office in San Jose, CA

Compensation: Paid 12-week internship

Company

hirify.global builds hardware, software, racks, and manufacturing systems for frontier AI inference.

What you will do

  • Port state-of-the-art models to the company’s inference architecture and build programming abstractions and testing capabilities.
  • Build and scale runtime systems for multi-node inference, intra-node execution, state management, and error handling.
  • Optimize routing and communication layers using collectives.
  • Use performance profiling and debugging tools to identify bottlenecks and correctness issues.
  • Co-design hardware instructions and model operations to maximize inference performance.
  • Implement high-performance software components for the Model Toolkit.

Requirements

  • Progress toward a Bachelor’s, Master’s, or PhD in computer science, computer engineering, applied mathematics, or a related field.
  • Proficiency in Python and C++.
  • Understanding of performance-sensitive or complex distributed software systems, Linux internals, accelerator architectures, compilers, or high-speed interconnects.
  • Experience porting applications to non-standard accelerator hardware or platforms.
  • Deep knowledge of transformer architectures and/or inference serving stacks such as vLLM or SGLang.

Nice to have

  • Proficiency in Rust.
  • Experience with low-latency, high-performance networking, distributed systems, SIMD optimization, PyTorch, or JAX.
  • Knowledge of Mixture-of-Experts transformer architectures or experience in math competitions.

Culture & Benefits

  • Fully in-person work with direct mentorship from experienced engineers and industry leaders.
  • Generous housing support for those relocating.
  • Daily lunch and dinner at the office.
  • Opportunity to work across engineering and research disciplines on frontier AI infrastructure.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →