Назад
Company hidden
5 часов Π½Π°Π·Π°Π΄

Software Engineer (AI Inference)

175Β 000 - 275Β 000$
Π€ΠΎΡ€ΠΌΠ°Ρ‚ Ρ€Π°Π±ΠΎΡ‚Ρ‹
onsite
Π’ΠΈΠΏ Ρ€Π°Π±ΠΎΡ‚Ρ‹
fulltime
Английский
b2
Π‘Ρ‚Ρ€Π°Π½Π°
US
РСлокация
US
Вакансия ΠΈΠ· списка Hirify.GlobalВакансия ΠΈΠ· Hirify Global, списка ΠΌΠ΅ΠΆΠ΄ΡƒΠ½Π°Ρ€ΠΎΠ΄Π½Ρ‹Ρ… tech-ΠΊΠΎΠΌΠΏΠ°Π½ΠΈΠΉ
Для мэтча ΠΈ ΠΎΡ‚ΠΊΠ»ΠΈΠΊΠ° Π½ΡƒΠΆΠ΅Π½ Plus

ΠœΡΡ‚Ρ‡ & Π‘ΠΎΠΏΡ€ΠΎΠ²ΠΎΠ΄

Для мэтча с этой вакансиСй Π½ΡƒΠΆΠ΅Π½ Plus

ОписаниС вакансии

ВСкст:
/
TL;DR
Software Engineer (AI Inference): Building and scaling runtime systems for frontier AI inference hardware, including multi-node inference, intra-node execution, state management, and communication layers with an accent on performance, reliability, and model porting. Focus on optimizing distributed execution, profiling bottlenecks, and supporting state-of-the-art models on non-standard accelerator architectures.

Location: San Jose, United States; fully on-site in San Jose (Santana Row)

Salary: $175,000–$275,000 per year plus significant equity

Company

hirify.global builds hardware, software, racks, and manufacturing systems for frontier AI inference, focusing on throughput and latency across prefill and decode workloads.

What you will do

  • Port state-of-the-art models to hirify.global's accelerator architecture.
  • Build programming abstractions and testing capabilities for rapid model porting.
  • Develop and scale the runtime for multi-node inference, intra-node execution, state management, and error handling.
  • Optimize routing and communication layers using collectives.
  • Use performance profiling and debugging tools to identify bottlenecks and correctness issues.

Requirements

  • Proficiency in C++ or Rust.
  • Understanding of performance-sensitive or complex distributed software systems.
  • Knowledge of Linux internals, accelerator architectures, compilers, or high-speed interconnects such as NVLink or InfiniBand.
  • Familiarity with PyTorch or JAX.
  • Experience porting applications to non-standard accelerator hardware or platforms.
  • Ability to work fully in person in San Jose.

Nice to have

  • Experience building low-latency, high-performance applications with kernel-level and user-space networking stacks.
  • Deep understanding of distributed systems, including consensus protocols, consistency models, and communication patterns.
  • Knowledge of Transformer architectures, especially Mixture-of-Experts.
  • Experience with extensive SIMD optimizations for performance-critical paths.

Culture & Benefits

  • Medical, dental, and vision coverage with generous premium support.
  • $500 monthly credit for waiving medical benefits.
  • $2,000 monthly housing subsidy for employees living within walking distance of the office.
  • Relocation support for moving to San Jose.
  • Wellness benefits, daily lunch and dinner at the office, and an ROI-justified unlimited compute budget.
  • Fully in-person, cross-disciplinary engineering and research environment.

Π‘ΡƒΠ΄ΡŒΡ‚Π΅ остороТны: Ссли Ρ€Π°Π±ΠΎΡ‚ΠΎΠ΄Π°Ρ‚Π΅Π»ΡŒ просит Π²ΠΎΠΉΡ‚ΠΈ Π² ΠΈΡ… систСму, ΠΈΡΠΏΠΎΠ»ΡŒΠ·ΡƒΡ iCloud/Google, ΠΏΡ€ΠΈΡΠ»Π°Ρ‚ΡŒ ΠΊΠΎΠ΄/ΠΏΠ°Ρ€ΠΎΠ»ΡŒ, Π·Π°ΠΏΡƒΡΡ‚ΠΈΡ‚ΡŒ ΠΊΠΎΠ΄/ПО, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡ‚Π΅ этого - это мошСнники. ΠžΠ±ΡΠ·Π°Ρ‚Π΅Π»ΡŒΠ½ΠΎ ΠΆΠΌΠΈΡ‚Π΅ "ΠŸΠΎΠΆΠ°Π»ΠΎΠ²Π°Ρ‚ΡŒΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡˆΠΈΡ‚Π΅ Π² ΠΏΠΎΠ΄Π΄Π΅Ρ€ΠΆΠΊΡƒ. ΠŸΠΎΠ΄Ρ€ΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β†’