Software Engineer (AI Inference)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Location: San Jose, United States; fully on-site in San Jose (Santana Row)
Salary: $175,000β$275,000 per year plus significant equity
Company
builds hardware, software, racks, and manufacturing systems for frontier AI inference, focusing on throughput and latency across prefill and decode workloads.
What you will do
- Port state-of-the-art models to 's accelerator architecture.
- Build programming abstractions and testing capabilities for rapid model porting.
- Develop and scale the runtime for multi-node inference, intra-node execution, state management, and error handling.
- Optimize routing and communication layers using collectives.
- Use performance profiling and debugging tools to identify bottlenecks and correctness issues.
Requirements
- Proficiency in C++ or Rust.
- Understanding of performance-sensitive or complex distributed software systems.
- Knowledge of Linux internals, accelerator architectures, compilers, or high-speed interconnects such as NVLink or InfiniBand.
- Familiarity with PyTorch or JAX.
- Experience porting applications to non-standard accelerator hardware or platforms.
- Ability to work fully in person in San Jose.
Nice to have
- Experience building low-latency, high-performance applications with kernel-level and user-space networking stacks.
- Deep understanding of distributed systems, including consensus protocols, consistency models, and communication patterns.
- Knowledge of Transformer architectures, especially Mixture-of-Experts.
- Experience with extensive SIMD optimizations for performance-critical paths.
Culture & Benefits
- Medical, dental, and vision coverage with generous premium support.
- $500 monthly credit for waiving medical benefits.
- $2,000 monthly housing subsidy for employees living within walking distance of the office.
- Relocation support for moving to San Jose.
- Wellness benefits, daily lunch and dinner at the office, and an ROI-justified unlimited compute budget.
- Fully in-person, cross-disciplinary engineering and research environment.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β