Назад
Company hidden
7 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄

Principal AI Systems Performance Engineer (AI)

Π’ΠΈΠΏ Ρ€Π°Π±ΠΎΡ‚Ρ‹
fulltime
Π“Ρ€Π΅ΠΉΠ΄
senior
Английский
b2
Π‘Ρ‚Ρ€Π°Π½Π°
US
Вакансия ΠΈΠ· списка Hirify.GlobalВакансия ΠΈΠ· Hirify Global, списка ΠΌΠ΅ΠΆΠ΄ΡƒΠ½Π°Ρ€ΠΎΠ΄Π½Ρ‹Ρ… tech-ΠΊΠΎΠΌΠΏΠ°Π½ΠΈΠΉ
Для мэтча ΠΈ ΠΎΡ‚ΠΊΠ»ΠΈΠΊΠ° Π½ΡƒΠΆΠ΅Π½ Plus

ΠœΡΡ‚Ρ‡ & Π‘ΠΎΠΏΡ€ΠΎΠ²ΠΎΠ΄

Для мэтча с этой вакансиСй Π½ΡƒΠΆΠ΅Π½ Plus

ОписаниС вакансии

ВСкст:
/
TL;DR
Principal AI Systems Performance Engineer (AI): Optimizing and scaling foundation models on SambaNova’s reconfigurable dataflow platform with an accent on throughput, latency, compiler-runtime-hardware collaboration, and efficient inference. Focus on profiling performance bottlenecks, applying quantization and scheduling optimizations, and building scalable single-node and distributed AI systems.

Location: San Jose, California, United States

Company

hirify.global develops a full-stack generative AI platform combining its SN40L chip, software stack, and open-source foundation models for enterprise and government use.

What you will do

  • Bring up and optimize foundation models such as DeepSeek, Llama, and Qwen on the hirify.global platform.
  • Profile and improve model performance across compiler, runtime, and hardware layers.
  • Collaborate with machine learning, compiler, runtime, and hardware teams on co-designed AI applications.
  • Integrate advances in model architecture, quantization, scheduling, and memory optimization.
  • Develop scalable end-to-end inference solutions aligned with customer needs.
  • Identify bottlenecks and optimize dataflow and scheduling for single-node and distributed systems.

Requirements

  • Bachelor’s or higher degree in computer science, electrical engineering, applied mathematics, physics, statistics, or a related field.
  • 3+ years of experience in deep learning performance optimization, compiler/runtime/kernel optimization, hardware-software co-design, or systems performance tuning.
  • Proficiency in Python or C++, with strong foundations in algorithms, data structures, and numerical computing.
  • Experience with PyTorch, TensorFlow, or JAX.
  • Demonstrated ability to analyze and optimize real-world machine learning pipelines.

Nice to have

  • Experience with LLM or multimodal model training and inference.
  • Background in distributed training, continuous batching, and high-throughput inference systems.
  • Knowledge of quantization, graph optimization, kernel fusion, model partitioning, memory hierarchy, caching, and scheduling.
  • Experience with DeepSpeed, Megatron, vLLM, TensorRT, CUDA, Triton, OpenCL, cuDNN, or cuBLAS.
  • Publications or open-source contributions in ML systems or performance optimization.

Culture & Benefits

  • Full-time US employment includes base salary, equity, and benefits.
  • Medical insurance premiums are covered at 95% for employees and 77% for dependents.
  • Health Savings Account with employer contribution, plus dental, vision, disability, life, AD&D, and flexible spending account options.
  • Well-being benefits include Headspace, Gympass+, One Medical, counseling, and an Employee Assistance Program.
  • Equal Opportunity/Affirmative Action employment practices apply.

Π‘ΡƒΠ΄ΡŒΡ‚Π΅ остороТны: Ссли Ρ€Π°Π±ΠΎΡ‚ΠΎΠ΄Π°Ρ‚Π΅Π»ΡŒ просит Π²ΠΎΠΉΡ‚ΠΈ Π² ΠΈΡ… систСму, ΠΈΡΠΏΠΎΠ»ΡŒΠ·ΡƒΡ iCloud/Google, ΠΏΡ€ΠΈΡΠ»Π°Ρ‚ΡŒ ΠΊΠΎΠ΄/ΠΏΠ°Ρ€ΠΎΠ»ΡŒ, Π·Π°ΠΏΡƒΡΡ‚ΠΈΡ‚ΡŒ ΠΊΠΎΠ΄/ПО, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡ‚Π΅ этого - это мошСнники. ΠžΠ±ΡΠ·Π°Ρ‚Π΅Π»ΡŒΠ½ΠΎ ΠΆΠΌΠΈΡ‚Π΅ "ΠŸΠΎΠΆΠ°Π»ΠΎΠ²Π°Ρ‚ΡŒΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡˆΠΈΡ‚Π΅ Π² ΠΏΠΎΠ΄Π΄Π΅Ρ€ΠΆΠΊΡƒ. ΠŸΠΎΠ΄Ρ€ΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β†’