7 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Principal AI Systems Performance Engineer (AI)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Π’Π΅ΠΊΡΡ:
TL;DR
Principal AI Systems Performance Engineer (AI): Optimizing and scaling foundation models on SambaNovaβs reconfigurable dataflow platform with an accent on throughput, latency, compiler-runtime-hardware collaboration, and efficient inference. Focus on profiling performance bottlenecks, applying quantization and scheduling optimizations, and building scalable single-node and distributed AI systems.
Location: San Jose, California, United States
Company
develops a full-stack generative AI platform combining its SN40L chip, software stack, and open-source foundation models for enterprise and government use.
What you will do
- Bring up and optimize foundation models such as DeepSeek, Llama, and Qwen on the platform.
- Profile and improve model performance across compiler, runtime, and hardware layers.
- Collaborate with machine learning, compiler, runtime, and hardware teams on co-designed AI applications.
- Integrate advances in model architecture, quantization, scheduling, and memory optimization.
- Develop scalable end-to-end inference solutions aligned with customer needs.
- Identify bottlenecks and optimize dataflow and scheduling for single-node and distributed systems.
Requirements
- Bachelorβs or higher degree in computer science, electrical engineering, applied mathematics, physics, statistics, or a related field.
- 3+ years of experience in deep learning performance optimization, compiler/runtime/kernel optimization, hardware-software co-design, or systems performance tuning.
- Proficiency in Python or C++, with strong foundations in algorithms, data structures, and numerical computing.
- Experience with PyTorch, TensorFlow, or JAX.
- Demonstrated ability to analyze and optimize real-world machine learning pipelines.
Nice to have
- Experience with LLM or multimodal model training and inference.
- Background in distributed training, continuous batching, and high-throughput inference systems.
- Knowledge of quantization, graph optimization, kernel fusion, model partitioning, memory hierarchy, caching, and scheduling.
- Experience with DeepSpeed, Megatron, vLLM, TensorRT, CUDA, Triton, OpenCL, cuDNN, or cuBLAS.
- Publications or open-source contributions in ML systems or performance optimization.
Culture & Benefits
- Full-time US employment includes base salary, equity, and benefits.
- Medical insurance premiums are covered at 95% for employees and 77% for dependents.
- Health Savings Account with employer contribution, plus dental, vision, disability, life, AD&D, and flexible spending account options.
- Well-being benefits include Headspace, Gympass+, One Medical, counseling, and an Employee Assistance Program.
- Equal Opportunity/Affirmative Action employment practices apply.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β
ΠΠΎΡ ΠΎΠΆΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
3 Π΄Π½Ρ Π½Π°Π·Π°Π΄
AI Engineer (World Models)
3 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Software Engineering - Distributed Agentic AI Systems (AI)
3 Π΄Π½Ρ Π½Π°Π·Π°Π΄
ML Platform Engineer (AI)
150Β 000 - 350Β 000$
3 Π΄Π½Ρ Π½Π°Π·Π°Π΄
AI Engineer
5 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Principal Software Engineer (AI)
160Β 200 - 425Β 000$
3 Π΄Π½Ρ Π½Π°Π·Π°Π΄