Solution Architect - GPU & HPC
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
TL;DR
Solution Architect - GPU & HPC (GPU cloud): Own the technical sales cycle for GPU cloud opportunities by translating customer workload requirements into technically sound, commercially viable solution designs with an accent on HPC/AI workload architecture, multi-GPU/multi-node performance tuning, and delivery feasibility. Focus on building reference architectures and producing detailed proposals, RFP responses, and BoMs that can be committed and delivered to the required standard.
Location: UK-based (Customer-site travel required)
Company
builds Hyperstack, an AI cloud delivering on-demand and private GPU infrastructure for compute-intensive workloads.
What you will do
- Lead the technical sales cycle end-to-end: customer brief, architecture design, proposal, and delivery handover as primary technical authority for GPU cloud solutions.
- Engage with customers to capture workload requirements, technical constraints, and commercial objectives; produce detailed solution designs (architecture diagrams, network topology, storage configurations, GPU resource allocation models).
- Collaborate with Pre-Sales Engineering and infrastructure/network teams to validate delivery feasibility before commitments are made.
- Maintain a library of reference architectures and solution templates across AI/ML training, inference, HPC, and rendering workloads.
- Create high-quality technical proposals, RFP responses, and statements of work with realistic estimates and risk assessments.
- Define and maintain Bills of Materials (BoMs) for proposed solutions and feed customer requirements/competitive intelligence into engineering leadership.
Requirements
- Proven experience designing and delivering HPC or AI software stacks at scale, including workload profiling, scheduler configuration (SLURM/PBS or equivalent), MPI/NCCL tuning, and distributed training frameworks (PyTorch/JAX/DeepSpeed).
- Deep understanding of GPU software environments: CUDA, cuDNN, NCCL, driver stacks, and production tooling for reliable large-scale training and inference.
- Hands-on performance optimization for AI/HPC workloads across multi-GPU and multi-node setups, including bottleneck identification and tuning at both application and infrastructure layers.
- Strong knowledge of containerization and orchestration in HPC/AI contexts: Docker, Kubernetes, NVIDIA GPU Operator, and container-native workload management.
- Background in an OEM, hyperscaler, neo-cloud, or enterprise/research HPC environment with exposure to the full design-to-deployment lifecycle for GPU-accelerated workloads.
- Ability to produce clear technical documentation and architecture diagrams for both engineering and board-level audiences.
Nice to have
- Experience with large-scale cluster benchmarking (e.g., NCCL tests, MLPerf) across GPU generations and topologies.
- Exposure to MLOps tooling and AI platform layers (MLflow/W&B, Triton/vLLM, Kubeflow/Airflow).
- Familiarity with InfiniBand and high-performance networking for distributed training performance.
- Commercial awareness from contributing to BoMs, proposals, or RFP responses in a pre-sales/customer-facing technical role.
Culture & Benefits
- Competitive salary with an annual discretionary bonus scheme.
- Employee wellbeing benefits and 25 days of holiday plus public holidays.
- Flexible working with regular customer-site travel as part of the role.
- Ownership and autonomy with direct influence over technical win rate on strategic opportunities.
- Work on cutting-edge GPU cloud infrastructure for AI/ML and HPC workloads.
- Collaborative international culture focused on trust, transparency, and ownership.
Hiring process
- Not specified in the provided text.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β