14 часов назад
AI Data Platform Solutions Architect
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Data Platform Solutions Architect (NVIDIA AI Enterprise, Kubernetes, GPUs, vector databases): Supporting and designing HyperPOD AI Data Platform solutions across NVIDIA AI Enterprise services, vector databases, RAG and agentic workflows, high-performance storage, and networking with an accent on end-to-end triage, diagnostics, and production supportability. Focus on building reproducible labs, golden-stack patterns, unified diagnostic bundles, and observability for complex GPU-accelerated AI platforms.
Location: Remote - California
Company
develops enterprise and sovereign AI Data Platform solutions built on storage, NVIDIA AI Enterprise, NVIDIA GPUs, and Supermicro reference hardware.
What you will do
- Serve as the primary NVIDIA AI Enterprise and vector database expert for HyperPOD customer environments.
- Lead complex end-to-end triage across GPUs, NVIDIA AI Enterprise services, vector databases, Kubernetes, containers, networking, and Infinia storage.
- Diagnose and resolve performance issues in RAG and agentic AI workflows, including retrieval latency, GPU utilization, and data access patterns.
- Define unified diagnostic bundles and collaborate on Prometheus, Grafana, ELK, and NetQ observability dashboards.
- Build hands-on labs, proofs of concept, known-good configurations, and reusable implementation and troubleshooting assets.
- Provide structured compatibility, upgrade, rollback, and observability feedback to Product Management, Engineering, NVIDIA, OEM, and vendor partners.
Requirements
- 5+ years of experience in Linux-based infrastructure roles such as SRE, MLOps, platform engineering, or L2/L3 production support; 8+ years of total technical experience preferred.
- Strong hands-on experience with Docker or containerd, Kubernetes, Helm, Operators, pods, DaemonSets, CSI, CNI, and ingress or load balancers.
- Production experience with NVIDIA GPUs, drivers, CUDA concepts, GPU performance triage, GPU Operator, and GPU cluster platforms such as DGX or HGX.
- Experience with high-performance storage, RDMA or InfiniBand and high-speed Ethernet networking, HPC/AI fabrics, and cloud-adjacent Kubernetes patterns.
- Experience operating vector databases such as Milvus, Qdrant, Pinecone, pgVector, or vector search in OpenSearch or Elasticsearch.
- Understanding of RAG, generative AI, embeddings, retrieval, reranking, prompt design, context management, NVIDIA AI Enterprise components, and MLOps or GenAI pipelines.
Nice to have
- Experience with scale-out storage and RDMA-accelerated HPC/AI clusters at scale.
- Hands-on experience with NVIDIA reference blueprints for Enterprise RAG, VSS, AIQ, or similar enterprise AI architectures.
- Familiarity with AI observability, responsible AI practices, guardrails, model drift, toxicity monitoring, and GDPR or HIPAA considerations.
- Experience tuning Prometheus, Grafana, Loki, ELK, or NetQ for AI workloads, service-level dashboards, and SLOs.
Culture & Benefits
- Work remotely from California as part of the Global Support Services product support organization.
- Collaborate with NVIDIA solutions architects, OEM architects, Professional Services, Support Innovation, Product Management, and Engineering.
- Operate across enterprise AI, storage, networking, observability, and customer support domains.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Software Architect (Data Intelligence)
12 часов назад
Senior Solutions Architect (AI)
213 000 - 255 000$
14 часов назад
AI Solutions Architect (LLMs & AI Agents)
3 дня назад
Solutions Architect (Data Platforms)
14 часов назад
Principal Solutions Architect (AI)
314 300 - 44 900CAD
11 часов назад
Senior Solutions Architect (Cloud Data Storage)
250 000 - 275 000$