Назад
Company hidden
14 часов назад

AI Data Platform Solutions Architect

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Data Platform Solutions Architect (NVIDIA AI Enterprise, Kubernetes, GPUs, vector databases): Supporting and designing HyperPOD AI Data Platform solutions across NVIDIA AI Enterprise services, vector databases, RAG and agentic workflows, high-performance storage, and networking with an accent on end-to-end triage, diagnostics, and production supportability. Focus on building reproducible labs, golden-stack patterns, unified diagnostic bundles, and observability for complex GPU-accelerated AI platforms.

Location: Remote - California

Company

hirify.global develops enterprise and sovereign AI Data Platform solutions built on hirify.global storage, NVIDIA AI Enterprise, NVIDIA GPUs, and Supermicro reference hardware.

What you will do

  • Serve as the primary NVIDIA AI Enterprise and vector database expert for HyperPOD customer environments.
  • Lead complex end-to-end triage across GPUs, NVIDIA AI Enterprise services, vector databases, Kubernetes, containers, networking, and Infinia storage.
  • Diagnose and resolve performance issues in RAG and agentic AI workflows, including retrieval latency, GPU utilization, and data access patterns.
  • Define unified diagnostic bundles and collaborate on Prometheus, Grafana, ELK, and NetQ observability dashboards.
  • Build hands-on labs, proofs of concept, known-good configurations, and reusable implementation and troubleshooting assets.
  • Provide structured compatibility, upgrade, rollback, and observability feedback to Product Management, Engineering, NVIDIA, OEM, and vendor partners.

Requirements

  • 5+ years of experience in Linux-based infrastructure roles such as SRE, MLOps, platform engineering, or L2/L3 production support; 8+ years of total technical experience preferred.
  • Strong hands-on experience with Docker or containerd, Kubernetes, Helm, Operators, pods, DaemonSets, CSI, CNI, and ingress or load balancers.
  • Production experience with NVIDIA GPUs, drivers, CUDA concepts, GPU performance triage, GPU Operator, and GPU cluster platforms such as DGX or HGX.
  • Experience with high-performance storage, RDMA or InfiniBand and high-speed Ethernet networking, HPC/AI fabrics, and cloud-adjacent Kubernetes patterns.
  • Experience operating vector databases such as Milvus, Qdrant, Pinecone, pgVector, or vector search in OpenSearch or Elasticsearch.
  • Understanding of RAG, generative AI, embeddings, retrieval, reranking, prompt design, context management, NVIDIA AI Enterprise components, and MLOps or GenAI pipelines.

Nice to have

  • Experience with scale-out storage and RDMA-accelerated HPC/AI clusters at scale.
  • Hands-on experience with NVIDIA reference blueprints for Enterprise RAG, VSS, AIQ, or similar enterprise AI architectures.
  • Familiarity with AI observability, responsible AI practices, guardrails, model drift, toxicity monitoring, and GDPR or HIPAA considerations.
  • Experience tuning Prometheus, Grafana, Loki, ELK, or NetQ for AI workloads, service-level dashboards, and SLOs.

Culture & Benefits

  • Work remotely from California as part of the Global Support Services product support organization.
  • Collaborate with NVIDIA solutions architects, OEM architects, Professional Services, Support Innovation, Product Management, and Engineering.
  • Operate across enterprise AI, storage, networking, observability, and customer support domains.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →