Назад
Company hidden
1 час назад

AI Infrastructure Architect

Тип работы
fulltime
Грейд
lead
Английский
b2
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Infrastructure Architect (AI Inference): Defining and scaling NeuReality’s NR-NEXUS next-generation AI inference platform with an accent on cloud-native architecture, distributed workloads, and production system design. Focus on profiling GenAI infrastructure, optimizing performance and scalability, and building reliable model-serving, observability, and deployment capabilities.

Company

hirify.global develops NR-NEXUS, a next-generation AI inference platform.

What you will do

  • Lead the software architecture and technical roadmap for NR-NEXUS.
  • Write system specifications and translate technical capabilities into product value.
  • Research AI infrastructure, SaaS platforms, model serving, and inference trends.
  • Define performance goals and lead profiling, benchmarking, and optimization for GenAI and distributed AI workloads.
  • Collaborate with engineering teams, customers, partners, and open-source communities on performance, compatibility, and adoption.
  • Mentor software engineers and provide technical leadership.

Requirements

  • 7+ years of software engineering experience, including 3+ years in software architecture or technical leadership.
  • Strong experience with Kubernetes-based platforms, cloud-native architecture, distributed systems, microservices, APIs, and automation.
  • Deep understanding of GenAI/LLM infrastructure and distributed workloads.
  • Experience designing management software or SaaS platforms for production systems.
  • Hands-on experience with observability, including monitoring, logging, alerting, and SLA/SLO tracking.
  • Experience with CI/CD, deployment automation, upgrades, rollback mechanisms, security, authentication, authorization, and customer data center integrations.

Nice to have

  • Experience with production AI inference clusters using GPUs, AI accelerators, or specialized compute infrastructure.
  • Knowledge of model-serving frameworks such as vLLM, Triton Inference Server, or TensorRT-LLM.
  • Experience with scheduling, load balancing, autoscaling, failover, cluster observability, and GPU/accelerator orchestration.
  • Familiarity with GPUDirect RDMA, NCCL, NVLink, UALink, Prometheus, Grafana, OpenTelemetry, Helm, Argo CD, Istio, KServe, or Kubeflow.
  • Experience deploying software in on-premises, edge, private-cloud, or hybrid environments.

Culture & Benefits

  • Work closely with engineering teams, customers, partners, and open-source communities.
  • Provide technical leadership and mentorship to software engineers.
  • Contribute to a next-generation AI inference platform and its ecosystem compatibility.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →