1 час назад
AI Infrastructure Architect
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Infrastructure Architect (AI Inference): Defining and scaling NeuReality’s NR-NEXUS next-generation AI inference platform with an accent on cloud-native architecture, distributed workloads, and production system design. Focus on profiling GenAI infrastructure, optimizing performance and scalability, and building reliable model-serving, observability, and deployment capabilities.
Company
develops NR-NEXUS, a next-generation AI inference platform.
What you will do
- Lead the software architecture and technical roadmap for NR-NEXUS.
- Write system specifications and translate technical capabilities into product value.
- Research AI infrastructure, SaaS platforms, model serving, and inference trends.
- Define performance goals and lead profiling, benchmarking, and optimization for GenAI and distributed AI workloads.
- Collaborate with engineering teams, customers, partners, and open-source communities on performance, compatibility, and adoption.
- Mentor software engineers and provide technical leadership.
Requirements
- 7+ years of software engineering experience, including 3+ years in software architecture or technical leadership.
- Strong experience with Kubernetes-based platforms, cloud-native architecture, distributed systems, microservices, APIs, and automation.
- Deep understanding of GenAI/LLM infrastructure and distributed workloads.
- Experience designing management software or SaaS platforms for production systems.
- Hands-on experience with observability, including monitoring, logging, alerting, and SLA/SLO tracking.
- Experience with CI/CD, deployment automation, upgrades, rollback mechanisms, security, authentication, authorization, and customer data center integrations.
Nice to have
- Experience with production AI inference clusters using GPUs, AI accelerators, or specialized compute infrastructure.
- Knowledge of model-serving frameworks such as vLLM, Triton Inference Server, or TensorRT-LLM.
- Experience with scheduling, load balancing, autoscaling, failover, cluster observability, and GPU/accelerator orchestration.
- Familiarity with GPUDirect RDMA, NCCL, NVLink, UALink, Prometheus, Grafana, OpenTelemetry, Helm, Argo CD, Istio, KServe, or Kubeflow.
- Experience deploying software in on-premises, edge, private-cloud, or hybrid environments.
Culture & Benefits
- Work closely with engineering teams, customers, partners, and open-source communities.
- Provide technical leadership and mentorship to software engineers.
- Contribute to a next-generation AI inference platform and its ecosystem compatibility.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
11 часов назад
Distinguished Engineer (AI Infrastructure)
6 дней назад
Principal Engineer (AI)
14 часов назад
Software Architect (Cloud Infrastructure)
10 часов назад
Senior Principal Engineer (AI Infrastructure)
16 часов назад
Software Architect (AI Cybersecurity)
4 дня назад