6 дней назад
AI Engineer - Inference (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Engineer - Inference (AI infrastructure): Build and operate self-hosted inference services and scalable model endpoints for internal applications, customer products, and future Inference-as-a-service offerings with an accent on model serving, GPU performance, security, and operational reliability. Focus on designing reusable inference recipes, optimizing latency and throughput, benchmarking distributed deployments, and connecting workload profiles to scheduling, capacity planning, and AI-factory operations.
Location: Singapore or Australia (Launceston, Hobart, Sydney, Melbourne)
Company
Technologies develops and operates energy-efficient AI infrastructure and the AI Cloud GPU platform across Asia Pacific.
What you will do
- Build, operate, and improve self-hosted AI inference services for internal applications, customer-facing products, and future Inference-as-a-service offerings.
- Design model-onboarding workflows covering compatibility validation, packaging, runtime selection, optimization, deployment, endpoint registration, testing, release, and lifecycle management.
- Provision secure, scalable endpoints for generation, RAG, embeddings, reranking, batch, multimodal, tool-calling, and agentic workloads.
- Develop deployment templates, APIs, SDKs, configuration standards, and self-service workflows for model endpoint management.
- Optimize inference performance through quantization, compilation, batching, caching, routing, memory optimization, distributed parallelism, and GPU profiling.
- Build benchmarking, observability, regression-testing, capacity-management, and security capabilities while collaborating with Platform, DevOps, Infrastructure, Security, Product, and AI applications teams.
Requirements
- 5+ years of software engineering experience, including 3+ years in AI inference, model serving, ML systems, high-performance computing, distributed systems, or comparable performance-critical environments.
- Production experience with model-serving platforms, inference APIs, GPU-backed services, AI developer platforms, or multi-tenant AI systems.
- Hands-on experience with inference frameworks such as TensorRT-LLM, TensorRT, SGLang, vLLM, Triton Inference Server, NVIDIA Dynamo, NVIDIA NIM, or equivalent technologies.
- Strong knowledge of the NVIDIA AI stack, including CUDA, cuDNN, NCCL, TensorRT, GPU profiling, distributed communication, and GPU performance analysis.
- Strong Python skills and working proficiency in C++ or Go, plus experience with Kubernetes, containers, CI/CD, GitOps, APIs, autoscaling, observability, and production platform operations.
- Experience with distributed inference, GPU topology, benchmarking, performance profiling, security, governance, tenant isolation, quotas, rate limiting, and data protection.
Culture & Benefits
- Permanent full-time employment.
- Founder-led environment with accessible leaders, fast decision-making, and limited bureaucracy.
- Early ownership and opportunities to grow into new technical domains.
- Work alongside experts in AI infrastructure, energy systems, and next-generation compute.
- Contribute to sustainable AI factories designed to strengthen the energy grid and surrounding communities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →