6 дней назад
Staff/Sr. ML Infrastructure / Platform Engineer (LLM)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff/Sr. ML Infrastructure / Platform Engineer (LLM): Building and operating a production-grade, GPU-accelerated LLM serving platform for enterprise AI products with an accent on Kubernetes operations, multi-GPU inference optimization, and infrastructure as code. Focus on tuning autoscaling for GPU cost and latency SLAs, managing model deployments across environments, and monitoring GPU utilization, cache occupancy, latency, and token costs.
Location: Taipei, Taiwan; onsite
Company
develops enterprise cybersecurity software and operates a global cloud security business with its global research base in Taiwan.
What you will do
- Design, build, and operate a production-grade, GPU-accelerated LLM serving platform for enterprise-scale AI products.
- Operate multi-model LLM serving infrastructure and optimize multi-GPU inference.
- Manage production Kubernetes clusters with NVIDIA GPU nodes, including driver setup and GPU node lifecycle.
- Develop Terraform and Terragrunt modules for AWS and GCP and package platform components and model deployments as Helm charts.
- Maintain Prometheus and Grafana monitoring and build dashboards for GPU utilization, KV cache occupancy, TTFT, ITL latency, and cost per token.
- Configure autoscaling and alerting for latency SLAs, SLA violations, and out-of-memory events.
Requirements
- Experience operating multi-model LLM serving infrastructure.
- Experience with production Kubernetes clusters and NVIDIA GPU infrastructure.
- Ability to manage GPU node lifecycle, including NVIDIA driver setup.
- Experience writing Terraform/Terragrunt modules for AWS or GCP and maintaining multi-environment configurations.
- Experience with Helm charts, Prometheus, and Grafana.
Nice to have
- Experience with LoRA or PEFT fine-tuning workflows.
- Experience using MLflow for experiment tracking, model registry, and automated adapter deployment.
- Experience building LoRA adapter CI/CD pipelines from training through registry and serving.
- Experience with SGLang or NVIDIA NIM and tuning continuous batching, KV cache management, or speculative decoding.
Culture & Benefits
- Work on enterprise-scale AI infrastructure supporting multiple products.
- Join a global cybersecurity company with research and development operations across five continents.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →