9 дней назад
Technology Enablement Engineer (AI Infrastructure)
129 600 - 190 067$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Technology Enablement Engineer (AI Infrastructure): Designing and operating production-grade GPU infrastructure for large-scale AI training and inference with an accent on Kubernetes clusters, distributed workloads, GPU scheduling, and performance engineering. Focus on building multi-tenant platforms, optimizing LLM workloads, automating cluster operations through GitOps, and debugging performance across compute, networking, memory, and storage layers.
Location: Ann Arbor, Michigan, United States; onsite
Salary: $129,600–$190,067 annually
Company
develops inspection tools, metrology systems, process solutions, and computational analytics used to manufacture advanced electronics.
What you will do
- Design and deploy scalable, multi-node GPU clusters on Kubernetes, including compute, networking, storage, scheduling, security, and observability.
- Build distributed training and reinforcement learning platforms with Ray, NVIDIA NeMo RL, PyTorch, and JAX.
- Deploy and optimize high-throughput LLM inference using vLLM, SGLang, and NVIDIA Dynamo.
- Implement GPU scheduling, quotas, isolation, autoscaling, health monitoring, and capacity management for multi-tenant environments.
- Profile and troubleshoot GPU workloads across compute, memory, networking, and storage layers using NVIDIA diagnostic tools.
- Automate provisioning, upgrades, deployments, and operational recovery with infrastructure-as-code and GitOps practices while contributing to open-source AI infrastructure.
Requirements
- Bachelor's degree and eight years of software engineering experience.
- At least four years in software, cloud, platform, HPC, or infrastructure engineering, including two years supporting distributed AI/ML workloads.
- Hands-on experience building GPU clusters from the ground up and operating them at production scale.
- Deep Kubernetes expertise, including operators, CRDs, Helm, networking, storage, scheduling, and cluster lifecycle management.
- Production experience with Ray and at least two of NVIDIA NeMo RL, vLLM, SGLang, or NVIDIA Dynamo.
- Strong knowledge of NVIDIA GPUs, CUDA, NCCL, GPU Operator, DCGM, MIG, RDMA, containers, CI/CD, GitOps, infrastructure-as-code, Python, and Linux.
Nice to have
- Experience with Google TPUs and TPU-oriented frameworks or distributed workloads.
- Experience with large language model training, fine-tuning, RLHF, or agentic reinforcement learning.
- Knowledge of TensorRT-LLM, Triton Inference Server, DeepSpeed, Megatron-LM, or similar performance-oriented frameworks.
- Experience operating secure, multi-tenant AI platforms in enterprise or regulated environments.
Culture & Benefits
- Full-time employment with medical, dental, vision, life, and other benefits.
- 401(k) with company matching and an employee stock purchase program.
- Paid time off, company holidays, and family care and bonding leave.
- Tuition reimbursement, student debt assistance, financial planning, wellness, and career development programs.
- Equal opportunity employment with reasonable accommodation available for qualified individuals with disabilities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Member of Technical Staff, Infrastructure (AI)
150 000 - 390 000$
10 дней назад
Staff Software Engineer (AI-Native Platform)
148 600 - 198 200$
11 дней назад
Senior Platform Engineer (AI)
228 000 - 279 000$
11 дней назад
Deployment Engineer (AI)
11 дней назад
Senior Platform Infrastructure Engineer (AI)
128 000 - 170 000$
11 дней назад
Staff Platform Engineer (AI/ML)
126 000 - 174 000$