Назад
Company hidden
9 дней назад

Technology Enablement Engineer (AI Infrastructure)

129 600 - 190 067$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Technology Enablement Engineer (AI Infrastructure): Designing and operating production-grade GPU infrastructure for large-scale AI training and inference with an accent on Kubernetes clusters, distributed workloads, GPU scheduling, and performance engineering. Focus on building multi-tenant platforms, optimizing LLM workloads, automating cluster operations through GitOps, and debugging performance across compute, networking, memory, and storage layers.

Location: Ann Arbor, Michigan, United States; onsite

Salary: $129,600–$190,067 annually

Company

hirify.global develops inspection tools, metrology systems, process solutions, and computational analytics used to manufacture advanced electronics.

What you will do

  • Design and deploy scalable, multi-node GPU clusters on Kubernetes, including compute, networking, storage, scheduling, security, and observability.
  • Build distributed training and reinforcement learning platforms with Ray, NVIDIA NeMo RL, PyTorch, and JAX.
  • Deploy and optimize high-throughput LLM inference using vLLM, SGLang, and NVIDIA Dynamo.
  • Implement GPU scheduling, quotas, isolation, autoscaling, health monitoring, and capacity management for multi-tenant environments.
  • Profile and troubleshoot GPU workloads across compute, memory, networking, and storage layers using NVIDIA diagnostic tools.
  • Automate provisioning, upgrades, deployments, and operational recovery with infrastructure-as-code and GitOps practices while contributing to open-source AI infrastructure.

Requirements

  • Bachelor's degree and eight years of software engineering experience.
  • At least four years in software, cloud, platform, HPC, or infrastructure engineering, including two years supporting distributed AI/ML workloads.
  • Hands-on experience building GPU clusters from the ground up and operating them at production scale.
  • Deep Kubernetes expertise, including operators, CRDs, Helm, networking, storage, scheduling, and cluster lifecycle management.
  • Production experience with Ray and at least two of NVIDIA NeMo RL, vLLM, SGLang, or NVIDIA Dynamo.
  • Strong knowledge of NVIDIA GPUs, CUDA, NCCL, GPU Operator, DCGM, MIG, RDMA, containers, CI/CD, GitOps, infrastructure-as-code, Python, and Linux.

Nice to have

  • Experience with Google TPUs and TPU-oriented frameworks or distributed workloads.
  • Experience with large language model training, fine-tuning, RLHF, or agentic reinforcement learning.
  • Knowledge of TensorRT-LLM, Triton Inference Server, DeepSpeed, Megatron-LM, or similar performance-oriented frameworks.
  • Experience operating secure, multi-tenant AI platforms in enterprise or regulated environments.

Culture & Benefits

  • Full-time employment with medical, dental, vision, life, and other benefits.
  • 401(k) with company matching and an employee stock purchase program.
  • Paid time off, company holidays, and family care and bonding leave.
  • Tuition reimbursement, student debt assistance, financial planning, wellness, and career development programs.
  • Equal opportunity employment with reasonable accommodation available for qualified individuals with disabilities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →