HPC AI Systems Administrator (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
HPC AI Systems Administrator (AI Infrastructure): Designing and maintaining a secure, scalable on-premises HPC cluster and multi-GPU architecture to support corporate data initiatives with an accent on platform enablement and compute efficiency. Focus on orchestrating developer sandboxes, optimizing GPU workloads, and ensuring enterprise-level security compliance.
Location: On-site in Houston, TX
Company
provides professional services and strategic infrastructure support for enterprise clients.
What you will do
- Lead the deployment, bare-metal configuration, and optimization of on-premises HPC clusters and multi-GPU architectures.
- Manage the end-to-end AI software stack, including Linux OS, GPU drivers, CUDA, NCCL, and containerization platforms.
- Implement workload scheduling and orchestration systems using Kubernetes, Slurm, or equivalent enterprise platforms.
- Establish automated telemetry and monitoring dashboards to track hardware utilization, thermal limits, and memory bandwidth.
- Operationalize strict security baselines and Zero Trust network architecture to ensure data privacy and compliance.
- Act as the primary technical interface for high-end hardware vendors and system integrators.
Requirements
- 3+ years of systems administration experience managing Linux-based HPC environments or enterprise-scale GPU infrastructure.
- Hands-on experience configuring and maintaining modern enterprise GPU hardware (e.g., NVIDIA Ampere or Hopper).
- Deep expertise in Linux system engineering, Docker, and Apptainer/Singularity.
- Knowledge of high-throughput networking fabrics (InfiniBand/RoCE) and distributed enterprise storage systems.
- Bachelor’s degree in Computer Science, Computer Engineering, System Administration, or equivalent experience.
- Must be based in or able to work on-site in Houston, TX.
Nice to have
- Professional certifications in enterprise AI infrastructure, virtualization, or cloud/hybrid architecture.
- Familiarity with infrastructure requirements for LLM fine-tuning pipelines and machine learning lifecycles.
Culture & Benefits
- Direct access to cutting-edge, top-tier enterprise compute infrastructure.
- Collaborative work environment with strategic, top-down execution support.
- Competitive salary and comprehensive benefits package.
- Professional development support.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →