13 дней назад
Senior IT Systems Administrator (Linux & HPC)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior IT Systems Administrator (Linux & HPC) (Linux/HPC infrastructure): Administering and securing enterprise Linux servers and cloud or co-located HPC platforms with an accent on SLURM workload scheduling, cluster management, storage, and network integration. Focus on diagnosing complex cross-platform issues, automating administration, maintaining resilience, and resolving performance and availability challenges across compute, storage, and infrastructure services.
Location: Leatherhead, Surrey, United Kingdom; hybrid work arrangement
Company
delivers professional services, technologies, engineering, scientific, and lifecycle solutions to government, defense, and industrial clients worldwide.
What you will do
- Administer, patch, harden, upgrade, monitor, and troubleshoot enterprise Linux platforms, especially Red Hat Enterprise Linux or comparable distributions.
- Operate HPC clusters across management, login, compute, and storage components, including SLURM queues, partitions, scheduling policies, accounting, and workload troubleshooting.
- Support NVIDIA Base Command Manager and Azure CycleCloud for cluster provisioning, node lifecycle management, monitoring, and SLURM integration.
- Manage physical and virtual infrastructure, Cisco compute platforms, NetApp file services, NFS, networking, and storage connectivity.
- Implement secure configuration, backup, recovery, disaster-recovery, monitoring, incident response, root-cause analysis, and controlled change processes.
- Provide technical guidance, peer review, documentation, knowledge transfer, and operational support, including on-call participation where required.
Requirements
- Substantial hands-on experience administering Linux in a complex enterprise or research-computing environment.
- Practical experience supporting HPC clusters and troubleshooting compute, scheduler, storage, and network layers.
- Strong SLURM administration and workload-troubleshooting skills, plus Bash or another relevant scripting language.
- Experience with Linux performance analysis, capacity management, patching, security hardening, vulnerability remediation, shared file services, and NFS.
- Knowledge of enterprise networking fundamentals, distributed-system dependencies, technical change management, and structured IT service management.
- Bachelor’s degree in computing, engineering, or a related field with relevant experience, or equivalent professional training and practical experience.
Nice to have
- Experience with NVIDIA Base Command Manager, Bright Cluster Manager, Cisco compute platforms, or NetApp storage.
- Knowledge of InfiniBand, MPI workloads, environment modules, compilers, engineering or scientific applications, and HPC software stacks.
- Experience with Ansible, Git, infrastructure as code, virtualization, containers, or cloud-hosted Linux workloads.
- Relevant Linux, HPC, Cisco, NVIDIA, or NetApp certifications.
Culture & Benefits
- Collaboration with government and industry clients on engineering, technology, science, and lifecycle projects.
- Professional development opportunities and competitive benefits.
- Work focused on safety, reliability, secure services, creativity, resourcefulness, and collaboration.
- Opportunities to contribute to projects spanning planning, design, operations, maintenance, and sustainability.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →