Назад
Company hidden
56 минут назад

High Performance Compute Systems Site Lead (Onsite - LANL)

105 500 - 243 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior/lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
High Performance Compute Systems Site Lead (HPC/AI): Leading onsite technical delivery for large-scale HPE Cray, Linux, and AI systems at Los Alamos National Laboratory with an accent on incident coordination, maintenance planning, hardware diagnostics, and customer-facing service leadership. Focus on troubleshooting complex interactions across compute, interconnect, storage, power, cooling, and operating-system infrastructure while preparing the site for next-generation HPC platforms.

Location: Daily onsite work in Los Alamos, New Mexico, United States. This is not a remote or hybrid position. Additional onsite support is required for planned maintenance, major incidents, and on-call responsibilities.

Salary: $105,500–$243,000 annual base salary in New Mexico, United States. Variable incentives may also be offered.

Company

hirify.global is a global edge-to-cloud technology company delivering enterprise infrastructure, high-performance computing, AI, and data solutions.

What you will do

  • Provide onsite technical leadership for HPE hardware engineers, Linux administrators, software analysts, inventory specialists, and remote engineering resources.
  • Coordinate daily priorities, support cases, maintenance windows, upgrades, installations, acceptance activities, and service-delivery commitments.
  • Lead onsite responses to major incidents, escalations, root-cause analyses, corrective actions, and customer operational reviews.
  • Diagnose and maintain large-scale HPE Cray HPC, Linux, high-speed interconnect, storage, power, cooling, and management infrastructure.
  • Perform hands-on rack, cabling, fiber, hardware replacement, out-of-band management, firmware, and data-center support activities.
  • Create procedures, incident records, maintenance plans, technical documentation, and automation using Bash, Python, Git, and related tools.

Requirements

  • US citizenship and the ability to obtain and maintain a DOE Q Clearance are required.
  • Onsite work in Los Alamos, New Mexico, Monday through Friday, with on-call, after-hours, and major-incident support.
  • High school diploma with at least 7 years of relevant experience, or a technical associate/bachelor’s degree with at least 5 years of relevant experience.
  • At least 5 years supporting enterprise server hardware and data-center infrastructure, including processors, memory, storage, power, BMCs, network adapters, and cabling.
  • At least 3 years supporting HPC or large-scale Linux environments, Linux system administration, and technical leadership or work coordination.
  • Experience with Bash or Python, structured troubleshooting, formal service processes, hardware tools, technical documentation, and 24x7 production support; ability to lift up to 50 pounds independently and up to 75 pounds with assistance.

Nice to have

  • DOE Q, DoD Top Secret, or comparable federal security-clearance experience.
  • Experience with HPE Cray EX, Slingshot, InfiniBand, high-speed Ethernet, parallel filesystems, workload managers, or HPC monitoring systems.
  • Experience with liquid-cooled infrastructure, NVIDIA GB200 or GB300 NVL72, NVIDIA DGX, GPU systems, Redfish, IPMI, or BMC technologies.
  • DOE national laboratory, DoD, government research, regulated-environment, project-management, ITIL, Git, or relevant technical certification experience.

Culture & Benefits

  • Health, financial, and emotional wellbeing benefits for employees and their families.
  • Professional development programs supporting technical expertise and career growth.
  • Inclusive workplace focused on varied backgrounds, collaboration, and personal flexibility.
  • Mission-critical work supporting scientific discovery and national security through HPC and AI infrastructure.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →