8 часов назад
Senior HPC Storage Engineer (HPC Storage)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior HPC Storage Engineer (HPC Storage): Architecting, modernizing, and optimizing secure high-performance storage and supporting infrastructure across classified HPC environments with an accent on scalability, reliability, performance, and cybersecurity. Focus on designing parallel file-system and data-movement solutions, automating infrastructure with Python and Bash, resolving complex production issues, and supporting AI/ML and other data-intensive workloads.
Location: Oak Ridge, Tennessee, USA
Company
Oak Ridge National Laboratory is a U.S. Department of Energy national laboratory focused on scientific breakthroughs and solutions for energy, environmental, and national security challenges.
What you will do
- Lead the architecture, design, modernization, and optimization of storage and supporting systems across classified HPC environments.
- Translate scientific and mission requirements into secure, scalable, resilient infrastructure solutions.
- Analyze storage performance, availability, scalability, data movement, and system reliability in mission-critical environments.
- Develop prototypes, automation, monitoring, metrics, and observability solutions using Python, Bash, Grafana, Nagios, and related tools.
- Lead troubleshooting and root-cause analysis, coordinate lifecycle planning, and resolve hardware and software issues with vendors and engineering teams.
- Advise on technical strategy, mentor staff, evaluate emerging HPC and AI/ML technologies, and shape infrastructure roadmaps.
Requirements
- BS degree in Computer Science, Information Technology, Engineering, or a related field, plus at least eight years of relevant experience; equivalent education, experience, and certifications may be considered.
- At least five years of experience supporting Intelligence Community environments or missions.
- Ability to independently lead complex technical efforts and translate requirements into scalable infrastructure solutions.
- Current Top Secret clearance with SCI eligibility and eligibility for Sensitive or Special Access Program information are required.
- Q Clearance with SCI must be obtained and maintained; pre-placement and random substance-abuse testing and possible polygraph testing apply.
- Experience with HPC storage, Linux engineering, high-performance networking, automation, and infrastructure supporting data-intensive workloads.
Nice to have
- Experience with Lustre, IBM Storage Scale/GPFS, BeeGFS, or similar parallel file systems.
- Experience with InfiniBand, TCP/IP, Fibre Channel, tiered storage, data lifecycle management, resiliency, and performance optimization.
- Experience with Ansible, Git-based workflows, CI/CD, infrastructure-as-code, or modern DevOps practices.
- Experience designing infrastructure for HPC, AI/ML, GPU computing, or other data-intensive workloads.
Culture & Benefits
- Collaboration with scientists, researchers, HPC engineers, cybersecurity professionals, technical leaders, vendors, and mission partners.
- Medical, prescription, dental, vision, retirement, pension, life insurance, and disability benefits.
- Vacation, holidays, parental leave, flexible spending and health savings accounts, wellness programs, and educational assistance.
- On-site fitness, banking, and cafeteria facilities, plus employee discounts and relocation assistance.
- Participation in a planned maintenance schedule and on-call rotation for mission-critical systems.
Hiring process
- The position remains open for at least five days and closes when a qualified candidate is identified or hired.
- Employment requires successful completion of applicable clearance, credentialing, substance-abuse testing, and eligibility procedures.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →