23 часа назад
HPC Operations Engineer
175 000 - 225 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
HPC Operations Engineer (Linux/HPC): Operating and supporting a large Research compute fleet across scheduling, compute, storage, access, provisioning, and service health with an accent on methodical troubleshooting and dependable day-to-day operations. Focus on resolving job and scheduler failures, provisioning Linux machines, maintaining runbooks, and identifying recurring issues for automation and permanent fixes.
Location: New York, United States
Salary: $175,000–$225,000 annual base salary, plus eligible discretionary bonus
Company
is a quantitative trading firm that develops high-performance electronic trading infrastructure and operates as a large global organization.
What you will do
- Provide first-line operational support for HPC users across scheduling, compute, storage, and access.
- Troubleshoot job failures, scheduler errors, resource constraints, and infrastructure incidents through resolution or documented escalation.
- Monitor queues, node status, storage, and service availability across the Research compute fleet.
- Perform maintenance, patching, configuration updates, provisioning, reinstalls, and decommissions of Linux machines.
- Create and maintain runbooks, knowledge-base articles, and user guides.
- Identify recurring issues and propose workflow improvements or automation candidates.
Requirements
- Bachelor’s degree in computer science, engineering, or a related field, or equivalent practical experience.
- 2+ years supporting Linux-based production environments.
- Solid Linux administration fundamentals, including RHEL-family and/or Ubuntu systems.
- Methodical troubleshooting skills and the ability to determine when escalation is required.
- Experience working directly with users in a technical support or operations role.
- Strong written communication and attention to operational detail.
Nice to have
- Bash or Python scripting for routine operational tasks.
- Experience with Slurm, HTCondor, or LSF batch schedulers.
- Knowledge of NFS, automounter, LDAP, PXE, kickstart, or Ansible.
- Prior HPC or large-scale compute experience.
- Familiarity with Prometheus and Grafana.
Culture & Benefits
- Generous paid time off policies.
- Hybrid working opportunities.
- Financial wellness tools and savings plans available by region.
- Free breakfast, lunch, and snacks in the office.
- Wellness reimbursements, sports teams, fitness events, volunteer opportunities, and charitable giving.
- Workshops and continuous learning opportunities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
1 день назад
Infrastructure Engineer (AI Hardware)
150 000 - 250 000$
Lambda
3 дня назад
Storage Engineer
267 000 - 356 000$
1 день назад
Sr. System Engineer (HPC/AI)
24 часа назад
Mid-Level Linux System Administrator (Linux)
95 000 - 135 000$
3 дня назад
Server Administrator 4 (HPC)
124 400 - 150 138$
23 часа назад
Cloud Support Engineer
145 000 - 175 000$