2 дня назад
Operations Engineering Manager (HPC/GPU)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Operations Engineering Manager (HPC/GPU) (GPU-accelerated HPC infrastructure): Owning operational reliability and leading an Operations Engineering team with an accent on Linux administration, incident management, automation, observability, and scaled Agile delivery. Focus on improving infrastructure availability and performance, driving automation and operational documentation, and coordinating incident response across Platform, Network, Infrastructure, and external support teams.
Location: London, United Kingdom; flexible work from home
Company
develops sustainable data-center infrastructure, cooling concepts, and software solutions for accelerated compute and HPC.
What you will do
- Lead, coach, and develop a team of Operations Engineers, including goal setting, workload allocation, performance reviews, and development planning.
- Own the reliability, performance, and availability of GPU-accelerated HPC infrastructure.
- Oversee monitoring, incident trend analysis, root cause analysis, operational metrics, runbooks, and change management.
- Champion the scaled Agile Framework for Operations Engineering and coordinate planning, backlogs, ceremonies, and delivery flow.
- Drive automation, scripting, configuration management, observability, and maintenance of operational documentation.
- Represent Operations in cross-functional planning and act as an escalation point during major incidents and complex technical issues.
Requirements
- 5+ years of experience in infrastructure or operations, including 2+ years managing a technical team.
- Advanced production Linux administration experience, ideally at scale.
- Experience with incident and problem management and external support teams.
- Hands-on experience with automation tools such as Ansible and monitoring or observability tools such as Grafana and Prometheus.
- Experience with Agile ways of working and exposure to scaled Agile frameworks.
- Strong communication and stakeholder management skills across Platform, Network, Infrastructure, and leadership teams.
Nice to have
- Experience with HPC or GPU-accelerated environments, including NVIDIA GPUs, InfiniBand/RDMA, or parallel file systems.
- Python or Bash scripting for automation and tooling.
- Experience with HPC/GPU performance tuning, CI/CD pipelines, DevOps tooling, on-call rotations, runbooks, or incident readiness.
Culture & Benefits
- Flexible work from home with an international and virtual team.
- Hardware support for HPC-focused work.
- Collaborative environment where ideas are valued over hierarchy.
- Opportunity to shape new departments, processes, and company culture.
- Focus on sustainability and carbon neutrality for data centers.
- Wellbeing, diversity, and inclusion initiatives.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →