Назад
4 дня назад

Reliability Engineer SME (M&E) (AI Infrastructure)

Формат работы
remote (только United_kingdom)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Reliability Engineer SME (M&E) (AI Infrastructure): Providing mechanical and electrical engineering support for Nscale’s EMEA data centre estate with an accent on high-density AI compute, cooling performance, power resilience, and engineering assurance. Focus on reviewing designs, approving changes on live systems, diagnosing complex incidents, leading RCA and CAPA, and modelling thermal and power capacity for GPU deployments.

Location: London; regular site travel required across EMEA.

Company

Nscale provides a GPU cloud infrastructure platform for AI start-ups and enterprise customers.

What you will do

  • Act as the technical escalation point for mechanical and electrical systems across owned, operated, and third-party data centres.
  • Set maintenance and engineering standards, review designs and vendor documentation, and provide technical input to procurement and service contracts.
  • Support AI infrastructure teams with high-density GPU rack cooling and power distribution systems.
  • Approve technical changes, audit sites and colocation facilities, and assess resilience, maintainability, compliance, and operational risk.
  • Support live incident response, lead root cause analysis and corrective and preventive actions, and drive resolution of recurring faults.
  • Maintain the standards library, train operations teams, support site mobilisation and due diligence, and contribute to capacity, thermal, lifecycle, and energy-efficiency planning.

Requirements

  • Strong building services engineering background in critical environments, with deep hands-on expertise in mechanical or electrical systems and working knowledge of the other discipline.
  • Experience with data centre cooling systems such as chillers, CRAC/CRAH units, air handling, piping, and secondary cooling loops, or with LV/HV distribution, UPS, generators, switchgear, and protection systems.
  • Experience conducting M&E design reviews, technical audits, and engineering standards compliance reviews in critical infrastructure or colocation environments.
  • Experience approving changes on live M&E systems and understanding the associated operational risks.
  • Experience with RCA and CAPA after M&E incidents, structured problem solving, incident support, and coaching operational engineers.
  • Strong communication skills for explaining technical electrical concepts and holding vendors and contractors accountable.

Nice to have

  • HV Authorised Person status or progress toward authorisation.
  • Experience with liquid cooling and high-density AI/GPU compute.
  • Chartered or Incorporated Engineer status, or an equivalent recognised qualification.
  • Experience with thermal and power capacity modelling for GPU deployments.
  • Experience across owned and operated data centres and third-party colocation environments.

Culture & Benefits

  • Collaborative, supportive, and innovative working environment focused on ownership and accountability.
  • Competitive package including base salary, bonus, and equity.
  • Compensation reviews every 12 months.
  • Flexible workplace focused on employee autonomy.
  • Progression plan tailored to professional ambitions.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →