Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Reliability Engineer SME (M&E) (AI Infrastructure): Providing mechanical and electrical engineering support for Nscale’s EMEA data centre estate with an accent on high-density AI compute, cooling performance, power resilience, and engineering assurance. Focus on reviewing designs, approving changes on live systems, diagnosing complex incidents, leading RCA and CAPA, and modelling thermal and power capacity for GPU deployments.
Location: London; regular site travel required across EMEA.
Company
Nscale provides a GPU cloud infrastructure platform for AI start-ups and enterprise customers.
What you will do
- Act as the technical escalation point for mechanical and electrical systems across owned, operated, and third-party data centres.
- Set maintenance and engineering standards, review designs and vendor documentation, and provide technical input to procurement and service contracts.
- Support AI infrastructure teams with high-density GPU rack cooling and power distribution systems.
- Approve technical changes, audit sites and colocation facilities, and assess resilience, maintainability, compliance, and operational risk.
- Support live incident response, lead root cause analysis and corrective and preventive actions, and drive resolution of recurring faults.
- Maintain the standards library, train operations teams, support site mobilisation and due diligence, and contribute to capacity, thermal, lifecycle, and energy-efficiency planning.
Requirements
- Strong building services engineering background in critical environments, with deep hands-on expertise in mechanical or electrical systems and working knowledge of the other discipline.
- Experience with data centre cooling systems such as chillers, CRAC/CRAH units, air handling, piping, and secondary cooling loops, or with LV/HV distribution, UPS, generators, switchgear, and protection systems.
- Experience conducting M&E design reviews, technical audits, and engineering standards compliance reviews in critical infrastructure or colocation environments.
- Experience approving changes on live M&E systems and understanding the associated operational risks.
- Experience with RCA and CAPA after M&E incidents, structured problem solving, incident support, and coaching operational engineers.
- Strong communication skills for explaining technical electrical concepts and holding vendors and contractors accountable.
Nice to have
- HV Authorised Person status or progress toward authorisation.
- Experience with liquid cooling and high-density AI/GPU compute.
- Chartered or Incorporated Engineer status, or an equivalent recognised qualification.
- Experience with thermal and power capacity modelling for GPU deployments.
- Experience across owned and operated data centres and third-party colocation environments.
Culture & Benefits
- Collaborative, supportive, and innovative working environment focused on ownership and accountability.
- Competitive package including base salary, bonus, and equity.
- Compensation reviews every 12 months.
- Flexible workplace focused on employee autonomy.
- Progression plan tailored to professional ambitions.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Manufacturing Engineer Manager (AI Infrastructure)
2 дня назад
Equipment Engineer (Nuclear)
40 000 - 50 000GBP
6 дней назад
Equipment Engineer, Technical Customer Support
2 дня назад
Controls Field Engineer (AI Data Centers)
3 дня назад
MEP Design Engineer (AI Data Center)
3 дня назад