Назад
2 дня назад

Design Reliability Engineer (Power and Energy)

130 000 - 200 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Design Reliability Engineer (Power and Energy): Building and maintaining availability models, FMEA programs, and reliability foundations for AI data center power systems with an accent on RAM analysis, common-mode failure detection, and digital twin development. Focus on allocating SLA targets across generation and distribution systems, validating reliability assumptions, and turning operational data into auditable availability predictions.

Location: Houston, TX; remote considered within the US, with regular site and vendor travel

Salary: $130,000–$200,000 USD base salary, plus potential bonus, equity, and/or commission

Company

Nscale develops sovereign generative AI infrastructure using high-performance GPUs and sustainable data centers across North America, Europe, and other markets.

What you will do

  • Own campus-level availability models covering generation, electrical distribution, fuel supply, cooling, and the full path to GPU capacity.
  • Allocate contractual SLA targets into subsystem availability budgets, redundancy requirements, MTTR assumptions, sparing, and maintenance strategies.
  • Lead FMEA/FMECA programs for engines, generators, BESS, switchgear, transformers, fuel gas systems, and cooling infrastructure.
  • Build the reliability core of the digital twin, including data structures, event capture, failure coding, downtime attribution, and run-hour tracking.
  • Support commissioning, reliability runs, root cause analyses, corrective actions, design reviews, HAZOPs, and vendor evaluations.
  • Direct RAM consultants, develop reliability standards, and communicate model confidence, sensitivities, and risks to executives.

Requirements

  • 5+ years of reliability engineering experience in power generation, process environments, or mission-critical facilities, including ownership of RAM analysis for large redundant systems.
  • Expertise in discrete-event simulation such as BlockSim or Raptor and closed-form analytical reliability methods.
  • Experience leading FMEA/FMECA for rotating, electrical, or process equipment and using IEEE 493, OREDA, IEEE 3006, or similar data sources.
  • Ability to allocate availability targets from commercial commitments to systems and defend the analysis to executives, customers, or insurers.
  • Software capability with Python or similar tools, including scripting analyses and structuring reliability data for automation.
  • Bachelor’s degree in Mechanical, Electrical, Chemical, Reliability Engineering, or a related field; regular travel to project sites and vendor facilities is required, typically 15–25%.

Nice to have

  • Experience with digital twins, telemetry-driven reliability, or operational availability measurement.
  • CRE certification or equivalent.
  • PE license.

Culture & Benefits

  • Flexible workplace with autonomy and a collaborative, supportive environment.
  • Competitive US compensation with base salary, bonus, and equity opportunities.
  • Medical, dental, and vision coverage.
  • 401(k) retirement plan with company match.
  • Generous paid time off and US federal holidays.
  • Opportunity to join an early-stage, rapidly growing AI infrastructure company.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →