Назад
Company hidden
19 часов Π½Π°Π·Π°Π΄

Reliability Engineer (AI)

122Β 440 - 232Β 190$
Π€ΠΎΡ€ΠΌΠ°Ρ‚ Ρ€Π°Π±ΠΎΡ‚Ρ‹
onsite
Π’ΠΈΠΏ Ρ€Π°Π±ΠΎΡ‚Ρ‹
fulltime
Π“Ρ€Π΅ΠΉΠ΄
middle/senior
Английский
b2
Π‘Ρ‚Ρ€Π°Π½Π°
US
Вакансия ΠΈΠ· списка Hirify.GlobalВакансия ΠΈΠ· Hirify RU Global, списка ΠΊΠΎΠΌΠΏΠ°Π½ΠΈΠΉ с восточно-СвропСйскими корнями
Для мэтча ΠΈ ΠΎΡ‚ΠΊΠ»ΠΈΠΊΠ° Π½ΡƒΠΆΠ΅Π½ Plus

ΠœΡΡ‚Ρ‡ & Π‘ΠΎΠΏΡ€ΠΎΠ²ΠΎΠ΄

Для мэтча с этой вакансиСй Π½ΡƒΠΆΠ΅Π½ Plus

ОписаниС вакансии

ВСкст:
/

TL;DR

Reliability Engineer (AI): Defining and owning pod-level reliability specifications for large-scale AI data center hardware with an accent on MTBF, AFR, and RAS features. Focus on driving FMEA, root-cause analysis, and thermal/power redundancy to ensure system resilience and availability.

Location: Must be based in the US (On-site presence required in Beaver Brook, MA or Santa Clara, CA)

Salary: $122,440–$232,190 USD

Company

A global leader in semiconductor design and manufacturing, driving innovation in AI hardware and computing solutions.

What you will do

  • Define and maintain pod-level reliability and availability targets for compute, memory, storage, and cooling subsystems.
  • Translate system SLA requirements into actionable reliability specifications for silicon and platform teams.
  • Lead FMEA, root-cause analysis, and fleet failure-data analytics to implement corrective actions.
  • Architect RAS features including ECC, memory mirroring, and predictive failure telemetry.
  • Partner with facilities teams to ensure pod power and cooling redundancy and disaster-recovery readiness.
  • Establish qualification processes like HALT/HASS and track field KPIs against specifications.

Requirements

  • BS/MS/PhD in EE/ME Reliability or related field or 4-6+ years of relevant experience.
  • Strong expertise in RAS, FMEA, and statistical reliability methods (Weibull, FIT).
  • Experience authoring and owning reliability specifications and requirement flow-down.
  • Proven experience with large-scale fleet telemetry and thermal/power redundancy.
  • Must be authorized to work in the United States.

Nice to have

  • Experience with AI cluster operations.
  • Data analytics skills using Python and SQL.

Culture & Benefits

  • Competitive total compensation package including pay and stock bonuses.
  • Comprehensive health, retirement, and vacation benefit programs.
  • Opportunity to work on cutting-edge AI hardware in an agile, startup-like environment.
  • Commitment to ethical hiring practices and RBA compliance.

Π‘ΡƒΠ΄ΡŒΡ‚Π΅ остороТны: Ссли Ρ€Π°Π±ΠΎΡ‚ΠΎΠ΄Π°Ρ‚Π΅Π»ΡŒ просит Π²ΠΎΠΉΡ‚ΠΈ Π² ΠΈΡ… систСму, ΠΈΡΠΏΠΎΠ»ΡŒΠ·ΡƒΡ iCloud/Google, ΠΏΡ€ΠΈΡΠ»Π°Ρ‚ΡŒ ΠΊΠΎΠ΄/ΠΏΠ°Ρ€ΠΎΠ»ΡŒ, Π·Π°ΠΏΡƒΡΡ‚ΠΈΡ‚ΡŒ ΠΊΠΎΠ΄/ПО, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡ‚Π΅ этого - это мошСнники. ΠžΠ±ΡΠ·Π°Ρ‚Π΅Π»ΡŒΠ½ΠΎ ΠΆΠΌΠΈΡ‚Π΅ "ΠŸΠΎΠΆΠ°Π»ΠΎΠ²Π°Ρ‚ΡŒΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡˆΠΈΡ‚Π΅ Π² ΠΏΠΎΠ΄Π΄Π΅Ρ€ΠΆΠΊΡƒ. ΠŸΠΎΠ΄Ρ€ΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β†’