Назад
Company hidden
2 месяца назад

Founding Engineering Manager, Production Engineering (AI Infrastructure)

Формат работы
onsite
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
Israel
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Founding Engineering Manager, Production Engineering (AI Infrastructure): Building the Production Engineering organization and reliability systems for Crusoe Cloud's large-scale GPU infrastructure with an accent on team leadership, software-defined operations, and incident management. Focus on automating physical-to-digital remediation, improving fleet observability, and scaling production systems across complex compute, storage, networking, and platform environments.

Location: Tel Aviv, Israel; on-site

Company

hirify.global builds vertically integrated AI infrastructure and cloud services spanning clean energy generation, data centers, GPUs, and software.

What you will do

  • Establish and lead the Production Engineering organization in Tel Aviv, including recruiting, mentoring, and setting operational standards.
  • Own incident response and partner with US and Dublin teams on a follow-the-sun global on-call rotation.
  • Drive alert reduction, runbook automation, predictive monitoring, and software-defined remediation using tools such as Temporal.
  • Govern Production Readiness Reviews and change control across compute, storage, networking, and platform teams.
  • Protect at least 30% of the team's capacity for strategic automation, tooling, and firmware optimization.
  • Remain hands-on in coding and incident response while converting repeated physical interventions into automated solutions.

Requirements

  • 8+ years of experience in infrastructure, SRE, or production engineering.
  • 2+ years of direct leadership experience with first-line engineering teams in a high-growth neocloud, hyperscaler, or large-scale distributed environment.
  • Strong hands-on coding skills in Go, Python, C++, or a comparable systems language.
  • Expertise in Linux internals, container orchestration at scale, and root-cause analysis across physical and virtual systems.
  • Experience operating tiered on-call models, defining SLIs, SLOs, and error budgets, and reducing paging fatigue.
  • Ability to work on-site in Tel Aviv, Israel.

Nice to have

  • Experience at a neocloud or AI infrastructure company operating large GPU clusters.
  • Exposure to InfiniBand, RoCEv2, BMC systems, firmware qualification, or hardware attestation.
  • Knowledge of NVIDIA or AMD accelerator failure modes and DCGM counters.
  • Experience scaling systems to support 10x fleet expansion.

Culture & Benefits

  • Blameless post-mortems focused on systemic failures rather than human error.
  • Cross-functional collaboration with energy, manufacturing, data center construction, and cloud services specialists.
  • Pension contributions and additional benefits aligned with local market standards.
  • Benefits supporting financial security, health, and work-life balance.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →