2 месяца назад
Founding Engineering Manager, Production Engineering (AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Founding Engineering Manager, Production Engineering (AI Infrastructure): Building the Production Engineering organization and reliability systems for Crusoe Cloud's large-scale GPU infrastructure with an accent on team leadership, software-defined operations, and incident management. Focus on automating physical-to-digital remediation, improving fleet observability, and scaling production systems across complex compute, storage, networking, and platform environments.
Location: Tel Aviv, Israel; on-site
Company
builds vertically integrated AI infrastructure and cloud services spanning clean energy generation, data centers, GPUs, and software.
What you will do
- Establish and lead the Production Engineering organization in Tel Aviv, including recruiting, mentoring, and setting operational standards.
- Own incident response and partner with US and Dublin teams on a follow-the-sun global on-call rotation.
- Drive alert reduction, runbook automation, predictive monitoring, and software-defined remediation using tools such as Temporal.
- Govern Production Readiness Reviews and change control across compute, storage, networking, and platform teams.
- Protect at least 30% of the team's capacity for strategic automation, tooling, and firmware optimization.
- Remain hands-on in coding and incident response while converting repeated physical interventions into automated solutions.
Requirements
- 8+ years of experience in infrastructure, SRE, or production engineering.
- 2+ years of direct leadership experience with first-line engineering teams in a high-growth neocloud, hyperscaler, or large-scale distributed environment.
- Strong hands-on coding skills in Go, Python, C++, or a comparable systems language.
- Expertise in Linux internals, container orchestration at scale, and root-cause analysis across physical and virtual systems.
- Experience operating tiered on-call models, defining SLIs, SLOs, and error budgets, and reducing paging fatigue.
- Ability to work on-site in Tel Aviv, Israel.
Nice to have
- Experience at a neocloud or AI infrastructure company operating large GPU clusters.
- Exposure to InfiniBand, RoCEv2, BMC systems, firmware qualification, or hardware attestation.
- Knowledge of NVIDIA or AMD accelerator failure modes and DCGM counters.
- Experience scaling systems to support 10x fleet expansion.
Culture & Benefits
- Blameless post-mortems focused on systemic failures rather than human error.
- Cross-functional collaboration with energy, manufacturing, data center construction, and cloud services specialists.
- Pension contributions and additional benefits aligned with local market standards.
- Benefits supporting financial security, health, and work-life balance.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →