4 дня назад
Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program (Ceph)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program (Ceph): Building an autonomous hardware-discovery and Ceph role-assignment pipeline that provisions real storage racks from bare metal to a healthy serving cluster with an accent on hardware inventory, declarative placement, and lifecycle automation. Focus on reconciling desired and observed state, managing cluster recovery and topology, and integrating multi-node bring-up with CI and lab hardware.
Location: Kfar Saba, Israel; onsite at the Kfar Saba office
Company
develops Flash and advanced memory technologies and delivers storage solutions for businesses and digital products.
What you will do
- Own automated hardware discovery for storage nodes, including DPU/SoC, SSD, DRAM, RNIC, and BMC inventory through Redfish, IPMI, PXE/DHCP, and first-boot agents.
- Design declarative Ceph role assignment and cluster composition for OSD, MON, MGR, MDS, and NFS gateway services.
- Drive bare-metal provisioning, OS and image deployment, cluster bootstrap, Ceph orchestrator deployment, SSD-based OSD provisioning, and convergence to HEALTH_OK.
- Support autonomous zero-touch operation alongside deterministic manual overrides for lab, bring-up, failure-injection, and customer-shaped configurations.
- Automate node lifecycle, drain and rebalance operations, daemon replacement, quorum recovery, failover validation, and re-discovery after re-imaging.
- Integrate reconciliation, drift reporting, and multi-node cluster creation and teardown into CI and laboratory testing.
Requirements
- Deep experience with lab hardware discovery, inventory, and fleet-scale automated provisioning.
- Hands-on bare-metal automation using PXE/DHCP boot, BMC out-of-band management with Redfish/IPMI, OS deployment, and cloud-init or first-boot agents.
- Experience with infrastructure-as-code and configuration automation using Ansible-, Terraform-, or equivalent tooling, with a declarative desired-state approach.
- Strong Python and shell scripting skills and experience with Jenkins, GitLab CI, or equivalent CI systems.
- Ability to design autonomous systems with reliable manual override controls.
- Work onsite from the Kfar Saba office in Israel.
Nice to have
- Operational knowledge of Ceph roles, cephadm-style orchestration, placement specifications, ceph-volume, CRUSH maps, failure domains, quorum, and cluster health and lifecycle.
- Distributed-systems experience, including quorum, Paxos, rebalance, recovery, and failure-domain reasoning.
- Experience with RoCEv2/Ethernet storage-fabric bring-up and port discovery.
- Familiarity with DPU/SoC-based nodes and constrained-node environments.
Culture & Benefits
- Full-time exempt employment.
- Inclusive environment focused on diversity, belonging, respect, and contribution.
- Accessibility support is available throughout the hiring process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →