Назад
Company hidden
4 дня назад

Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program (Ceph)

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
Israel
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program (Ceph): Building an autonomous hardware-discovery and Ceph role-assignment pipeline that provisions real storage racks from bare metal to a healthy serving cluster with an accent on hardware inventory, declarative placement, and lifecycle automation. Focus on reconciling desired and observed state, managing cluster recovery and topology, and integrating multi-node bring-up with CI and lab hardware.

Location: Kfar Saba, Israel; onsite at the Kfar Saba office

Company

hirify.global develops Flash and advanced memory technologies and delivers storage solutions for businesses and digital products.

What you will do

  • Own automated hardware discovery for storage nodes, including DPU/SoC, SSD, DRAM, RNIC, and BMC inventory through Redfish, IPMI, PXE/DHCP, and first-boot agents.
  • Design declarative Ceph role assignment and cluster composition for OSD, MON, MGR, MDS, and NFS gateway services.
  • Drive bare-metal provisioning, OS and image deployment, cluster bootstrap, Ceph orchestrator deployment, SSD-based OSD provisioning, and convergence to HEALTH_OK.
  • Support autonomous zero-touch operation alongside deterministic manual overrides for lab, bring-up, failure-injection, and customer-shaped configurations.
  • Automate node lifecycle, drain and rebalance operations, daemon replacement, quorum recovery, failover validation, and re-discovery after re-imaging.
  • Integrate reconciliation, drift reporting, and multi-node cluster creation and teardown into CI and laboratory testing.

Requirements

  • Deep experience with lab hardware discovery, inventory, and fleet-scale automated provisioning.
  • Hands-on bare-metal automation using PXE/DHCP boot, BMC out-of-band management with Redfish/IPMI, OS deployment, and cloud-init or first-boot agents.
  • Experience with infrastructure-as-code and configuration automation using Ansible-, Terraform-, or equivalent tooling, with a declarative desired-state approach.
  • Strong Python and shell scripting skills and experience with Jenkins, GitLab CI, or equivalent CI systems.
  • Ability to design autonomous systems with reliable manual override controls.
  • Work onsite from the Kfar Saba office in Israel.

Nice to have

  • Operational knowledge of Ceph roles, cephadm-style orchestration, placement specifications, ceph-volume, CRUSH maps, failure domains, quorum, and cluster health and lifecycle.
  • Distributed-systems experience, including quorum, Paxos, rebalance, recovery, and failure-domain reasoning.
  • Experience with RoCEv2/Ethernet storage-fabric bring-up and port discovery.
  • Familiarity with DPU/SoC-based nodes and constrained-node environments.

Culture & Benefits

  • Full-time exempt employment.
  • Inclusive environment focused on diversity, belonging, respect, and contribution.
  • Accessibility support is available throughout the hiring process.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →