Назад
2 дня назад

Senior SRE Platform Architect (AI)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior/lead
Английский
c1
Страна
US
vacancy_detail.hirify_telegram_tooltipВакансия из Telegram канала -

Мэтч & Сопровод

Покажет вашу совместимость и напишет письмо

Описание вакансии

Sr. SRE Platform Architect

Company

Bitdeer

Conditions

6 days agoHead San Jose, CA / Austin, TX - Hybrid Hybrid Full Time Devops Jobs by Bitdeer

Bitdeer Bitdeer is a NASDAQ-listed (BTDR) high-performance computing and Bitcoin mining company headquartered in Singapore. It provides end-to-end Bitcoin mining solutions (mining hardware like SEALMINER, Minerbase containers, cloud mining, hosting, and mining farm/data center operations) as well as AI Cloud services offering GPU compute (NVIDIA GB200 NVL72, B200, H200, H100) for AI training and deployment. Its customers range from individual and institutional Bitcoin miners to enterprises and developers needing scalable AI/ML compute infrastructure. Singapore, SG Funding Loans & Equity ($179M) Unknown ($179T) Investors Matrixport Ventures Projects Bitdeer AI About Bitdeer Bitdeer Technologies Group is a NASDAQ-listed (ticker: BTDR) technology company headquartered in Singapore that describes itself as a "world-leading" high-performance computing platform and Bitcoin mining services provider. The company is vertically integrated across the value chain, spanning IC design and hardware manufacturing (its own SEALMINER ASIC miners and Minerbase mobile cooling containers), infrastructure construction and cloud mining, and artificial intelligence/high-performance computing. Bitdeer handles the full range of mining-related processes for its customers, including equipment procurement, transport logistics, datacenter design and construction, equipment management, and daily operations, and offers institutional services, a hash rate market, and a miner rights trading marketplace via its mobile apps (Bitdeer App and Minerplus App). Since 2013 Bitdeer has built more than 30 data centers globally and currently operates 9 large-scale data centers (including one of North America's largest) with roughly 3GW of diversified energy capacity and tens of exahashes of managed hash rate, with major operations in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia. Since 2023 it has also been expanding a global AI infrastructure business (Bitdeer AI Cloud), powered by thousands of NVIDIA GPUs (including GB200 NVL72 and B200, with GB300 NVL72 and B300 planned), offering turnkey AI datacenter solutions and GPU cloud compute for AI training and deployment starting at around $2/hour. Bitdeer serves both individual/retail Bitcoin miners and institutional clients, as well as AI developers and enterprises seeking scalable, energy-efficient compute. View jobs by Bitdeer

Skills

Ai Infrastructure Bios/Firmware Lifecycle Bmc Ceph Cloud Infrastructure Control Plane Data Center Operations Dcgm Ddn Design Review Distributed Storage Domain Driven Design Firmware Gitops Gpu Gpu Operator Hpc Storage Infiniband Ipmi Kuberay Kubernetes Kubernetes Operators Kueue Lustre Mig Multi-Region Architecture Multi-Region Deployment Nccl Netapp Networking Nvidia Nvlink Nvme-Of Nvswitch Observability Platform Architecture Plugin Framework Design Postmortem Post-Mortem Facilitation Provisioning Pure Rancher Ray Redfish Rma Roce Runbook Authorship Runbooks Site Reliability Engineering Sli Slo Slurm Sre Technical Writing Vast Vgpu Volcano Weka Ztp

About the Role

You will lead the design, development, and evolution of Bitdeer's next-generation public cloud platform, owning the end-to-end architecture across CPU, GPU, RDS, storage, networking, serverless, and AI services. You'll be the single point of architectural accountability for the NeoCloud SRE platform, a multi-region GPU rental fleet spanning self-built and OEM-rented data centers, with dozens of bounded contexts, frameworks, and three operational tiers. You'll write and defend the design under review, and shepherd it through the engineering squads that build it. You'll collaborate with cross-functional teams and global partners to define the cloud technology roadmap, optimize multi-region deployments, and deliver world-class infrastructure and platform solutions that power large-scale AI and enterprise workloads.

Requirements

  • 10+ years of production SRE, platform-engineering, or infra-architecture experience, including 3+ years at architect level
  • Hands-on experience with GPU / AI-compute infrastructure including NVIDIA GPU ops (DCGM, MIG, vGPU, NVLink/NVSwitch, XID semantics, NCCL), InfiniBand or RoCE fabrics, and HPC storage (Lustre, NetApp/Pure/DDN/VAST, NVMe-oF)
  • Multi-region observability at scale including metrics, logs, traces, profiles, analytics-lake substrate, recording rules, MWMBR burn-rate alerting, and SLI/SLO discipline
  • Experience with cluster platforms including Kubernetes (control plane, GPU Operator, topology-aware scheduling) and at least one of Slurm, Volcano, Kueue, Ray, or KubeRay
  • Data-center operations experience including ZTP, BMC/IPMI/Redfish, BIOS/firmware lifecycle, RMA, and multi-vendor OEM management
  • Strong DDD instincts including bounded contexts, public contracts, and one-context-one-repo discipline
  • Experience designing a plugin framework with a uniform manifest and lifecycle
  • Writing fluency to author and maintain a multi-thousand-line architecture document as well as executive one-pagers
  • Cross-team operating tempo including design reviews, runbook authorship, on-call shadowing, and post-mortem facilitation
  • Hyperscale or NeoCloud experience
  • BS/MS in Computer Science or similar

Responsibilities

  • Write and maintain the platform architecture document, keeping the design coherent across all sections, frameworks, and tiers
  • Review every framework-level change including new bounded contexts, plugin kinds, tier-deployment shifts, schema changes, naming changes, and cross-context contract changes
  • Set design invariants such as residency rules, Tier 2 self-sufficiency budget, survival-uplink contracts, naming conventions, SLO catalogues, and redaction-at-boundary rules
  • Run the plugin framework and author and evolve the uniform contract for every extension
  • Decide tier placement for Edge DC versus Regional Controller versus Global Hub, making data-residency, compliance, and availability tradeoffs explicit
  • Coordinate with cloud-service teams and tenants who author plugins, SDKs, dashboards, and agent recipes riding the platform
  • Coordinate with Security on joint ownership of vulnerability management, exposure management, and joint operations
  • Pre-flight roadmap items by producing one-page designs that fit the existing layered model, tier topology, naming conventions, and extension contracts before implementation
  • Defend the design under review, rejecting scope creep and one-off integrations while approving genuinely needed new plugin kinds

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →

Текст вакансии взят без изменений

Источник -