Назад
5 дней назад

Distributed Software Engineer

Формат работы
hybrid
Тип работы
fulltime
Грейд
lead
Страна
Canada
vacancy_detail.hirify_telegram_tooltipВакансия из Telegram канала -

Мэтч & Сопровод

Покажет вашу совместимость и напишет письмо

Описание вакансии

TL;DR
Distributed Software Engineer (Go/Kubernetes): Building and operating reliable, observable clusters across large fleets of systems, servers, and switches with an accent on bare-metal automation, Kubernetes operators, control-plane services, and observability. Focus on designing CRD-driven infrastructure, automating installation and recovery, and exposing cluster capabilities through APIs, CLIs, and an MCP gateway.

Distributed Software Engineer

Company

Cerebras Systems, Inc.

Conditions

6 days ago

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build and operate software that turns large fleets of Cerebras systems, servers, and switches into reliable, observable clusters. You will automate bare-metal infrastructure, create Kubernetes operators and control-plane services, improve fleet reliability, and expose cluster capabilities through APIs, CLIs, and an MCP gateway.

Requirements

  • 5+ years building and operating production distributed systems or infrastructure software
  • Production-quality Go and Python skills
  • Experience writing or debugging Kubernetes controllers and operators
  • Knowledge of CRDs, reconciliation semantics, informer caches, admission webhooks, and RBAC
  • Debugging skills across distributed systems, Linux, and networking
  • Experience with Prometheus, Grafana, PromQL, exporter design, and alerting
  • Active use of coding agents and rigor in verifying their output

Responsibilities

  • Build declarative CRD-driven automation for bare-metal networking, operating systems, and application software across clusters
  • Deliver push-button cluster installation, upgrades, and security patching with canaries and downtime budgets
  • Develop Kubernetes operators for scheduling large inference workloads
  • Build gRPC control-plane services, authorization, admission webhooks, and quota policies
  • Create metrics and log pipelines, exporters, SLOs, and alerting for systems, servers, and network fabric
  • Implement failure detection, highly available control planes, and automated recovery
  • Develop CLIs, APIs, and an MCP gateway for users, operators, and AI agents

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →

Текст вакансии взят без изменений

Источник -