Назад
3 дня назад

AI Fleet Platform Software Engineer

Формат работы
hybrid
Тип работы
fulltime
Грейд
head
Страна
US/Canada
vacancy_detail.hirify_telegram_tooltipВакансия из Telegram канала -

Мэтч & Сопровод

Покажет вашу совместимость и напишет письмо

Описание вакансии

TL;DR
AI Fleet Platform Software Engineer (AI clusters/Go/Python): Building and operating software that manages large fleets of AI clusters with an accent on control planes, fleet management, observability, and operational automation. Focus on designing reliable distributed systems, handling asynchronous workflows and partial failures, and automating incident investigation and service restoration.

AI Fleet Platform Software Engineer

Company

Cerebras Systems, Inc.

Conditions

3 days ago

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build and operate software that manages large fleets of AI clusters. You will create services, integrations, operational tools, and user-facing applications that help operators monitor health, capacity, performance, and incidents. You will lead projects from design through production, automate operational workflows, and improve reliability as the fleet grows.

Requirements

  • 12+ years of industry experience building and operating production software for distributed systems or large-scale infrastructure
  • Strong Go or Python skills
  • Experience designing services and APIs
  • Expertise in control planes, fleet management systems, or operational platforms
  • Experience with Linux, containers, Kubernetes, and distributed-system failures
  • Experience designing for asynchronous work, retries, and partial failures
  • Experience with event streaming, workflow automation, or time-series telemetry
  • Strong judgment in reliability, security, and observability
  • Ability to lead ambiguous projects and collaborate across engineering and operations teams

Responsibilities

  • Build and operate software for managing large fleets of AI clusters
  • Provide operators with actionable views of cluster health, capacity, performance, and issues
  • Develop services and integrations across infrastructure systems
  • Automate incident investigation and service-restoration workflows
  • Design reliable systems that withstand component and site failures
  • Gather platform-user needs and make practical product and engineering decisions
  • Lead projects from design through production and use operational feedback to improve them

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →

Текст вакансии взят без изменений

Источник -