Назад
Company hidden
3 дня назад

Senior Software Engineer (AI)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Software Engineer (AI): Building and operating software platforms for managing large fleets of AI clusters with an accent on backend services, operational tooling, and user-facing applications. Focus on designing reliable distributed systems, automating incident response, and maintaining fleet health through component and site failures.

Location: Hybrid in Sunnyvale, California, or Toronto, Canada

Company

hirify.global builds large-scale AI computing systems and software for high-speed model training and inference.

What you will do

  • Build and operate software for managing large fleets of AI clusters.
  • Develop operator-facing views of cluster health, capacity, performance, and active issues.
  • Build services and integrations that combine data and workflows across infrastructure systems.
  • Automate incident investigation, response, and service restoration workflows.
  • Design systems that remain reliable through component and data center failures as the fleet scales.
  • Lead projects from design through production, measure impact, and use operational feedback to improve the platform.

Requirements

  • 12+ years of industry experience building and operating production software for distributed systems or large-scale infrastructure.
  • Strong Go or Python skills, including service and API design.
  • Expertise in control planes, fleet management systems, or operational platforms.
  • Experience with Linux, containers, Kubernetes, asynchronous work, retries, and partial failures.
  • Experience with event streaming, workflow automation, or time-series telemetry.
  • Strong judgment in reliability, security, observability, and cross-functional technical leadership.

Nice to have

  • Experience building software for incident response, hardware health, or capacity management.
  • Experience developing dashboards and applications for infrastructure operators.
  • Experience working with AI clusters and their compute, networking, and hardware systems.

Culture & Benefits

  • Work on an AI platform designed beyond the constraints of traditional GPUs.
  • Opportunities to publish and open-source AI research.
  • Work with one of the fastest AI supercomputers in the world.
  • Startup vitality with job stability and a non-corporate work culture.
  • Continuous learning, growth, and support in an inclusive work environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →