Назад
Company hidden
3 дня назад

Infrastructure Software Engineering – Platform & Build (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Infrastructure Platform Engineer (AI hardware and build systems): Building and scaling CI infrastructure, observability systems, and a multi-language Bazel monorepo that supports AI hardware, software, and research workloads with an accent on reliability, build performance, and developer productivity. Focus on designing reproducible pipelines, debugging failures across large compute clusters, and operating infrastructure with SRE practices.

Location: London, United Kingdom

Company

hirify.global is an AI hardware startup developing chips and systems for high-speed inference of frontier AI models.

What you will do

  • Own CI infrastructure end to end, including reproducible multi-language pipelines, debugging, and performance optimization across large compute clusters.
  • Build and maintain observability, alerting, runbooks, and incident response workflows for core infrastructure.
  • Scale and maintain a multi-language Bazel monorepo supporting Python, C++, Rust, SystemVerilog, and machine learning workloads.
  • Improve the infrastructure and platform experience for engineers working on hardware design, verification, kernel development, ML compilers, and simulators.
  • Apply an SRE mindset to reliability, build performance, fault isolation, and root cause analysis.

Requirements

  • 5+ years of experience in software, platform, or infrastructure engineering.
  • 5+ years of hands-on experience with CI/CD for large-scale products, including monitoring, performance debugging, and resolving pipeline failures.
  • 3+ years of experience with build systems such as Bazel, Buck, Pants, or similar cross-language systems.
  • Experience with infrastructure as code using Terraform, OpenTofu, or Pulumi.
  • Experience with monitoring and observability tools such as Prometheus or Grafana.
  • Working knowledge of DevOps and SRE practices, including SLOs, alerting, incident management, and on-call work, plus strong proficiency in Python, Go, or Rust.

Culture & Benefits

  • Meaningful equity and ownership participation.
  • Private medical, dental, and vision coverage.
  • Contributory pension.
  • 25 days of holiday plus bank holidays.
  • Life and critical illness insurance.
  • Inclusive and diverse office environment.

Hiring process

  • Some roles may require additional eligibility checks related to UK, US, and international export control regulations.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →