Назад
Company hidden
обновлено 14 часов назад

Member of Technical Staff, AI Infrastructure Team (AI)

Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
UK/Finland
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff, AI Infrastructure Team (AI): Improving networking and communication layers for large-scale LLM training across distributed GPU clusters with an accent on collective communication performance, benchmarking, and reliability. Focus on identifying NCCL bottlenecks, profiling distributed training workloads, building regression infrastructure, and implementing fault-tolerance mechanisms to improve cluster utilization.

Location: Helsinki, Finland or London, UK. Hybrid work from the Helsinki or London office three days per week.

Company

hirify.global is a full-stack AI cloud company that installs, operates, and optimizes compute infrastructure for AI model training and inference.

What you will do

  • Improve networking and communication layers for large-scale LLM training workloads.
  • Optimize collective communication across distributed GPU clusters to improve throughput, utilization, and reliability.
  • Debug networking-stack bottlenecks and build tooling for benchmarking, profiling, and regression testing.
  • Collaborate with training, infrastructure, hardware, and networking teams to improve workload scaling across clusters.
  • Build tools and dashboards for identifying NCCL bottlenecks, slow ranks, communication tail latency, and training network health.
  • Implement fault-tolerance mechanisms that reduce cluster idle time and improve training efficiency.

Requirements

  • Experience with distributed systems, networking, or large-scale ML training infrastructure.
  • Experience with communication libraries such as NCCL, MPI, NVSHMEM, or similar technologies.
  • Experience with profiling and debugging tools such as Nsight Systems, NCCL logs, PyTorch Profiler, or perf.
  • Strong systems thinking and ability to analyze performance bottlenecks across distributed environments.
  • Ability to independently define and drive technical projects.
  • Curiosity about low-level systems, networking, and large-scale AI infrastructure.

Culture & Benefits

  • Full-time, permanent employment.
  • Highly collaborative, research-adjacent work across technical teams.
  • Opportunity to contribute while the AI cloud infrastructure is still being built.

Hiring process

  • Submit an application through the Careers page; applications sent by email are not accepted.
  • Hiring proceeds without an artificial deadline and moves forward when a suitable candidate is found.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →