Назад
Company hidden
2 месяца назад

Performance Engineer (Containers/Serverless)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Finland
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Performance Engineer (Containers/Serverless): Profiling and optimizing containerized ML/AI workloads across image distribution, runtime startup, storage, model loading, and GPU execution with an accent on cold-start latency, inference throughput, and training performance. Focus on designing caching and storage layers, benchmarking real workloads, and removing bottlenecks across networking, filesystems, container runtimes, and GPUs.

Location: Helsinki, Finland; hybrid work with presence in the Helsinki office at least 3 days per week

Company

hirify.global builds cloud infrastructure for AI, operating GPU clusters across Europe, the US, and Asia for demanding production ML workloads.

What you will do

  • Profile and optimize the end-to-end execution path for containerized ML/AI workloads, including image distribution, runtime startup, weight loading, and inference.
  • Design and tune storage layers between S3-compatible object stores and GPU nodes, including prefetchers, caching tiers, network filesystems, and local NVMe layouts.
  • Improve time-to-first-token, training step time, inference throughput, and cold-start latency for internal and customer workloads.
  • Benchmark real workloads and translate findings into measurable platform improvements.
  • Collaborate across compute, networking, and platform teams to remove end-to-end bottlenecks.
  • Document performance findings and trade-offs in internal and occasional external technical write-ups.

Requirements

  • Production experience making ML/AI workloads measurably faster.
  • Strong knowledge of Linux storage and container runtime internals.
  • Hands-on experience with at least one distributed or network filesystem and S3-compatible object storage.
  • Experience with object-storage performance issues such as small-object overhead, range requests, eventual consistency, and multipart tuning.
  • Familiarity with model-serving runtimes such as vLLM or SGLang and formats including safetensors, GGUF, and sharded checkpoints.
  • Ability to reason across NICs, switches, filesystems, caches, container runtimes, and GPUs as one system.

Nice to have

  • Systems programming experience with Go, Rust, or Python.
  • Experience with checkpoint/restore, container image acceleration, caching layers, RDMA, GPUDirect Storage, or NVMe-oF.
  • Experience with serverless GPU platforms, model registries, or Kubernetes-based ML infrastructure.
  • Open-source contributions to storage, ML runtime, container, or kernel projects.
  • Bare-metal performance engineering experience.

Culture & Benefits

  • Full-time, permanent employment.
  • Cash and equity compensation.
  • Healthcare, lunch, wellbeing, and other fringe benefits.
  • Access to real GPUs for performance testing.
  • Low-hierarchy environment with pragmatic delivery and engineering ownership.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →