2 месяца назад
Performance Engineer (Containers/Serverless)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Performance Engineer (Containers/Serverless): Profiling and optimizing containerized ML/AI workloads across image distribution, runtime startup, storage, model loading, and GPU execution with an accent on cold-start latency, inference throughput, and training performance. Focus on designing caching and storage layers, benchmarking real workloads, and removing bottlenecks across networking, filesystems, container runtimes, and GPUs.
Location: Helsinki, Finland; hybrid work with presence in the Helsinki office at least 3 days per week
Company
builds cloud infrastructure for AI, operating GPU clusters across Europe, the US, and Asia for demanding production ML workloads.
What you will do
- Profile and optimize the end-to-end execution path for containerized ML/AI workloads, including image distribution, runtime startup, weight loading, and inference.
- Design and tune storage layers between S3-compatible object stores and GPU nodes, including prefetchers, caching tiers, network filesystems, and local NVMe layouts.
- Improve time-to-first-token, training step time, inference throughput, and cold-start latency for internal and customer workloads.
- Benchmark real workloads and translate findings into measurable platform improvements.
- Collaborate across compute, networking, and platform teams to remove end-to-end bottlenecks.
- Document performance findings and trade-offs in internal and occasional external technical write-ups.
Requirements
- Production experience making ML/AI workloads measurably faster.
- Strong knowledge of Linux storage and container runtime internals.
- Hands-on experience with at least one distributed or network filesystem and S3-compatible object storage.
- Experience with object-storage performance issues such as small-object overhead, range requests, eventual consistency, and multipart tuning.
- Familiarity with model-serving runtimes such as vLLM or SGLang and formats including safetensors, GGUF, and sharded checkpoints.
- Ability to reason across NICs, switches, filesystems, caches, container runtimes, and GPUs as one system.
Nice to have
- Systems programming experience with Go, Rust, or Python.
- Experience with checkpoint/restore, container image acceleration, caching layers, RDMA, GPUDirect Storage, or NVMe-oF.
- Experience with serverless GPU platforms, model registries, or Kubernetes-based ML infrastructure.
- Open-source contributions to storage, ML runtime, container, or kernel projects.
- Bare-metal performance engineering experience.
Culture & Benefits
- Full-time, permanent employment.
- Cash and equity compensation.
- Healthcare, lunch, wellbeing, and other fringe benefits.
- Access to real GPUs for performance testing.
- Low-hierarchy environment with pragmatic delivery and engineering ownership.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →