Назад
1 день назад

Staff Software Engineer (AI Inference)

116 000 - 155 000GBP
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Staff Software Engineer (AI Inference): Driving architecture and reliability for a Kubernetes-native inference platform with an accent on request routing, adaptive scheduling, and GPU resource management. Focus on implementing advanced inference optimizations like speculative decoding and KV-cache reuse to optimize cost-per-token and P99 SLAs.

Location: London, England

Salary: £116,000 – £155,000

Company

CoreWeave is a specialized cloud platform providing massive-scale GPU infrastructure for AI/ML, VFX, and real-time inference.

What you will do

  • Lead cross-cutting design initiatives for request routing, adaptive scheduling, and GPU resource management.
  • Implement advanced inference optimizations such as speculative decoding and KV-cache reuse.
  • Establish performance benchmarking frameworks and optimize cost-per-token under strict P99 SLAs.
  • Drive architecture, performance, and reliability for the Kubernetes-native inference platform.
  • Mentor senior and mid-level engineers while raising engineering rigor and observability practices.

Requirements

  • Must be a U.S. person or eligible to access export-controlled information per U.S. Government regulations.
  • 8–12+ years of experience building large-scale distributed systems or cloud platforms.
  • Strong coding proficiency in Go, Python, or C++.
  • Deep production-scale expertise in Kubernetes orchestration, scheduling, and service design.
  • Hands-on experience with inference systems (batching, caching, memory optimization, mixed precision BF16/FP8).
  • Bachelor’s degree in Computer Science, Engineering, or a related technical field.

Nice to have

  • Contributions to modern inference frameworks like vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe.
  • Expertise in GPU systems engineering (CUDA, NCCL, RDMA, NUMA).
  • Experience with hyperscale cloud environments or large-scale AI/ML infrastructure.

Culture & Benefits

  • Comprehensive health, dental, and vision insurance (100% paid).
  • Company-paid life, short-term, and long-term disability insurance.
  • 401(k) with generous employer match and ESPP participation.
  • Flexible PTO and mental wellness benefits through Spring Health.
  • Catered daily lunch in office and data center locations.
  • Support for family-forming (Carrot) and flexible childcare (Kinside).

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →