4 дня назад
Senior AI Platform Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior AI Platform Engineer (LLM/AWS): Defining and delivering the infrastructure strategy for large-scale LLM programmes across compute, data, model lifecycle management, evaluation frameworks, and platform engineering with an accent on GPU infrastructure, distributed training, and production-ready AI systems. Focus on optimising training and inference workloads, building reproducible model and data lifecycle capabilities, evolving knowledge graph infrastructure, and coordinating research, product, MLOps, and infrastructure delivery.
Location: London, United Kingdom
Company
provides clinical research services, commercial insights, and healthcare intelligence for the life sciences and healthcare industries.
What you will do
- Own the AI platform and infrastructure roadmap for large language model initiatives.
- Design and deliver high-performance compute environments across AWS and on-premises platforms, including GPU infrastructure and Slurm clusters.
- Optimise LLM training and inference workloads for performance, scalability, and reliability.
- Build model and data lifecycle capabilities covering dataset versioning, lineage, reproducibility, and model registries.
- Lead the evolution and integration of knowledge graph infrastructure with AI workflows.
- Coordinate AI Research, Data Engineering, MLOps, Product, and Infrastructure teams while mentoring engineers and guiding technology partnerships.
Requirements
- Significant experience building and operating large-scale AI, machine learning, or distributed computing platforms in enterprise environments.
- Deep understanding of LLM architectures, GPU infrastructure, CUDA, cuDNN, NCCL, PyTorch, and distributed training.
- Experience with tensor, pipeline, data, and expert parallelism, plus high-performance serving frameworks such as vLLM, TensorRT-LLM, NVIDIA NIM, or SGLang.
- Expertise in quantisation, mixed precision, FP8, GPTQ, AWQ, LoRA, GPU profiling, and performance tuning.
- Strong background in AWS, high-performance computing, distributed systems, containers, infrastructure automation, Kubernetes, Slurm, Ray, or equivalent technologies.
- Proven ability to lead complex cross-functional initiatives and bridge research and production while maintaining governance, security, and reliability.
Culture & Benefits
- Work with large healthcare datasets, advanced analytics tools, and modern AI technologies.
- Contribute to healthcare solutions intended to improve patient care and population health.
- Access career development opportunities across diverse geographies, capabilities, and therapeutic areas.
- Join an inclusive workplace that values diversity, respect, integrity, and professional growth.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →