1 месяц назад
Senior Site Reliability Engineer, AI Inference
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer, AI Inference (LLM serving and hardware optimization): Building reliable, scalable inference systems for large language models across GPU-rich data centers and resource-constrained edge devices with an accent on low latency, high throughput, and hardware acceleration. Focus on designing Kubernetes-based orchestration, optimizing inference across NVIDIA GPUs and AI accelerators, and monitoring TTFT, throughput, memory utilization, and cost efficiency during traffic spikes.
Location: Dublin, Ireland
Company
develops cybersecurity and application delivery solutions that help organizations create, secure, and operate digital applications.
What you will do
- Build and maintain high-performance LLM inference engines using vLLM, TGI, and NVIDIA Triton.
- Optimize inference deployments for low latency, high throughput, and reliable service across data centers and edge devices.
- Profile and optimize models for NVIDIA GPUs, Apple Silicon, TPUs, and LPUs in collaboration with hardware teams.
- Design auto-scaling architectures for real-time and batch inference pipelines using Kubernetes for routing and orchestration.
- Establish observability for TTFT, tokens per second, memory bandwidth, SLAs, and cost per 1K tokens.
- Build performance and load-testing suites, identify bottlenecks, and maintain reliability during traffic spikes.
Requirements
- Senior-level experience in SRE, MLOps, or high-performance AI endpoint reliability.
- Proficiency in Python, C++, Rust, or Golang for high-performance AI workflows.
- Hands-on experience with vLLM, TensorRT, Llama.cpp, and Ollama.
- Strong knowledge of Docker, Kubernetes, and cloud platforms including AWS, GCP, and Azure.
- Expertise in GPU and AI accelerator profiling and performance optimization, including NVIDIA GPUs and TPUs.
- Ability to design scalable inference systems that remain stable and performant during demand surges.
Nice to have
- Experience with LLM deployment techniques such as Speculative Decoding or PagedAttention.
- Contributions to open-source inference libraries or hardware-level kernel development, including CUDA or Triton kernels.
- Experience optimizing high-throughput inference environments for traffic bursts.
Culture & Benefits
- Work with advanced AI optimization, hardware acceleration, and real-time AI technologies.
- Collaborate across multidisciplinary and cross-functional teams.
- Contribute to enterprise AI systems focused on scalability, reliability, and cost efficiency.
- Participate in knowledge sharing and long-term architectural improvements.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →