AI Platform Support Engineer (EMEA)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Location: Hybrid from the London office, with at least 2 in-office days per week. EMEA shifts run from 9AM–7PM CET/CEST on either Saturday–Tuesday or Thursday–Sunday. Occasional team and company offsites are required. Visa sponsorship is not available.
Annual base salary: £75,000–£95,000 GBP, plus discretionary bonus, equity, and benefits.
Company
builds an end-to-end platform for developing, training, and deploying AI systems, combining developer-first software with large-scale AI compute.
What you will do
- Partner directly with customer ML engineering teams running production training and inference workloads.
- Diagnose distributed systems and ML infrastructure issues involving Kubernetes, GPU allocation, networking, storage, and inference serving.
- Troubleshoot PyTorch, CUDA, NCCL, containerized workloads, and multi-node GPU systems.
- Analyze logs, metrics, traces, and system behavior to identify root causes and performance bottlenecks.
- Drive reliability improvements through post-incident reviews, observability enhancements, documentation, runbooks, and automation.
- Collaborate with infrastructure, networking, and platform engineering teams to improve customer troubleshooting workflows.
Requirements
- Strong software engineering and systems troubleshooting experience.
- Experience with Kubernetes, containers, cloud infrastructure, and distributed systems.
- Linux expertise covering networking, storage, process management, and performance tuning.
- Experience operating and troubleshooting production or research machine learning workloads, including distributed ML systems such as PyTorch, CUDA, or NCCL.
- Experience with GPU infrastructure, orchestration, observability, and debugging tools such as Prometheus, Grafana, or OpenTelemetry.
- Strong communication skills and the ability to work directly with technical customers and engineering teams.
Nice to have
- Experience with large-scale model training, distributed inference, Ray, Kubeflow, Slurm, or similar scheduling platforms.
- Knowledge of InfiniBand, RDMA, high-performance networking, bare-metal infrastructure, or ML storage systems.
- Experience at an AI infrastructure, cloud, MLOps, or developer tooling company.
- Contributions to platform engineering, developer infrastructure, or operational tooling projects.
- Experience writing automation, tooling, or scripts in Python or similar languages.
Culture & Benefits
- Fast-moving environment focused on ownership, open communication, continuous improvement, and long-term scalable systems.
- Medical, dental, and vision coverage for employees and eligible dependents.
- Equity, retirement savings support, unlimited PTO, company holidays, and a two-week winter break.
- Paid parental and family leave, professional development allowance, wellness benefits, and work-from-home stipends.
- Four weeks of paid sabbatical after four years of service, flexible schedules, and complimentary office meals.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →