9 дней назад
Senior Platform Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Platform Engineer (AI) (GCP/Kubernetes/GPU): Building and operating a GPU platform and self-service compute infrastructure for AI model training and serving with an accent on Kubernetes, infrastructure-as-code, GitOps, and high-performance workloads. Focus on designing reliable production clusters, automating deployment workflows, controlling cloud costs, and enabling data science teams to move from proof of concept to production.
Location: Paris, France; office-based
Company
A global gaming organization creating original interactive entertainment experiences across major studios and franchises.
What you will do
- Design, build, and operate a GPU platform on GCP for AI workloads using infrastructure-as-code.
- Deliver self-service compute for data science teams, including Ray on Kubernetes, job submission, queuing, quotas, and cost visibility.
- Establish golden paths for deploying services so product teams can ship without performing recurring operations manually.
- Support the deployment, hosting, and production operation of machine learning models.
- Control infrastructure costs across GPUs, cloud resources, and related services.
- Collaborate with data science teams from proof of concept through production deployment while improving system performance, reliability, and efficiency.
Requirements
- Significant experience as a Platform, DevOps, or MLOps Engineer, with strong software engineering fundamentals and distributed systems knowledge.
- Hands-on experience with Kubernetes, Terraform, GitOps, CRDs and operators, scheduling, autoscaling, node pool design, state management, ArgoCD, and Helm.
- Solid experience with GCP; AWS experience is an advantage.
- Platform-as-a-product mindset, autonomy, comfort with ambiguity, and ability to treat data scientists as platform users.
- Fluent English, written and spoken, is required.
Nice to have
- Production experience running or optimizing GPU-based workloads for machine learning training or inference.
- Experience with distributed job scheduling such as KubeRay, Slurm, or Kubeflow.
- Experience with cross-cloud networking between AWS and GCP.
- Familiarity with MLOps tools such as MLflow, Weights & Biases, and model registries.
Culture & Benefits
- Permanent full-time employment in an office-based environment.
- Access to an internal e-learning platform from the first day.
- Access to a game library, consoles, competitor games, and board games.
- Works council discounts for entertainment, fitness, cultural activities, and leisure.
- Career and development planning after one year, plus clubs, sports activities, and organized weekends.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →