20 часов назад
Software Engineer (Compute Platform), London
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineer (Compute Platform) (AI/ML infrastructure): Building and operating large-scale GPU/TPU compute infrastructure for foundation models in biotech with an accent on accelerator strategy, cloud systems, and deployment reliability. Focus on designing and optimizing Kubernetes-based cluster environments, integrating new hardware, and improving the efficiency and reliability of machine learning runs.
Location: London, United Kingdom; hybrid work with attendance at the office 3 days a week
Company
develops AI models and drug design systems to accelerate scientific discovery and the development of medicines.
What you will do
- Define the end-to-end GPU and TPU accelerator strategy, including infrastructure design, performance optimization, hardware integration, and cluster deployment.
- Build, monitor, manage, and support research, development, and production cloud infrastructure and systems.
- Drive infrastructure efficiency improvements up to the service layer used by machine learning platform teams.
- Improve the reliability and consistency of large-scale machine learning runs.
- Support hardware acquisition and deployment decisions and contribute to tooling, infrastructure, and architecture choices.
- Collaborate with machine learning, science, research, product, business development, and operations teams.
Requirements
- Real-world experience with large-scale AI and machine learning workloads.
- Experience designing cloud compute infrastructure, preferably on GCP.
- Strong programming skills.
- Significant experience deploying and working with Kubernetes.
- Familiarity with Nvidia GPU generations.
- Ability to work in a hybrid model and attend the London office 3 days per week.
Nice to have
- Background in machine learning software engineering or infrastructure SRE.
- Experience leading and delivering projects with multidisciplinary stakeholders.
- Familiarity with Google TPU generations, workload scheduling, hardware benchmarking, and machine learning efficiency research.
- Understanding of machine learning-driven research and development cycles.
Culture & Benefits
- Collaborative interdisciplinary environment spanning science, research, engineering, product, and operations.
- Values centered on thoughtful work, initiative, integrity, determination, and collaboration.
- Shared learning and an environment designed to support employees and diverse perspectives.
- Hybrid working model with regular in-person collaboration.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
IBM Watson
21 час назад
Software Engineer, Kubernetes Platform (Kubernetes)
DeepL
5 дней назад
Platform Engineer (AI)
Elliptic
6 дней назад
Lead DevOps Engineer (Web3)
20 часов назад
Infrastructure Engineer (AI)
Anthropic
3 дня назад
Senior & Staff Software Engineer, Continuous Integration (AI)
325 000 - 390 000GBP
6 дней назад
Senior Technical Implementation Engineer (Kubernetes)
70 000 - 90 000GBP