3 дня назад
Senior Infrastructure Operations Engineer (AI)
116 000 - 128 000GBP
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Infrastructure Operations Engineer (AI): Operating and scaling large-scale GPU infrastructure platforms with an accent on Linux systems, bare metal environments, provisioning, observability, and reliability automation. Focus on designing resilient platforms, responding to incidents, automating customer provisioning, and reducing operational toil across complex infrastructure.
Location: Fully remote within the UK, or hybrid from the London office hub; occasional team and company offsites. Work visa sponsorship is not available.
Annual base salary: £116,000–£128,000 GBP
Company
builds an end-to-end platform for developing, training, and deploying AI systems, combining developer-focused software with large-scale AI compute infrastructure.
What you will do
- Design, build, and roll out platforms and operational patterns that reduce incidents and support customer-facing and internal features.
- Deploy infrastructure updates and improvements for internal and end-customer use cases.
- Operate large-scale GPU environments, Linux systems, bare metal infrastructure, and provisioning workflows.
- Improve reliability, observability, and automation while reducing manual operational work.
- Collaborate with Infrastructure Engineering, Network Operations, Customer Success, and Software and Platform Development teams.
- Participate in a primary/secondary on-call rotation.
Requirements
- 8+ years of experience with Linux as a server or hosting platform; Ubuntu experience is beneficial.
- 5+ years of experience with AWS.
- At least 2 years of experience with Kubernetes, container fundamentals, Terraform, and Ansible.
- Experience managing network-attached storage using NFS, Ceph, or similar protocols.
- Experience with monitoring systems such as Prometheus and the ELK stack, plus familiarity with GitOps workflows.
- Software development experience with Python, Go, Bash, or other languages for automation and system or API integration.
Nice to have
- Experience troubleshooting and provisioning bare-metal hardware, including Dell systems.
- Experience with GPU servers in bare-metal or virtualized environments.
- Experience with datacenter networks, 400Gb Ethernet, InfiniBand, network switches, routers, or firewalls.
- Experience with SONiC switches, Palo Alto firewalls, Juniper Networks, or VAST storage systems.
Culture & Benefits
- Ownership-focused environment with open communication, urgency, and continuous improvement.
- Medical, dental, and vision coverage for employees and eligible dependents.
- Equity, retirement or pension contributions, unlimited PTO, company holidays, and a two-week winter break.
- Paid parental and family leave, professional development allowance, wellness benefits, and work-from-home stipends.
- Four-week paid sabbatical after four years of service, flexible schedules, and office meals.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Senior Cloud Operations Engineer (AWS)
4 дня назад
Senior DevOps Engineer (AWS)
50 000 - 65 000GBP
9 дней назад
Senior Systems Engineer (AWS)
5 дней назад
Senior Infrastructure Engineer (AWS/Terraform)
7 дней назад
Software Engineer (Application Engineer or Cloud Engineer)
60 000GBP
3 дня назад