Назад
Company hidden
3 дня назад

Senior Infrastructure Operations Engineer (AI)

116 000 - 128 000GBP
Формат работы
remote (только United_kingdom)/hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Infrastructure Operations Engineer (AI): Operating and scaling large-scale GPU infrastructure platforms with an accent on Linux systems, bare metal environments, provisioning, observability, and reliability automation. Focus on designing resilient platforms, responding to incidents, automating customer provisioning, and reducing operational toil across complex infrastructure.

Location: Fully remote within the UK, or hybrid from the London office hub; occasional team and company offsites. Work visa sponsorship is not available.

Annual base salary: £116,000–£128,000 GBP

Company

hirify.global builds an end-to-end platform for developing, training, and deploying AI systems, combining developer-focused software with large-scale AI compute infrastructure.

What you will do

  • Design, build, and roll out platforms and operational patterns that reduce incidents and support customer-facing and internal features.
  • Deploy infrastructure updates and improvements for internal and end-customer use cases.
  • Operate large-scale GPU environments, Linux systems, bare metal infrastructure, and provisioning workflows.
  • Improve reliability, observability, and automation while reducing manual operational work.
  • Collaborate with Infrastructure Engineering, Network Operations, Customer Success, and Software and Platform Development teams.
  • Participate in a primary/secondary on-call rotation.

Requirements

  • 8+ years of experience with Linux as a server or hosting platform; Ubuntu experience is beneficial.
  • 5+ years of experience with AWS.
  • At least 2 years of experience with Kubernetes, container fundamentals, Terraform, and Ansible.
  • Experience managing network-attached storage using NFS, Ceph, or similar protocols.
  • Experience with monitoring systems such as Prometheus and the ELK stack, plus familiarity with GitOps workflows.
  • Software development experience with Python, Go, Bash, or other languages for automation and system or API integration.

Nice to have

  • Experience troubleshooting and provisioning bare-metal hardware, including Dell systems.
  • Experience with GPU servers in bare-metal or virtualized environments.
  • Experience with datacenter networks, 400Gb Ethernet, InfiniBand, network switches, routers, or firewalls.
  • Experience with SONiC switches, Palo Alto firewalls, Juniper Networks, or VAST storage systems.

Culture & Benefits

  • Ownership-focused environment with open communication, urgency, and continuous improvement.
  • Medical, dental, and vision coverage for employees and eligible dependents.
  • Equity, retirement or pension contributions, unlimited PTO, company holidays, and a two-week winter break.
  • Paid parental and family leave, professional development allowance, wellness benefits, and work-from-home stipends.
  • Four-week paid sabbatical after four years of service, flexible schedules, and office meals.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →