Назад
Company hidden
1 день назад

Infrastructure Operations Engineer (AI)

160 000 - 200 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK/US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Infrastructure Operations Engineer (AI infrastructure): Building and operating large-scale GPU infrastructure platforms with an accent on Linux systems, provisioning workflows, reliability, and automation. Focus on designing incident-reducing platforms, automating infrastructure operations, and troubleshooting complex bare-metal, networking, storage, and containerized environments.

Location: Based in New York City, San Francisco, Seattle, or London, with a minimum of 2 in-office days per week; occasional team and company offsites.

Annual base salary: $160,000–$200,000 USD, plus discretionary bonus, equity, and benefits.

Company

hirify.global builds an end-to-end platform for developing, training, and deploying AI systems, combining developer-first software with large-scale GPU infrastructure.

What you will do

  • Design, build, and deploy platforms and operational patterns that reduce incidents and enable customer-facing and internal features.
  • Deploy updates and improvements for internal and end-customer infrastructure use cases.
  • Operate large-scale GPU environments, Linux systems, bare-metal infrastructure, and provisioning workflows.
  • Participate in break/fix operations, incident response, customer provisioning, observability, and the on-call rotation.
  • Collaborate with Infrastructure Engineering, Network Operations, Customer Success, and Software and Platform Development teams.

Requirements

  • 8+ years of experience with Linux as a server or hosting platform.
  • 5+ years of experience with AWS.
  • At least 2 years of experience with Kubernetes, container fundamentals, Terraform, and Ansible.
  • At least 2 years of experience managing network-attached storage using NFS, Ceph, or similar protocols.
  • Experience with Prometheus, the ELK stack, GitOps workflows, and automation using Python, Go, Bash, or other languages.
  • Deep networking fundamentals and experience building complex systems, with strong written and oral communication.

Nice to have

  • Experience troubleshooting and provisioning bare-metal hardware, including Dell hardware.
  • Experience with GPU servers, datacenter-level networks, 400Gb Ethernet, or InfiniBand.
  • Experience with VAST storage, SONiC switches, Palo Alto firewalls, or Juniper Networks equipment.

Culture & Benefits

  • Hybrid work model with flexible schedules for office-based teams.
  • Medical, dental, and vision coverage, with benefits varying by location, team, and role.
  • RSUs, 401(k) matching in the U.S., and pension contributions in the U.K.
  • Unlimited PTO, company holidays, floating holidays, and a two-week winter company closure.
  • Paid parental and family leave, professional development allowance, wellness and work-from-home stipends, and complimentary office meals.
  • Four weeks of paid sabbatical leave after four years of service.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →