1 день назад
Infrastructure Operations Engineer (AI)
160 000 - 200 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Infrastructure Operations Engineer (AI infrastructure): Building and operating large-scale GPU infrastructure platforms with an accent on Linux systems, provisioning workflows, reliability, and automation. Focus on designing incident-reducing platforms, automating infrastructure operations, and troubleshooting complex bare-metal, networking, storage, and containerized environments.
Location: Based in New York City, San Francisco, Seattle, or London, with a minimum of 2 in-office days per week; occasional team and company offsites.
Annual base salary: $160,000–$200,000 USD, plus discretionary bonus, equity, and benefits.
Company
builds an end-to-end platform for developing, training, and deploying AI systems, combining developer-first software with large-scale GPU infrastructure.
What you will do
- Design, build, and deploy platforms and operational patterns that reduce incidents and enable customer-facing and internal features.
- Deploy updates and improvements for internal and end-customer infrastructure use cases.
- Operate large-scale GPU environments, Linux systems, bare-metal infrastructure, and provisioning workflows.
- Participate in break/fix operations, incident response, customer provisioning, observability, and the on-call rotation.
- Collaborate with Infrastructure Engineering, Network Operations, Customer Success, and Software and Platform Development teams.
Requirements
- 8+ years of experience with Linux as a server or hosting platform.
- 5+ years of experience with AWS.
- At least 2 years of experience with Kubernetes, container fundamentals, Terraform, and Ansible.
- At least 2 years of experience managing network-attached storage using NFS, Ceph, or similar protocols.
- Experience with Prometheus, the ELK stack, GitOps workflows, and automation using Python, Go, Bash, or other languages.
- Deep networking fundamentals and experience building complex systems, with strong written and oral communication.
Nice to have
- Experience troubleshooting and provisioning bare-metal hardware, including Dell hardware.
- Experience with GPU servers, datacenter-level networks, 400Gb Ethernet, or InfiniBand.
- Experience with VAST storage, SONiC switches, Palo Alto firewalls, or Juniper Networks equipment.
Culture & Benefits
- Hybrid work model with flexible schedules for office-based teams.
- Medical, dental, and vision coverage, with benefits varying by location, team, and role.
- RSUs, 401(k) matching in the U.S., and pension contributions in the U.K.
- Unlimited PTO, company holidays, floating holidays, and a two-week winter company closure.
- Paid parental and family leave, professional development allowance, wellness and work-from-home stipends, and complimentary office meals.
- Four weeks of paid sabbatical leave after four years of service.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Nscale
1 день назад
Infrastructure Software Engineer (AI)
150 000 - 215 000$
6 часов назад
Senior SRE (GPU Infrastructure)
168 000 - 252 000$
3 дня назад
Infrastructure Engineer (AI)
128 000 - 189 000$
18 часов назад
Infrastructure Engineer
4 дня назад
Platform Engineer II (Cloud Infrastructure)
115 000 - 130 000$
4 дня назад
SRE Engineer II (Cloud/DevOps)
141 000 - 162 000$