Назад
Company hidden
6 дней назад

Senior Cloud Support Engineer (AI Infrastructure)

145 000 - 175 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Cloud Support Engineer (AI Infrastructure): Providing technical support for sustainable GPU cloud infrastructure used in AI/ML, physics simulations, and computational biology with an accent on Linux troubleshooting, HPC technologies, and customer success. Focus on diagnosing VM and hardware failures, managing alert triage and maintenance testing, and collaborating with SRE, Networking, and Storage teams on root cause analysis.

Location: On-site in New York, NY or Denver, CO, US

Salary: $145,000–$175,000 annually plus bonus; restricted stock units included.

Company

hirify.global builds vertically integrated AI infrastructure, including energy systems, data centers, and cloud services for demanding AI workloads.

What you will do

  • Provide technical customer support through Zendesk while meeting SLAs and maintaining 95%+ CSAT.
  • Participate in a 24/7 on-call rotation and resolve critical incidents.
  • Diagnose VM, hardware, and scaling-test issues using Linux CLI and internal tools.
  • Manage alert triage, maintenance-window preparation, and node delivery testing.
  • Collaborate with SRE, Networking, and Storage teams from initial triage through root cause analysis.
  • Create onboarding materials, knowledge-base documentation, and standard operating procedures.

Requirements

  • Bachelor's degree in IT, Computer Science, Engineering, or a related field, or 4+ years of equivalent technical experience.
  • 5+ years of customer support experience, ideally in cloud, storage, or networking environments.
  • Strong Linux command-line interface skills and proficiency with Git.
  • Experience with Kubernetes, Slurm or Terraform, and monitoring tools such as Grafana.
  • Familiarity with AWS, Azure, or GCP and strong customer communication skills.
  • Understanding of HPC technologies including InfiniBand, RDMA, RoCE, and software-defined networking.

Nice to have

  • CKA, CKAD, CKS, KCNA, AWS, NVIDIA, InfiniBand, Linux Foundation, or system administration certifications.
  • Deep expertise in cloud platforms, automation tools, and scripting languages.
  • Experience mentoring, training, and onboarding colleagues.
  • Interest in using technology to support a more sustainable future.

Culture & Benefits

  • Restricted stock units and competitive compensation.
  • Paid time off, paid holidays, and paid parental leave.
  • Health, dental, vision, life, and disability insurance, plus HSA contributions.
  • Professional development, tuition reimbursement, and mental health and wellness support.
  • 401(k) retirement plan with company matching up to 4% of salary.
  • Commuter benefits, cell phone stipend, and volunteer time off.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →