Назад
Company hidden
5 дней назад

Senior Cloud Support Engineer (AI Infrastructure)

105 000 - 125 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Cloud Support Engineer (AI Infrastructure): Supporting customers using sustainable GPU cloud infrastructure for AI/ML, physics simulations, and computational biology with an accent on Linux troubleshooting, Kubernetes, HPC technologies, and customer success. Focus on diagnosing VM and hardware issues, resolving critical incidents through 24/7 on-call rotation, and collaborating with SRE, networking, and storage teams on root cause analysis.

Location: Denver, Colorado, United States; on-site

Salary: $105,000–$125,000 per year plus bonus; restricted stock units included.

Company

hirify.global builds vertically integrated energy and AI infrastructure, including sustainable GPU cloud computing for demanding workloads.

What you will do

  • Provide technical customer support through Zendesk while meeting service-level agreements and maintaining 95%+ customer satisfaction.
  • Participate in a 24/7 on-call rotation and resolve critical incidents.
  • Diagnose and resolve virtual machine, hardware, scaling, and cloud infrastructure issues using command-line and internal tools.
  • Manage alert triage, maintenance-window preparation, and node delivery testing.
  • Collaborate with SRE, networking, and storage teams from initial triage through root cause analysis.
  • Create onboarding materials, knowledge base documentation, and standard operating procedures.

Requirements

  • Bachelor’s degree in IT, computer science, engineering, or a related field, or 4+ years of equivalent technical experience.
  • 2–5 years of customer support experience, preferably in cloud, storage, or networking environments.
  • Strong Linux command-line skills and proficiency with Git.
  • Experience with Kubernetes, Slurm, Terraform, Grafana, and other cloud or workload-management tools.
  • Familiarity with AWS, Azure, or GCP and strong communication skills for prioritizing escalations.
  • Understanding of HPC technologies including InfiniBand, RDMA, RoCE, and software-defined networking.

Nice to have

  • Cloud, Kubernetes, AWS, NVIDIA, Linux Foundation, or InfiniBand certifications.
  • Automation and scripting experience.
  • Experience mentoring, training, and onboarding colleagues.
  • Interest in using technology to support a more sustainable future.

Culture & Benefits

  • Competitive compensation, bonus, equity, and restricted stock units.
  • Paid time off, holidays, leave programs, and paid parental leave.
  • Health, dental, vision, life, disability, and mental health coverage.
  • Employer HSA contributions and a 401(k) plan with a company match of up to 4% of salary.
  • Professional development, tuition reimbursement, commuter benefits, daily meal allowance, and cell phone stipend.
  • Global travel insurance, emergency assistance, volunteer time off, and location-specific programs.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →