Назад
Company hidden
1 день назад

Technical Support Engineer (AI Infrastructure)

Формат работы
onsite
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Technical Support Engineer (AI Infrastructure): Supporting enterprise customers running workloads on GPU and AI infrastructure, troubleshooting compute, networking, platform, scheduling, and performance issues with an accent on incident response and customer communication. Focus on diagnosing infrastructure problems, escalating hardware and data-center issues, improving runbooks and tooling, and providing reliable 24/7 coverage.

Location: San Francisco, California, United States

Company

hirify.global builds and operates GPU data centers and supercomputers, leasing compute infrastructure and enabling customers to sublease capacity.

What you will do

  • Own inbound support tickets and live escalations for enterprise customers running workloads on GPU and AI infrastructure.
  • Triage and resolve issues across compute, networking, platform, scheduling, orchestration, and performance layers.
  • Use monitoring and alerting tools to diagnose incidents and act as a first responder through resolution.
  • Escalate hardware, data-center, facility, and colocation issues with clear documentation to engineering, partners, and hardware vendors.
  • Communicate status and expectations to customers and coordinate handoffs across time zones.
  • Improve runbooks, SOPs, knowledge-base content, and support workflows while tracking CSAT, first-response, and resolution metrics.

Requirements

  • 3–5 years of experience in technical support, customer support engineering, or a similar customer-facing technical role.
  • Experience troubleshooting infrastructure-level issues with Linux administration, basic shell or Python scripting, and monitoring or observability tools such as Prometheus, Grafana, or Datadog.
  • Ability to explain technical problems clearly to customers and internal engineering teams.
  • Excellent written communication skills and the ability to remain calm during outages or escalations.
  • Availability for shift-based work in a 24/7 coverage model, including occasional after-hours, weekend, and holiday incident coverage.
  • Ability to build support processes and coverage from scratch and maintain a strong focus on customer outcomes.

Nice to have

  • Experience with GPU or AI infrastructure, specialized hardware, or managed services.
  • Familiarity with data-center or colocation operations.
  • Experience with support tooling, AI-assisted workflows, scripting, diagnostics automation, or incident management.

Culture & Benefits

  • Visa and work permit sponsorship available.
  • Competitive salary and company equity.
  • 401(k) matching up to 4%.
  • Medical, dental, and vision insurance for employees and dependents, with 100% of premiums covered.
  • Unlimited paid time off, 10+ observed holidays, and paid leave for biological, adoptive, and foster parents.
  • Daily lunch and an office book allowance.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →