1 месяц назад
Infrastructure Support Lead (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Infrastructure Support Lead (AI): Owning technical support for deployed AI infrastructure and building the support structure for a growing global installed base with an accent on customer incident resolution, hardware service operations, and cross-functional service delivery. Focus on designing escalation workflows, diagnosing hardware/software boundary issues, and operating support across Asia hours and international time zones.
Location: Based in Korea, with up to 25% international travel. Support coverage includes Asia hours, on-call, and out-of-hours escalation coverage across time zones.
Company
builds high-performance AI systems with full-stack software supporting PyTorch and vLLM for production inference workloads.
What you will do
- Manage the support team and maintain consistent service delivery standards.
- Act as the senior technical point of contact for customer issues, driving resolution or routing issues to the appropriate engineer.
- Design and operate the support model, including intake, triage, escalation paths, and workflows connecting customers, partners, and internal teams.
- Own internal and customer-facing knowledge bases, runbooks, troubleshooting guides, and documentation.
- Implement and operate ticketing and field service management tools.
- Drive hardware service workflows covering RMA, spares, dispatch, and coordination with global support partners.
Requirements
- 8+ years of experience in technical support, field service, or service delivery for hardware deployed at customer sites.
- 2+ years working with AI/ML infrastructure or accelerated compute environments and 2+ years leading support or field engineering teams.
- Ability to lead technical debugging calls, read logs, reason across the hardware/software boundary, and distinguish configuration, firmware, and silicon issues.
- Hands-on experience with Linux, Kubernetes, containerized inference serving, and Python automation.
- Experience supporting external customers with hardware faults, software issues, performance problems, and major enterprise or government incidents.
- Professional Korean and English fluency required; based in Korea.
Nice to have
- Background in AI/ML, HPC, data center infrastructure, accelerators, or GPU systems.
- Experience with distributed LLM inference orchestration, including vLLM disaggregated serving, NVIDIA Dynamo, llm-d, or equivalent systems.
- Experience with PyTorch, inference optimization, performance tuning, data center integration, and multi-layer performance diagnosis.
- Experience implementing field service systems, KCS knowledge management, AI-assisted support operations, or supporting air-gapped sites.
- Experience with controlled hardware movement, structurally different markets, or product supportability requirements before release.
Culture & Benefits
- Work across global delivery and support partners and technical and business teams.
- Support production systems deployed across multiple regions, enterprises, and governments.
- International travel of up to 25% is part of the role.
Hiring process
- Application screening followed by an online interview.
- On-site interview including an assignment, followed by management and culture-fit interviews.
- Compensation discussion and final offer.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
CrowdStrike
13 дней назад
Technical Support Engineer, Korean-speaking (Cybersecurity)
Replit
12 дней назад
Support Engineer
110 000 - 140 000$
14 дней назад
IT Support Engineer / Junior Technologist (AI)
13 дней назад
Technical Support Manager (AI)
Writer
13 дней назад
Senior Support Engineer (AI)
Replit
12 дней назад
Support Engineer I
110 000 - 140 000$