7 дней назад
Designated Service Engineer (AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Designated Service Engineer (AI Infrastructure): Troubleshooting and supporting enterprise AI data infrastructure for customers in South Korea with an accent on distributed storage, Linux administration, networking, and system monitoring. Focus on diagnosing hardware failures and performance bottlenecks, resolving complex issues across distributed environments, and bridging customer needs with Engineering.
Location: South Korea
Company
develops NeuralMesh, a cloud- and AI-native data infrastructure platform that accelerates AI model training, inference, and other compute-intensive workloads.
What you will do
- Act as the primary technical liaison between customers and Engineering on feature gaps, reliability issues, and documentation improvements.
- Troubleshoot hardware failures, network congestion, performance bottlenecks, and other enterprise infrastructure issues.
- Monitor systems remotely, identify potential issues, and manage customer cases through the ticketing system.
- Provide technical expertise to customers, account teams, pre-sales engineers, partners, and resellers.
- Create internal and customer-facing documentation, including FAQs and knowledge-base articles.
- Participate in follow-the-sun on-call rotations, alternative work hours, and potential regional or international travel.
Requirements
- 10+ years of experience in customer-facing technical roles supporting complex enterprise infrastructure.
- Experience with L3 or higher support for Linux-based storage, networking, virtualization, cloud, or related infrastructure.
- Strong troubleshooting skills in multi-platform, distributed environments and a deep understanding of distributed storage systems.
- Expertise in Linux/Unix administration and networking technologies including InfiniBand, Ethernet, DPDK, and UCX.
- Proficiency in Python and Bash, including automation scripting for monitoring and troubleshooting.
- Knowledge of POSIX, NFS, S3, log management, Prometheus, and Grafana.
Nice to have
- Experience with JIRA, Confluence, Slack, Kubernetes, containers, LXC, AWS, Azure, OCI, or GCP.
- Experience collaborating between customer support and product development teams.
- Experience managing large-scale HPC clusters.
- Strong technical writing and creative problem-solving skills.
Culture & Benefits
- Accountability, ownership, integrity, and high standards are emphasized.
- Collaboration, empathy, transparency, and constructive conflict resolution are core values.
- Customer success is a central priority in decision-making and service delivery.
- The role includes follow-the-sun support rotations and potential regional or international travel.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →