4 дня назад
Forward Deployed Engineer (AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Forward Deployed Engineer (AI Infrastructure): Building and deploying high-performance distributed storage and AI infrastructure solutions for strategic enterprise customers with an accent on systems optimization, networking, and production reliability. Focus on debugging complex infrastructure issues, developing automation and APIs, and translating customer requirements into core product improvements.
Location: Singapore; ability to work on-site at customer locations and travel up to 25% as required
Company
Data platform company building enterprise infrastructure for capturing, managing, protecting, and analyzing massive datasets used in AI training and inference.
What you will do
- Work directly with customer architects, operators, and developers to design and operate strategic deployments.
- Optimize infrastructure throughput and debug complex distributed systems, storage, and networking issues.
- Develop custom API endpoints, automation tools, and scripts for critical customer workflows.
- Translate operational insights and customer requirements into precise technical requirements for core product teams.
- Advocate for customer needs while protecting product quality and roadmap priorities.
- Coordinate with engineering teams on builds, hotfixes, roadmap updates, and code reviews.
Requirements
- 5+ years of experience in a highly technical engineering role with customer-facing collaboration.
- Deep knowledge of distributed systems, high-performance file systems, storage protocols, and performance tuning.
- Strong understanding of InfiniBand, RoCE/RDMA, and 100GbE+ networking.
- Advanced knowledge of Linux architecture, memory management, and kernel mechanics.
- Proficiency in Python and C/C++, including debugging and contributing to complex enterprise codebases.
- Ability to manage high-stakes technical issues, communicate with both executives and developers, and work independently.
Nice to have
- Experience operating large-scale AI training or HPC clusters.
- Experience with Docker, Kubernetes, PyTorch, or TensorFlow.
Culture & Benefits
- Direct impact on infrastructure supporting AI training, inference, and real-time data analysis.
- Close collaboration with global engineering teams and strategic enterprise customers.
- Hands-on exposure to advanced distributed storage, networking, and AI infrastructure.
- Up to 25% travel to customer locations.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →