Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Software Engineer, Infrastructure (Data & AI) (Kubernetes, Kafka, Spark, Flink): Building and operating highly available infrastructure for high-throughput data processing, real-time workloads, and production AI applications with an accent on distributed systems, cloud platforms, and platform reliability. Focus on designing AI gateways and model-serving infrastructure, improving scalability and observability, and creating self-service capabilities for engineering teams.
Location: Seoul, South Korea
Company
Airwallex is an AI-native financial operating system providing infrastructure for global payments and financial products.
What you will do
- Design, build, and operate highly available data and AI infrastructure on Kubernetes and public cloud platforms.
- Develop scalable streaming, batch-processing, and real-time data platforms using Kafka, Spark, and Flink.
- Build AI gateways, model-routing layers, traffic management, rate limiting, authentication, observability, and usage controls.
- Create self-service platform capabilities that allow data, AI, and application teams to deploy and operate workloads independently.
- Collaborate across application, data, machine learning, security, and infrastructure teams to define durable platform capabilities.
- 5+ years of experience in DevOps, SRE, or platform engineering with end-to-end ownership of production systems.
- Strong experience designing, operating, and troubleshooting production Kubernetes environments.
- Experience with distributed data infrastructure such as Kafka, Spark, or Flink, or with AI gateways and model-routing platforms.
- Hands-on experience with AWS, Google Cloud, or Microsoft Azure.
- Strong knowledge of cloud and container networking, distributed systems, infrastructure-as-code, automated delivery, and observability.
- Production programming experience in Go, Python, or Java, along with strong debugging, communication, and ownership skills.
- Experience building internal developer platforms and paved-road workflows.
- Experience with SGLang, vLLM, or NVIDIA Triton Inference Server.
- Knowledge of GPU scheduling, batching, model parallelism, memory management, autoscaling, and inference optimization.
- Experience improving the cost efficiency of large-scale data processing or AI inference workloads.
- Contributions to infrastructure, data-platform, Kubernetes, or AI-serving open-source projects.
- Work on infrastructure supporting global payments, data platforms, and production AI applications.
- Collaborate across multiple engineering disciplines and international offices.
- Build reusable platform capabilities adopted by engineering teams.
- Focus on improving scalability, reliability, operational efficiency, observability, and cost visibility.
Requirements
Nice to have
Culture & Benefits
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →