3 дня назад
Software Golang Engineer (Slurm)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Golang Engineer (Slurm) (Go/Kubernetes/GPU Infrastructure): Design and build a managed Slurm service on Kubernetes with an accent on reliable Go development, workload scheduling, and GPU-intensive distributed systems. Focus on building observability and automated remediation for GPU, node, network, and control-plane failures across high-performance infrastructure.
Location: Poland, Serbia, Georgia, or Cyprus. Hybrid or remote options may be available depending on the role. Work from anywhere in the world is available for up to 45 days per year.
Company
provides infrastructure and software solutions for AI, cloud, networking, and security, operating edge and cloud infrastructure for digital products worldwide.
What you will do
- Design and build a managed Slurm service running on Kubernetes.
- Write clean, reliable, and maintainable Go code.
- Develop scheduling and orchestration capabilities for GPU-intensive and distributed workloads.
- Build observability and automated remediation for GPU, node, network, and control-plane failures using VictoriaMetrics, Grafana, and DCGM.
- Take end-to-end ownership of complex distributed-system challenges and collaborate with technology partners and customers.
Requirements
- Production experience using Slurm from a user perspective, including sbatch, srun, squeue, and sinfo.
- Strong Go proficiency and experience building production-grade Kubernetes operators, controllers, CRDs, and reconciliation loops.
- Experience preserving traditional Slurm cluster behavior while running infrastructure on Kubernetes.
- Experience diagnosing performance and reliability issues across GPUs, schedulers, hardware, high-performance networks, and distributed storage systems.
- Product mindset, customer empathy, excellent communication skills, and the ability to own complex technical challenges end to end.
Nice to have
- Experience operating large-scale HPC or GPU clusters for external customers.
- Experience with PyTorch distributed training and other large-scale AI/ML frameworks.
- Experience with InfiniBand, RoCE, RDMA, GPUDirect, Lustre, WEKA, Ceph, or similar high-performance infrastructure.
- Experience building unified job-submission workflows across Kubernetes and Slurm.
- Contributions to Slurm, Kubernetes, Soperator, or other cloud-native and HPC open-source projects.
Culture & Benefits
- Flexible working hours and hybrid or remote options depending on the role.
- Private medical insurance for employees and families, where applicable.
- Extra paid vacation and sick leave days, depending on location.
- Language courses, support for important life events, and team sports and social activities.
- Modern offices with snacks, drinks, and entertainment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →