14 дней назад
Infrastructure Operations Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Infrastructure Operations Engineer (AI) (OpenStack/Kubernetes/GPU infrastructure): Designing, deploying, and operating scalable OpenStack and Kubernetes environments for business-critical GPU workloads with an accent on infrastructure automation, GPU scheduling, security, and platform resilience. Focus on building servers and racks, operating physical data-centre infrastructure, leading incident response, and improving reliability across large-scale compute environments.
Location: Portugal; the role requires hands-on work in data centres and travel to Quebec sites as required.
Company
develops Hyperstack, a full-stack AI cloud providing on-demand and private GPU infrastructure for researchers and enterprises.
What you will do
- Design, deploy, and operate OpenStack and Kubernetes environments for GPU workloads.
- Build infrastructure using infrastructure-as-code and GitOps, automating provisioning, deployment, and operational workflows.
- Optimise GPU workload scheduling with Kubernetes and NVIDIA tooling.
- Implement monitoring, logging, and alerting to maintain platform stability.
- Lead incident response and continuous reliability improvements.
- Maintain RBAC, network policies, tenant isolation, and other infrastructure and container security controls while collaborating with Platform, DevOps, AI, Product, and Support teams.
Requirements
- Extensive hands-on Linux systems administration experience.
- Proven experience physically building, cabling, commissioning, stacking, and racking servers and hardware.
- Direct hands-on data-centre experience and willingness to travel to Quebec sites as required.
- Strong understanding of networking and storage systems.
Nice to have
- Experience installing, racking, and configuring GPU hardware, especially NVIDIA platforms.
- Production experience running OpenStack or Kubernetes at scale.
- Experience with infrastructure automation, CI/CD, and Git-based workflows.
- Exposure to HPC or large-scale compute environments.
- Open-source contributions.
Culture & Benefits
- Flexible working arrangements, remote or hybrid depending on the role and location.
- Competitive salary and annual discretionary bonus scheme.
- Employee wellbeing benefits, 25 days of holiday, and public holidays.
- Significant ownership, autonomy, and opportunities to experiment.
- Career progression and growth opportunities in a collaborative, international environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →