4 дня назад
Industrial AI Cloud - Infrastructure Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Industrial AI Cloud - Infrastructure Engineer (AI) (NVIDIA/Linux/Ansible/Terraform): Building and operating compute, network, and storage infrastructure for a large-scale industrial AI cloud with an accent on GPU clusters, bare-metal systems, automation, and reliable operations. Focus on provisioning NVIDIA GPU nodes, automating infrastructure with Ansible and Terraform, operating AI and HPC workloads, and monitoring mission-critical services.
Location: Budapest, Debrecen, Pécs, or Szeged, Hungary. Hybrid working is available, with remote work restricted to Hungary due to European taxation regulations.
Company
is a Deutsche Telekom Group subsidiary providing IT and telecommunications services to large customers across Germany and Europe.
What you will do
- Build, automate, and operate compute, network, and storage infrastructure for an industrial AI cloud.
- Provision and maintain bare-metal servers and NVIDIA GPU nodes, including PXE boot, operating system installation, firmware updates, and hardware lifecycle activities.
- Develop and maintain Ansible and Terraform automation for infrastructure provisioning, configuration, and deployments.
- Operate Debian-based environments, Kubernetes and bare-metal AI/HPC workloads, identity and access management integrations, and high-performance WEKA storage.
- Manage Prometheus and Grafana monitoring, alerting, documentation, and ITIL-based incident, problem, and change processes.
- Coordinate with data center, NVIDIA, Deutsche Telekom, and other project teams to deliver reliable infrastructure services.
Requirements
- Experience with hardware installation, maintenance, and operations.
- Advanced production experience with Linux, preferably Debian.
- Hands-on Infrastructure-as-Code experience with Ansible and Terraform.
- Knowledge of NVIDIA GPU-accelerated servers, GPU orchestration, and AI cloud platform dependencies.
- Solid networking knowledge covering IP, routing, VLANs, DNS, firewalls, and L1/L2 operations.
- Experience with Keycloak, Entra ID, LDAP, Prometheus, Grafana, ITIL processes, and operational troubleshooting in 24/7 mission-critical environments.
Nice to have
- Experience with large GPU clusters, HPC, or data center environments.
- Knowledge of sovereign cloud, data security, and compliance requirements.
- Experience with Redfish, WEKA by Hitachi, and GitOps for infrastructure changes.
Culture & Benefits
- Work on an industrial AI cloud developed jointly with NVIDIA.
- Direct collaboration with NVIDIA and Deutsche Telekom experts.
- Hybrid working model within Hungary.
- Training opportunities and career progression.
- Work in accordance with data security and privacy requirements for European Union customers.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior DevOps / Platform Engineer (Kubernetes)
15 - 20$
6 дней назад
Senior DevSecOps Engineer (AWS/Kubernetes)
70 000 - 85 000€
6 дней назад
Staff/Senior DevOps Engineer - SCAYLE Platform (all genders) (AWS)
75 000 - 100 000€
7 дней назад
Senior Golang Engineer - DevOps Enablement (all genders)
70 000 - 85 000€
Cohere
5 дней назад
Deployment Engineer (AI)
4 дня назад