обновлено 16 часов назад
Network Reliability Engineer (AI)
200 - 250PLN
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Network Reliability Engineer (AI): Building and operating large-scale AI infrastructure with an accent on monitoring, incident diagnosis, observability, and service continuity. Focus on GPU and HPC environments, infrastructure lifecycle management, networking, automation, and improving stability, resiliency, scalability, and security.
Location: Remote, Warsaw, Poland
Salary: PLN 200–250 per hour
Company
's Consulting – Polska Team works on AI infrastructure and related engineering services.
What you will do
- Build and operate large-scale AI infrastructure across different environments and countries.
- Monitor infrastructure and application health, diagnose production incidents, and implement remediation.
- Troubleshoot high-impact production issues with engineering teams and participate in an on-call rotation.
- Implement and maintain observability solutions and contribute to system and tool evolution.
- Apply best practices for stability, resiliency, scalability, and security.
- Maintain technical documentation, collaborate with development teams, and participate in knowledge-sharing activities.
Requirements
- Experience with Go or Python and strong Bash/Python scripting skills.
- Hands-on experience with Linux systems, including Ubuntu or Debian.
- Knowledge of networking technologies such as VLAN/LAN, TCP/IP, DNS, BGP, load balancing, and IPv6.
- Familiarity with monitoring and logging tools such as Prometheus, Grafana, and Elastic.
- Experience with Infrastructure-as-Code tools such as Ansible, Salt, or AWX.
- Experience managing MariaDB and understanding of GitLab CI/CD pipelines; comfortable communicating in English.
Nice to have
- Hands-on experience with GPU and HPC infrastructure.
Culture & Benefits
- Remote work arrangement for the Warsaw-based Polska Team.
- Permanent employment contract or B2B engagement.
- Proactive, solution-oriented, collaborative, and knowledge-sharing work environment.
- Opportunities to contribute to automation and continuous improvement.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Site Reliability Engineer (AI Infrastructure)
175 000 - 265 000$
Инвольта
17 часов назад
DevOps-инженер (Highload RTB)
500 000₽
Valletta.Software | AI-Care
13 минут назад
Senior DevOps / SRE Support Engineer (LATAM)
5 000 - 5 500$
4 дня назад
Site Reliability Engineer Engineer
5 дней назад
Senior Site Reliability Engineer (AWS)
4 дня назад