Site Reliability Engineer II (Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Site Reliability Engineer II (Infrastructure): Building and maintaining automation tools for physical security infrastructure and cloud environments with an accent on reliability, scalability, and reducing operational toil. Focus on developing IaC configurations, API-driven network management, and advanced monitoring systems for large-scale device fleets.
Location: Remote, Ireland
Salary: €50,640 – €63,300 (Annual)
Company
provides specialized IT infrastructure support and managed services for large-scale enterprise environments.
What you will do
- Establish and enforce SLIs/SLOs for infrastructure tooling, configuration compliance, and deployment metrics.
- Automate repetitive manual tasks across managed infrastructure to measurably reduce operational overhead and MTTR.
- Develop and maintain Infrastructure-as-Code (IaC) for Windows and Linux roles using Ansible, Terraform, or Puppet.
- Build API-driven tools for network configuration management, zero-touch provisioning, and real-time health monitoring.
- Deploy standardized monitoring agents, centralized log collection (ELK), and custom dashboards with Prometheus and Grafana.
- Participate in a 24x5 on-call rotation to ensure continuity for mission-critical physical security infrastructure.
Requirements
- 6+ years of experience in Infrastructure or Automation Engineering.
- Strong proficiency in Python, Bash, and PowerShell, with experience in Go for backend services.
- Hands-on experience with Terraform, Ansible, Chef, or Puppet, including drift detection and version control.
- Advanced knowledge of Linux/Windows server environments and Cisco device administration.
- Experience with monitoring solutions (Prometheus, Grafana, Datadog) and CMDB platforms like NetBox.
- Must be based in Ireland.
Culture & Benefits
- Full-time employment with a focus on enterprise-grade server and network management.
- Opportunity to work within a large-scale, cloud-hosted environment using Kubernetes and Helm.
- Culture of blameless postmortems and continuous systemic improvement for infrastructure resilience.
- Engagement with cutting-edge internal toolchains for development and code review.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →