Назад
Company hidden
1 месяц назад

Senior Infrastructure Engineer (SRE)

55 000 - 68 000€
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Italy
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Infrastructure Engineer (SRE) (cloud infrastructure/Kubernetes): Operating and improving business-critical production services across cloud, identity, networking, databases, APIs, and distributed systems with an accent on incident response, observability, and resilient infrastructure. Focus on diagnosing complex failures, automating operational controls, leading recovery decisions, and engineering lasting reliability improvements.

Location: Milan or Turin, Italy

Salary: €55,000–€68,000 per year, plus a 10% bonus based on personal and company objectives.

Company

hirify.global develops service-driven supply chain planning software that helps companies improve inventory management, customer satisfaction, planner productivity, and financial performance.

What you will do

  • Lead major incidents from impact assessment and containment through recovery, stakeholder communication, root-cause analysis, and corrective actions.
  • Troubleshoot distributed production services across Windows and Linux, Kubernetes, containers, networking, DNS, cloud infrastructure, identity, databases, storage, APIs, and service dependencies.
  • Operate and improve Azure, OCI, or comparable cloud environments, including monitoring, access controls, backup and recovery, reliability, and cost-aware scaling.
  • Define and improve SLIs, SLOs, observability, alerting, capacity planning, resilience, dependency mapping, and recovery readiness.
  • Automate operational tasks and controls using PowerShell, Python, infrastructure as code, and CI/CD pipelines with testing, secure credential handling, validation, and rollback.
  • Lead and mentor the IT/Ops team while aligning Engineering, Product, Security, Support, and business stakeholders on service ownership, priorities, risks, and commitments.

Requirements

  • Typically 5+ years of SRE or production engineering experience operating business-critical, customer-facing, or high-availability services.
  • Recent hands-on ownership of high-severity incidents, including technical triage, recovery decisions, communication, and measurable follow-through.
  • Strong troubleshooting fundamentals across Windows, Linux, TCP/IP, DNS, routing, firewalls, proxies, load balancers, and multiple service layers.
  • Hands-on experience with Azure, OCI, or a comparable cloud platform, plus production Kubernetes and container operations.
  • Deep experience with observability, distributed tracing, actionable alerting, SLI/SLO design, capacity analysis, failure-mode thinking, and post-incident engineering.
  • Experience with Active Directory, Microsoft Entra ID or equivalent IAM, PowerShell or Python automation, Terraform or comparable infrastructure-as-code tooling, CI/CD, and cross-functional technical leadership.

Nice to have

  • Experience with Microsoft 365, endpoint management, EDR, device compliance, and hybrid workplace operations.
  • Formal ITIL, cloud, security, or infrastructure certifications.

Culture & Benefits

  • Work in a rapidly growing global company whose supply chain planning software is deployed in more than 44 countries.
  • Participate in a culture focused on problem solving, deep care, durable creativity, and finding the right answer rather than the first answer.
  • Receive a salary of €55,000–€68,000 per year plus a 10% performance-based bonus.
  • Equal employment opportunities are provided, and employment practices comply with applicable privacy and data-protection requirements.

Hiring process

  • Complete a scenario-based SRE technical discussion covering realistic production incidents involving cloud, Kubernetes, identity, networking, databases, storage, APIs, and service dependencies.
  • Explain the logs, metrics, traces, commands, tools, trade-offs, and recovery criteria used in incident diagnosis and resolution.
  • Discuss how to align technical and business stakeholders when priorities, risks, and customer commitments conflict.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →