Site Reliability Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Site Reliability Engineer (SRE): Ensuring reliability, resiliency, and innovation in mission-critical information systems with an accent on operational management, incident response, and enterprise-level infrastructure design. Focus on building robust systems, implementing monitoring strategies, and driving continuous improvement across hybrid cloud and on-premise environments.
Location: Must be based in Madrid, Spain, as the role requires regular office attendance.
Company
is a global technology services provider that designs, builds, manages, and modernizes the mission-critical technology systems that the world depends on.
What you will do
- Analyze business needs and provide strategic designs for robust, scalable information systems.
- Manage the full software lifecycle, from deployment and testing to maintenance and incident resolution.
- Define and implement service level indicators (SLIs) and objectives (SLOs) in collaboration with cross-functional teams.
- Design and implement application monitoring to ensure performance meets business goals.
- Develop strategies to manage operational load and handle overflow using metrics and automation.
- Collaborate with global teams to drive innovation and operational excellence.
Requirements
- Spanish: Native speaker required.
- Must be based in Madrid, Spain, for regular office attendance.
- 10+ years of experience in operational management, including incident management and escalations.
- Strong experience in enterprise environments with Windows, Linux (RHEL), and UNIX (AIX, Solaris).
- Proficiency in public cloud platforms (AWS, Azure, GCP) and container orchestration (OpenShift).
- Experience with scripting and data formats including JSON, YAML, Bash, and PowerShell.
Nice to have
- BS degree in Computer Science, Engineering, or a related technical discipline.
- Expertise in automation tools such as Ansible and Terraform.
- Proficiency in Python.
- Experience with distributed technologies and Kubernetes.
- Expertise in open-source monitoring tools like Prometheus, Grafana, or Loki.
Culture & Benefits
- Dynamic, hybrid-friendly work culture that prioritizes well-being and professional growth.
- Access to comprehensive Be Well programs supporting financial, mental, physical, and social health.
- Continuous learning opportunities, including certifications with Microsoft, Google, and Amazon.
- Inclusive environment focused on belonging, empathy, and shared success.
- Global collaboration opportunities within a large-scale, forward-thinking organization.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →