Назад
Company hidden
1 день назад

Site Reliability Engineer (AI SaaS)

Тип работы
fulltime
Грейд
junior
Английский
b2
Страна
Ireland
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (AI SaaS): Running, supporting, and scaling an AI-powered public SaaS platform for AI inference across distributed architectures with an accent on observability, automation, incident management, and operational security. Focus on improving SLOs, reducing operational toil with Terraform and Python, resolving complex customer-facing production issues, and strengthening reliability through cloud-native systems.

Location: Dublin, Ireland

Company

hirify.global develops cybersecurity and application delivery solutions that help organizations create, secure, and operate digital applications, including an AI-powered public SaaS platform.

What you will do

  • Monitor system behavior and ensure SLOs through metrics, logging, distributed tracing, and other observability tools.
  • Participate in a 24/7 support model, managing incidents, restoring services, analyzing root causes, and authoring postmortems.
  • Resolve high-priority customer escalations by analyzing logs, application traces, and system metrics.
  • Build automated workflows for infrastructure, monitoring, deployments, configuration management, and repetitive operational tasks using Terraform and scripting.
  • Collaborate with engineering teams to improve system architecture, reliability, logging, reporting, and alerting.
  • Apply SRE, high-availability, service mesh, container orchestration, security-as-code, and continuous improvement practices.

Requirements

  • Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent demonstrable experience.
  • 1–3+ years of experience in technical support, system administration, or cloud operations.
  • Foundational knowledge of public cloud environments such as AWS, Google Cloud, or OpenStack.
  • Proficiency in Python or Bash and familiarity with Infrastructure as Code tools such as Terraform.
  • Understanding of web technologies, protocols, and APIs, including HTTP, REST, and JSON.
  • Basic knowledge of networking, PostgreSQL or other databases, and Linux server administration, combined with strong troubleshooting and communication skills.

Nice to have

  • Experience with Kubernetes or other container orchestration systems.
  • Knowledge of Prometheus, Grafana, or equivalent observability tools.
  • Exposure to SLO management, automated delivery pipelines, and fault-tolerant architectures.
  • Familiarity with Ansible, Chef, or Puppet.

Culture & Benefits

  • Participation in an on-call rotation for out-of-hours incident response.
  • Collaborative work across engineering teams with opportunities to mentor peers and promote SRE practices.
  • Focus on automation, operational excellence, system reliability, and continuous service improvement.
  • Work in a cybersecurity environment focused on protecting applications and improving customer outcomes.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →