8 дней назад
Senior Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (AI): Building reliability standards, infrastructure automation, and observability tooling across on-premises, private cloud, and AWS environments with an accent on SLOs, incident response, and production readiness. Focus on designing scalable infrastructure, leading complex P1/P2 incident resolution, automating preventive fixes, and mentoring engineers in AI-assisted operations.
Location: Bielsko-Biała, Poland
Company
is a global data technology company focused on data quality, data enrichment, location intelligence, and AI-enabled products.
What you will do
- Define and maintain reliability standards, including SLOs, SLIs, error budgets, alerting, logging, and tracing requirements.
- Lead complex reliability engineering work across on-premises, private cloud, and AWS environments.
- Build infrastructure-as-code, deployment automation, monitoring, alerting, and observability tooling with Terraform, Ansible, Datadog, and scripting languages.
- Lead Operational Readiness Reviews, production and disaster recovery readiness validation, release triage, and deployment coordination.
- Lead complex P1/P2 incident response, incident command, root cause analysis, and preventive automation to reduce toil and improve MTTR.
- Mentor SREs and engineering teams, maintain operational documentation, coordinate on-call coverage, and participate in after-hours and holiday support rotations.
Requirements
- Bachelor’s degree in Computer Science, Information Systems, Engineering, or equivalent practical experience.
- At least 5 years of systems or infrastructure engineering experience in an enterprise production environment.
- Advanced Linux proficiency across on-premises and cloud environments, with strong experience in Terraform, Ansible, AWS, and at least one scripting language such as Python or Bash.
- Experience designing monitoring architecture, alerting strategies, CI/CD pipelines, deployment automation, and reliability standards.
- Strong knowledge of TCP/IP networking, DNS, load balancing, distributed systems, incident analysis, SLOs, and production readiness reviews.
- Proficient use of -provided AI tools such as GitHub Copilot or Claude for automation, incident analysis, runbook authoring, solution testing, and architecture documentation is required.
Nice to have
- Experience with Docker, ECS, Kubernetes, GitOps, GitLab, enterprise virtualization, ITIL, or security tools such as Qualys, CrowdStrike, and Rapid7.
- AWS Solutions Architect, SysOps Administrator, or DevOps Engineer certification.
- Wireshark, protocol analysis, mentoring, or on-call leadership experience.
Culture & Benefits
- Work across on-premises, private cloud, and AWS-managed service environments.
- Collaborate with engineering teams on reliability, security, compliance, vulnerability remediation, and data protection.
- Participate in a rotating on-call schedule covering after-hours, weekends, and holidays within defined SLA windows.
- A B2B contract is available for this position.
Hiring process
- The offer is dependent on positive candidate verification under ’s internal procedures.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →