обновлено 19 дней назад
Site Reliability Engineer (Python)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Python) (Infrastructure/Trading Platforms): Operating and improving distributed trading platforms across Windows and Linux environments, with an accent on automation, observability, incident response, and safe production delivery. Focus on building self-healing systems, designing CI/CD pipelines, tuning actionable monitoring, and solving complex reliability issues across regional and global infrastructure.
Location: Dublin, Ireland; rotating early/late weekday shifts, occasional weekend work, and participation in an on-call rota.
Company
operates technology platforms supporting trading activities, including distributed services, real-time news infrastructure, and production Python environments.
What you will do
- Operate production platforms across Windows and Linux environments, including deployments, capacity, certificates, health checks, and tenant operations.
- Triage incidents, perform root-cause analysis, and implement permanent fixes to prevent recurring failures.
- Build auto-remediation, replace manual runbooks with code, and eliminate recurring operational toil.
- Manage monitoring and alerting with Checkmk, Elasticsearch/Kibana, Grafana, and Prometheus, moving configuration into code and improving signal quality.
- Design and maintain GitLab CI/CD pipelines and Octopus Deploy orchestration for safe production releases.
- Build tested, documented production tooling and collaborate with development, infrastructure, and global platform-engineering teams.
Requirements
- Experience in site reliability, DevOps, platform, or production engineering.
- Strong scripting and automation skills with Python, PowerShell, or Bash.
- Practical experience working with both Windows and Linux.
- Hands-on experience with GitLab, Octopus Deploy, configuration management such as Ansible, and monitoring platforms such as Checkmk, Elasticsearch/Kibana, Grafana, or Prometheus.
- Understanding of infrastructure systems including compute, networking, storage, identity, RBAC, and certificate management.
- Strong troubleshooting, root-cause analysis, communication skills, and a technical degree or equivalent experience.
Culture & Benefits
- Hands-on engineering work at the boundary of development and infrastructure.
- Business-facing collaboration with developers, production engineers, infrastructure teams, and global platform-engineering groups.
- Automation-first practices focused on monitoring-as-code, self-healing systems, shared tooling, testing, and documentation.
- Rotating early/late weekday schedule with occasional weekend work and on-call participation.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Senior Site Reliability Engineer
6 дней назад
Senior Site Reliability Engineer (SRE) - Guidewire Cloud Platform (Application)
Twilio
6 дней назад
Staff Software Engineer (L4)
Anthropic
8 дней назад
Incident Response Manager (AI)
290 000 - 365 000$
11 дней назад
Senior Site Reliability Engineer - Platform Reliability (Resilience)
98 400 - 126 900€