5 часов назад
Database Reliability Engineer (PostgreSQL)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Database Reliability Engineer (PostgreSQL): Keeping critical PostgreSQL, ClickHouse, MongoDB, and Redis services reliable with an accent on high availability, disaster recovery, automation, and observability. Focus on designing failover and restore workflows, building DBaaS-style self-service capabilities, and responding to complex production incidents.
Location: Fully remote; work from any location worldwide
Company
is hiring for an infrastructure-focused database reliability role supporting production database operations and engineering teams.
What you will do
- Own production PostgreSQL reliability, including HA design, Patroni, PgBouncer, replication, failover, upgrades, query tuning, backups, PITR, and restore validation.
- Improve disaster recovery through tested restores, documented recovery paths, measurable RTO/RPO targets, runbooks, and safe maintenance plans.
- Support ClickHouse, MongoDB, and Redis operations by troubleshooting incidents, reviewing access and data-safety changes, and improving monitoring.
- Automate provisioning, grants, backups, restores, health checks, and ownership metadata with Ansible, Terraform/OpenTofu, GitLab CI/CD, scripts, and reproducible runbooks.
- Build DBaaS-style self-service capabilities for database requests, access, credentials, and operational checks.
- Improve observability and incident response with Grafana, metrics, logs, SLOs, alert rules, Opsgenie routing, and clear production communications.
Requirements
- Senior-level, hands-on PostgreSQL experience in business-critical production environments, typically 5+ years or equivalent depth.
- Strong knowledge of PostgreSQL internals and operations, including MVCC, WAL, transactions, locks, indexes, query planning, replication, autovacuum, bloat, upgrades, backups, PITR, and restore testing.
- Experience with highly available databases and failover, quorum, split-brain risk, rollback, and recovery.
- Strong Linux and infrastructure fundamentals covering systemd, networking, storage, filesystems, resource bottlenecks, TLS, DNS, firewalls, and root-cause troubleshooting.
- Automation experience with Ansible and scripting; Terraform/OpenTofu, GitLab CI/CD, and merge-request-based delivery are advantageous.
- English at upper-intermediate level or higher is required for clear team communication. Ability to learn and operate multiple database engines, including ClickHouse, is also required.
Nice to have
- ClickHouse operations experience with replication, Keeper/ZooKeeper, MergeTree engines, distributed DDL, grants, row policies, backups, query troubleshooting, and cluster recovery.
- MongoDB replica sets and Percona Backup for MongoDB.
- Redis/Sentinel and broker or cache failure modes.
- Database observability, SLOs, golden signals, alert tuning, and executable incident runbooks.
- Experience building internal platforms, self-service portals, or DBaaS workflows.
Culture & Benefits
- Fully remote work with flexible working hours.
- Professional development focus and an education budget.
- 24 paid vacation days, 10 national holidays, and unlimited sick leave.
- Private medical insurance compensation.
- Co-working and gym or sports reimbursement.
- Reward opportunity for an innovative idea that the company can patent.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
MySQL Database Administrator
150 000 - 170 000$
Devhunt
7 часов назад
Database Administrator (Middle)
3 часа назад
Staff Platform Database Reliability Engineer (MySQL)
2 дня назад
Database Administrator (PostgreSQL/ClickHouse)
2 дня назад
Senior Storage Infrastructure Engineer (AI)
100 000 - 300 000$
Lambda
7 дней назад
Site Reliability Engineer (AI Infrastructure)
240 000 - 356 000$