8 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Senior Database Reliability Engineer (AI)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Π’Π΅ΠΊΡΡ:
TL;DR
Senior Database Reliability Engineer (AI): Owning and evolving the operational database platform for Firmus AI Cloud and internal services across bare-metal, self-hosted, and cloud infrastructure with an accent on PostgreSQL reliability, automation, recovery, and performance. Focus on designing HA and failover standards, building database self-service through infrastructure as code, and solving complex production issues involving replication, locking, capacity, and recovery.
Location: Singapore
Company
Technologies develops and operates sustainable AI infrastructure across Asia Pacific, including the AI Cloud GPU platform and AI Factory infrastructure.
What you will do
- Own and evolve database architecture standards for relational, document, NoSQL, and cache workloads across bare-metal, self-hosted, and cloud infrastructure.
- Automate database provisioning, configuration, upgrades, backups, restores, failover, and decommissioning with infrastructure as code.
- Define RPO, RTO, SLO, high availability, failover, replication, sharding, partitioning, tenant isolation, and multi-site standards.
- Resolve production issues involving query performance, lock contention, replication lag, memory and storage pressure, and network bottlenecks.
- Own database security, access control, encryption, audit logging, change control, and compliance evidence for SOC 2 Type 2 and ISO 27001.
- Partner with software, platform, data engineering, and observability teams on schema design, migrations, CDC, telemetry, and customer escalations.
Requirements
- 7+ years of database reliability engineering experience, including at least 3 years owning PostgreSQL in large-scale production.
- Deep PostgreSQL troubleshooting experience covering planner behavior, locking, bloat, connection storms, replication, failover, backup, restore, and query tuning.
- Hands-on experience with PostgreSQL HA and backup tools such as Patroni-style HA, pgBackRest or WAL-G, and pgBouncer or a similar pooler.
- Production ownership of Redis or Memcached and at least one document or NoSQL database such as MongoDB, Cassandra, or DynamoDB.
- Infrastructure-as-code automation with Python or Go; understanding of Linux storage and networking plus database monitoring with Prometheus, Grafana, OpenTelemetry, or similar tools.
- Clear written and verbal communication in English, willingness to join the database on-call rotation, and occasional overseas travel are required.
Nice to have
- Experience operating stateful databases on Kubernetes and using database operators.
Culture & Benefits
- Full-time role within the Engineering and Technology team.
- Work on sustainable, energy-efficient AI infrastructure and GPU cloud services.
- Participate in incident response, post-mortems, design reviews, and reliability improvement initiatives.
- Inclusive workplace that encourages applications from candidates of all backgrounds.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β
ΠΠΎΡ ΠΎΠΆΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
10 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Senior Site Reliability Engineer (AI)
11 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Senior/Lead Site Reliability Engineer (AI/LLM)
Airwallex
10 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Senior Site Reliability Engineer (Fintech)
13 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄
Site Reliability Engineer (Azure SaaS)
11 Π΄Π½Π΅ΠΉ Π½Π°Π·Π°Π΄