Engineering 4 min read

The Engineering Blueprint for Multi-Region Database Replication

Building an application that survives a full region failure is one of the hardest problems in distributed systems. This is the blueprint we use for enterprises that need genuine multi-region resilience.

By Ananya Reddy · Published

The Engineering Blueprint for Multi-Region Database Replication

On April 20, 2021, a single misconfigured network device at an AWS US-EAST-1 data centre caused a cascade that took down a substantial portion of the internet, including services that had invested millions in cloud infrastructure. The organisations that maintained service continuity during that outage shared a common characteristic: they had genuinely multi-region architectures with battle-tested failover procedures. Building that level of resilience at the database layer , the hardest part , is what this post covers.

Understanding the CAP Theorem in Practice

Brewer's CAP theorem states that a distributed system can guarantee at most two of three properties: consistency (all nodes see the same data), availability (every request receives a response), and partition tolerance (the system operates despite network failures). In a multi-region deployment, network partitions are a near-certainty. You must choose between consistency and availability when a partition occurs. This choice , CP or AP , should drive your entire database architecture decision.

The Three Patterns for Multi-Region Databases

Three patterns dominate multi-region database architecture. Active-Passive (single-leader replication): one region handles all writes; the other regions replicate asynchronously and serve reads. Simplest to implement, easiest to reason about, but RPO (recovery point objective) is bounded by replication lag , potentially seconds to minutes of data loss on failover. Active-Active (multi-leader or leaderless): writes are accepted in multiple regions simultaneously. Eliminates write latency for geographically distributed users but introduces write conflicts that require careful resolution logic. Geo-Partitioned: different data partitions are primary in different regions. A European user's data lives and writes primarily in the EU region. Maximises performance and data sovereignty compliance but requires careful partition key design.

Implementing Active-Passive with PostgreSQL

For most enterprise applications, active-passive replication is the right starting point. In AWS, we implement this with PostgreSQL on RDS Multi-AZ for within-region redundancy, plus a cross-region read replica in the DR region. Replication is asynchronous , the primary doesn't wait for the replica to acknowledge writes. This means replication lag exists, typically 1-30 seconds. For failover, AWS RDS supports automated promotion of the cross-region replica to a new primary, but it requires DNS update propagation. We implement Route 53 health checks with a 60-second failover TTL. The realistic RTO (recovery time objective) is 3-8 minutes.

Write Conflict Resolution in Active-Active Architectures

Active-active databases are increasingly accessible via products like CockroachDB, YugabyteDB, and Amazon Aurora Global Database. Aurora Global Database provides sub-second replication lag with automated failover, making it the most operationally straightforward option for PostgreSQL-compatible workloads. The fundamental challenge in any active-active setup is write conflicts: two clients in different regions updating the same record simultaneously. Most systems implement last-write-wins (LWW) semantics using logical timestamps. For financial systems where last-write-wins is unacceptable, we implement application-level conflict resolution: optimistic locking with version vectors, or CRDT (conflict-free replicated data types) for specific data structures like counters.

Testing Failover: The Step Most Teams Skip

The most expensive database mistake we see in enterprise systems is an untested failover procedure. Teams invest in replication infrastructure but never validate that failover actually works , until a production incident forces the test. We implement quarterly chaos engineering exercises for every multi-region system. We trigger failover in a production-like staging environment, measure actual RTO and RPO against targets, and validate that application connection poolers handle the new primary endpoint correctly. Prisma, Sequelize, and other ORMs often cache database connections; validating they reconnect correctly after a primary switch is non-trivial.

Operational Considerations: The Day-Two Problem

Multi-region databases are significantly more expensive and operationally complex than single-region databases. Replication traffic crosses region boundaries , which has both cost and latency implications. Cross-region network egress is typically $0.08-0.09/GB. For write-heavy workloads with large row sizes, this cost can be substantial. Plan your replication traffic budget carefully. Additionally, schema migrations in multi-region active-active systems require extreme care: migrations must be backward-compatible with the previous version because you'll briefly have mixed-version nodes during deployment. We enforce expand-contract migration patterns and never drop columns or tables in the same migration that removes references to them.

DatabasesDistributed SystemsArchitectureHigh Availability