Skip to content

Widhian Bramantya

coding is an art form

Menu
  • About Me
Menu
postgresql

PostgreSQL Replication Deep Dive: From High Availability to Multi-Master Clusters

Posted on October 8, 2025October 8, 2025 by admin

Failover and High Availability Tools

Manual failover takes time, so automation is important.
Popular tools for PostgreSQL HA include:

Patroni

  • Uses Etcd or Consul for cluster coordination.
  • Detects failures and promotes a replica automatically.
flowchart TB
    subgraph D[etcd or Consul]
        E[Cluster State Store]
    end

    A[Clients] --> B[HAProxy - Routes to Current Primary]
    B --> C1[Primary Node - Patroni + PostgreSQL]
    B --> C2[Replica Node 1 - Patroni + PostgreSQL]
    B --> C3[Replica Node 2 - Patroni + PostgreSQL]

    C1 -. Heartbeat and Status .-> E
    C2 -. Heartbeat and Status .-> E
    C3 -. Heartbeat and Status .-> E

    E -. Leader Election .-> C1
    C1 -- WAL Stream --> C2
    C1 -- WAL Stream --> C3
    C2 -. Failover Promotion .-> C1

Here’s what each part of the diagram represents:

  • Clients → HAProxy:
    Clients connect through HAProxy (or Pgpool-II), which automatically routes traffic to whichever node is currently the primary.
  • Primary Node:
    This is the active PostgreSQL instance that handles all write operations.
    Patroni runs alongside PostgreSQL and communicates with the coordination service (etcd or Consul).
  • Replica Nodes:
    These nodes continuously stream WAL (Write-Ahead Log) data from the primary to stay synchronized.
    They are in standby mode but can be promoted if the primary fails.
  • etcd / Consul (Cluster Coordination):
    This key-value store keeps track of cluster state — which node is the leader, which replicas are healthy, etc.
    Patroni nodes use it for leader election and heartbeat monitoring.
  • Heartbeat and Leader Election:
    Every Patroni node sends heartbeat updates to etcd.
    If etcd detects that the current leader (primary) has stopped sending heartbeats, it triggers a leader election to promote a replica as the new primary.
  • Failover:
    Once the new leader is elected, Patroni automatically updates etcd, and HAProxy redirects client connections to the new primary.
    This entire process happens automatically — typically in a few seconds.
See also  Partitions, Replication, and Fault Tolerance in Kafka

Stolon

  • Works well with Kubernetes.
  • Separates cluster management and data state.
flowchart TB
    subgraph D[Stolon Keeper and Sentinel - Cluster Management]
        E[Cluster State Store - etcd or Consul]
    end

    A[Clients] --> B[Stolon Proxy - Routes to Current Primary]
    B --> C1[Keeper 1 - PostgreSQL Primary]
    B --> C2[Keeper 2 - PostgreSQL Replica]
    B --> C3[Keeper 3 - PostgreSQL Replica]

    C1 -. Heartbeat and Status .-> E
    C2 -. Heartbeat and Status .-> E
    C3 -. Heartbeat and Status .-> E

    E -. Leader Election .-> C1
    C1 -- WAL Stream --> C2
    C1 -- WAL Stream --> C3
    C2 -. Failover Promotion .-> C1

Stolon is another PostgreSQL high-availability (HA) solution, similar in goal to Patroni, but with a more container- and Kubernetes-friendly design.
It splits responsibilities into separate components: Keeper, Sentinel, Proxy, and a Cluster Store (etcd or Consul).

Here’s how it works:

  • Clients → Stolon Proxy:
    Clients connect through the Stolon Proxy.
    The proxy always knows which Keeper currently holds the primary PostgreSQL instance and routes all write traffic there automatically.
  • Keeper Nodes:
    Each Keeper manages one PostgreSQL instance.
    One Keeper runs as the primary, while others run as replicas (streaming WAL from the primary).
    If the primary Keeper fails, another Keeper can be promoted.
  • Sentinel Nodes:
    Sentinels are lightweight components that continuously monitor the health of Keepers.
    If a Sentinel detects a failure, it triggers a leader election and coordinates failover.
  • Cluster Store (etcd or Consul):
    This is the central state registry of the cluster.
    It records which Keeper is the current primary, which ones are replicas, and other cluster metadata.
    Both Keepers and Sentinels communicate with it.
  • Heartbeat and Failover:
    Keepers send heartbeat signals to the store.
    If the current primary stops sending heartbeats, Sentinels agree on a new leader and update the cluster state.
    The Proxy then redirects connections to the new primary automatically.
See also  What is Debezium? – An Introduction to Change Data Capture

pg_auto_failover

  • Built by Citus Data, easy to set up.
  • Uses a monitor node to manage primary/replica roles.

Simplified flow:

Primary  →  Replica  →  Failover →  Replica becomes new Primary

Use HAProxy or pgpool-II to automatically redirect connections to the new primary.

flowchart TB
    subgraph D[Monitor Node Cluster Management]
        E[Cluster State and Health Tracking]
    end

    A[Clients] --> B[Auto Failover Service - Virtual Endpoint]
    B --> C1[PostgreSQL Node 1 - Primary]
    B --> C2[PostgreSQL Node 2 - Secondary]

    C1 -. Heartbeat and Sync State .-> E
    C2 -. Heartbeat and Sync State .-> E

    E -. Promotion Decision .-> C1
    C1 -- WAL Stream --> C2
    C2 -. Automatic Failover .-> C1

pg_auto_failover is a simple and integrated PostgreSQL high-availability (HA) solution, developed by the Citus Data team (now part of Microsoft).
It’s designed to be easy to set up while still providing automatic monitoring and failover capabilities — ideal for small to medium clusters or single-region deployments.

Here’s how the components work:

  • Monitor Node:
    The Monitor is a lightweight control service that keeps track of the health and state of every data node. It decides when a failover should happen and which node should become the new primary.
    The monitor stores cluster metadata (for example, which node is primary, which is secondary, and their sync states).
  • Primary Node:
    The active PostgreSQL instance where write transactions occur.
    It continuously streams WAL (Write-Ahead Log) data to the secondary node.
    The monitor regularly checks its health via heartbeats.
  • Secondary Node:
    A standby PostgreSQL instance that receives WAL streams from the primary.
    If the primary fails, the monitor promotes this node automatically.
    After failover, the old primary can be reattached as a replica.
  • Auto Failover Service (Virtual Endpoint):
    Applications connect to the database through a load balancer or service endpoint.
    This connection is automatically redirected to the current primary after failover.
See also  Understanding PostgreSQL WAL, Slot, Publication, LSN, and Replication Lag

Related posts:

PostgreSQL Write-Ahead Log (WAL): Durability, Performance Tuning, and Recovery Explained

Understanding PostgreSQL WAL, Slot, Publication, LSN, and Replication Lag

Smart Automation in PostgreSQL: Managing Time-Based Data with pg_partman and pg_cron

Pages: 1 2 3 4
Category: PostgreSQL

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Linkedin

Widhian Bramantya

Recent Posts

  • Smart Automation in PostgreSQL: Managing Time-Based Data with pg_partman and pg_cron
  • Understanding PostgreSQL WAL, Slot, Publication, LSN, and Replication Lag
  • PostgreSQL Write-Ahead Log (WAL): Durability, Performance Tuning, and Recovery Explained
  • PostgreSQL Replication Deep Dive: From High Availability to Multi-Master Clusters
  • Finding Nearby Merchants in a Ride-Hailing App Using Elasticsearch Polygon Search
  • Advanced Text Search in Elasticsearch: N-Gram, Reverse, Fuzzy, and Search-as-you-type
  • Understanding and Customizing Analyzers in Elasticsearch
  • Log Management at Scale: Integrating Elasticsearch with Beats, Logstash, and Kibana
  • Index Lifecycle Management (ILM) in Elasticsearch: Automatic Data Control Made Simple
  • Blue-Green Deployment in Elasticsearch: Safe Reindexing and Zero-Downtime Upgrades
  • Maintaining Super Large Datasets in Elasticsearch
  • Elasticsearch Best Practices for Beginners
  • Implementing the Outbox Pattern with Debezium
  • Production-Grade Debezium Connector with Kafka (Postgres Outbox Example – E-Commerce Orders)
  • Connecting Debezium with Kafka for Real-Time Streaming
  • Debezium Architecture – How It Works and Core Components
  • What is Debezium? – An Introduction to Change Data Capture
  • Offset Management and Consumer Groups in Kafka
  • Partitions, Replication, and Fault Tolerance in Kafka
  • Delivery Semantics in Kafka: At Most Once, At Least Once, Exactly Once

Recent Comments

No comments to show.

Archives

  • October 2025
  • September 2025
  • August 2025
  • November 2021
  • October 2021
  • August 2021
  • July 2021
  • June 2021
  • March 2021
  • January 2021

Categories

  • Debezium
  • Devops
  • ElasticSearch
  • Golang
  • Kafka
  • Lua
  • NATS
  • PostgreSQL
  • Programming
  • RabbitMQ
  • Redis
  • VPC
© 2026 Widhian Bramantya | Powered by Minimalist Blog WordPress Theme