Monitoring and Tuning Replication Lag
Replication lag means your replica is behind the primary.
Too much lag can cause stale reads or failed failovers.
Check lag:
SELECT application_name, write_lag, flush_lag, replay_lag FROM pg_stat_replication;
Common causes:
- Slow disk on replica
- Network latency
- WAL retention too low (
wal_keep_sizetoo small) - High write load on primary
Tuning tips:
| Setting | Description |
|---|---|
wal_keep_size | Keep enough WAL files so replicas don’t fall behind |
max_wal_senders | Allow enough replicas to connect |
checkpoint_timeout | Adjust to control WAL file rotation |
hot_standby_feedback | Prevent conflicts with vacuum |
You can also monitor replication lag using:
pg_stat_replicationview- Prometheus + Grafana dashboards
pg_stat_activityfor long queries on replicas
Backup and PITR (Point-In-Time Recovery)
Replication isn’t a backup.
If you delete data by mistake, the replica will delete it too.
Always keep backups with pg_basebackup and WAL archiving for Point-In-Time Recovery (PITR).
Example:
restore_command = 'cp /var/lib/postgresql/wal_archive/%f %p'
Summary
| Goal | Technique | Tool |
|---|---|---|
| High Availability | Replication + Failover | Patroni, Stolon, pg_auto_failover |
| Read Scaling | Streaming replication | pgpool-II, HAProxy |
| Multi-Master | Logical replication (BDR) | BDR, pglogical |
| Monitoring | Replication lag & metrics | Prometheus, Grafana |
| Disaster Recovery | PITR | WAL archiving |
With the right setup, PostgreSQL can achieve near-zero downtime, reliable failover, and even global data consistency.
Category: PostgreSQL
