agora inbox for pgsql-hackers@postgresql.org
help / color / mirror / Atom feedOffline data checksum changes can cause incorrect checksum state on standbys
43+ messages / 6 participants
[nested] [flat]
* Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-12 07:55 Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 1 reply; 43+ messages in thread
From: Bertrand Drouvot @ 2026-08-12 07:55 UTC (permalink / raw)
To: pgsql-hackers@lists.postgresql.org; +Cc: Daniel Gustafsson <daniel@yesql.se>
Hi hackers,
while working on [1], I hit 2 issues involving offline checksum changes.
First one, is a case where a standby can enable checksum verification without its
own pages having been checksummed, making the standby unreadable.
The issue is due to f19c0eccae96 as the standby can enable checksum verification
from primary WAL.
Repro 1:
1/ create a primary and a standby both with checksum set to false
2/ stop the standby and the primary
3/ enable checksums only on the primary
4/ restart the primary and run checkpoint: this checkpoint XLOG_CHECKPOINT_REDO
record carries data_checksum_version=on. It gives the standby a WAL record that
makes it enable checksum verification.
5/ restart the standby: that will replay the primary’s new checksum state, despite
never having its own pages checksummed.
6/ try to connect to the standby: FATAL: invalid page in block 0 of relation "global/1260"
The second issue occurs when combining online and offline checksum transitions.
Repro 2:
1/ create a primary and a standby both with checksum set to false
2/ enable checksums online on the primary and wait that data_checksums is on
on the primary and on the standby
3/ stop only the standby
4/ disable checksums offline only on the standby
5/ restart the standby and check data_checksums. You'll see that it's still on
despites that we disabled it in step 4/
I initially considered detecting the mismatch during WAL replay and reporting an
error. Although this produces a clear error message, it does not help much in
practice because the standby cannot recover and has to be recreated.
Therefore, I think a simpler fix is to preserve the pre-f19c0eccae96 behavior
for offline checksum changes: they are not propagated through WAL. In the second
repro, the standby therefore remains off, honoring its local offline change.
This is what the attached patch proposes: it marks offline changes as local and
tracks the latest WAL-logged transition, so recovery ignores remote offline
states while still applying newer online transitions.
If this looks like too much code changes so close to the v19 release, another
option could be to remove pg_checksums --enable and --disable while keeping --check
and require checksum state changes to be done online.
As this is a 19 regression, I think it should be added as an open item.
Thoughts?
[1]: https://postgr.es/m/ajAAwSFy0WVMroyk%40bdtpg
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
Attachments:
[text/x-diff] v1-0001-Keep-offline-data-checksum-changes-local.patch (29.5K, ../../anwm6UPxoVS41QA2@bdtpg/2-v1-0001-Keep-offline-data-checksum-changes-local.patch)
download | inline diff:
From 6d3424c2fd0f20ddc73075bea22f9ad82d2b96d0 Mon Sep 17 00:00:00 2001
From: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Date: Tue, 11 Aug 2026 11:41:28 +0000
Subject: [PATCH v1] Keep offline data checksum changes local
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Commit f19c0eccae96 made checkpoint records propagate the data checksum state.
This lets recovery restore online checksum transitions. However, pg_checksums
changes that state while the server is stopped and rewrites only the selected
data directory. Replaying a checkpoint from that node could enable verification
on a standby whose pages were never checksummed.
Mark checksum states written by pg_checksums as node-local. Also retain the end
LSN of the newest applied WAL-logged checksum transition. Before saving a primary’s
replayed checkpoint as a possible standby restartpoint, recovery replaces the
checkpoint’s checksum fields with the standby’s current local checksum state.
It ignores offline state from another node and transitions no newer than the
local watermark, while newer online transitions still take effect.
Preserve localized checkpoint fields across startup and backups taken from a
standby. Keep the current control-file state separate from the historical state
needed to restart replay.
Serialize checksum transition WAL insertion and publication with redo records.
Thus a transition cannot fall before the checkpoint redo point while leaving stale
state in that record.
Add TAP coverage for both cross node offline cases. Verify that an unchecked
standby neither adopts the primary offline transition nor becomes unreadable,
and that a standby local offline disable survives restart after an online
transition. Also exercise primary crash recovery plus standby clean and crash
restarts.
XXX: Bump control file format
XXX: Bump WAL format
Author: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by:
Discussion:
Backpatch-through: 19
---
src/backend/access/transam/xlog.c | 211 +++++++++++++++---
src/backend/access/transam/xlogrecovery.c | 7 +-
.../utils/activity/wait_event_names.txt | 1 +
src/bin/pg_checksums/pg_checksums.c | 1 +
src/bin/pg_resetwal/pg_resetwal.c | 3 +
src/include/access/xlog_internal.h | 2 +
src/include/access/xlogrecovery.h | 3 +-
src/include/catalog/pg_control.h | 8 +-
src/include/storage/lwlocklist.h | 1 +
src/test/modules/test_checksums/meson.build | 1 +
.../test_checksums/t/010_offline_standby.pl | 87 ++++++++
11 files changed, 285 insertions(+), 40 deletions(-)
72.6% src/backend/access/transam/
3.4% src/include/catalog/
19.1% src/test/modules/test_checksums/t/
4.7% src/
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index b23d8bbbdad..b8414a176a9 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -559,6 +559,10 @@ typedef struct XLogCtlData
/* last data_checksum_version we've seen */
uint32 data_checksum_version;
+ /* current checksum state and latest online transition */
+ XLogRecPtr data_checksum_transition_lsn;
+ bool data_checksum_state_is_local;
+
slock_t info_lck; /* locks shared variables shown above */
} XLogCtlData;
@@ -676,6 +680,17 @@ static bool updateMinRecoveryPoint = true;
*/
static ChecksumStateType LocalDataChecksumState = 0;
+/*
+ * This node's checksum state after replaying the latest XLOG_CHECKPOINT_REDO
+ * record. Preserve it until the matching XLOG_CHECKPOINT_ONLINE record is
+ * replayed, then store it in the restartpoint instead of the checksum state
+ * received from the primary.
+ */
+static ChecksumStateType replayedCheckpointDataChecksumState = 0;
+static XLogRecPtr replayedCheckpointDataChecksumTransitionLSN = InvalidXLogRecPtr;
+static bool replayedCheckpointDataChecksumStateIsLocal = false;
+static bool replayedCheckpointDataChecksumStateValid = false;
+
/*
* Variable backing the GUC, keep it in sync with LocalDataChecksumState.
* See SetLocalDataChecksumState().
@@ -749,7 +764,7 @@ static void WALInsertLockAcquireExclusive(void);
static void WALInsertLockRelease(void);
static void WALInsertLockUpdateInsertingAt(XLogRecPtr insertingAt);
-static void XLogChecksums(uint32 new_type);
+static XLogRecPtr XLogChecksums(uint32 new_type);
/*
* Insert an XLOG record represented by an already-constructed chain of data
@@ -4283,12 +4298,16 @@ InitControlFile(uint64 sysidentifier, uint32 data_checksum_version)
ControlFile->wal_log_hints = wal_log_hints;
ControlFile->track_commit_timestamp = track_commit_timestamp;
ControlFile->data_checksum_version = data_checksum_version;
+ ControlFile->data_checksum_transition_lsn = InvalidXLogRecPtr;
+ ControlFile->data_checksum_state_is_local = false;
/*
* Set the data_checksum_version value into XLogCtl, which is where all
* processes get the current value from.
*/
XLogCtl->data_checksum_version = data_checksum_version;
+ XLogCtl->data_checksum_transition_lsn = InvalidXLogRecPtr;
+ XLogCtl->data_checksum_state_is_local = false;
}
static void
@@ -4747,6 +4766,7 @@ DataChecksumsNeedVerify(void)
void
SetDataChecksumsOnInProgress(void)
{
+ XLogRecPtr transition_lsn;
uint64 barrier;
/*
@@ -4756,14 +4776,12 @@ SetDataChecksumsOnInProgress(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
- SpinLockRelease(&XLogCtl->info_lck);
+ transition_lsn = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
+ ControlFile->data_checksum_transition_lsn = transition_lsn;
+ ControlFile->data_checksum_state_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4800,6 +4818,7 @@ SetDataChecksumsOnInProgress(void)
void
SetDataChecksumsOn(void)
{
+ XLogRecPtr transition_lsn;
uint64 barrier;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4824,11 +4843,7 @@ SetDataChecksumsOn(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_VERSION);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- SpinLockRelease(&XLogCtl->info_lck);
+ transition_lsn = XLogChecksums(PG_DATA_CHECKSUM_VERSION);
/*
* Update the controlfile before waiting since if we have an immediate
@@ -4836,6 +4851,8 @@ SetDataChecksumsOn(void)
*/
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
+ ControlFile->data_checksum_transition_lsn = transition_lsn;
+ ControlFile->data_checksum_state_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4864,6 +4881,7 @@ SetDataChecksumsOn(void)
void
SetDataChecksumsOff(void)
{
+ XLogRecPtr transition_lsn;
uint64 barrier;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4890,14 +4908,12 @@ SetDataChecksumsOff(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ transition_lsn = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
+ ControlFile->data_checksum_transition_lsn = transition_lsn;
+ ControlFile->data_checksum_state_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4928,14 +4944,12 @@ SetDataChecksumsOff(void)
/* Ensure that we don't incur a checkpoint during disabling checksums */
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ transition_lsn = XLogChecksums(PG_DATA_CHECKSUM_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_OFF;
+ ControlFile->data_checksum_transition_lsn = transition_lsn;
+ ControlFile->data_checksum_state_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -5426,6 +5440,8 @@ XLOGShmemInit(void *arg)
/* Use the checksum info from control file */
XLogCtl->data_checksum_version = ControlFile->data_checksum_version;
+ XLogCtl->data_checksum_transition_lsn = ControlFile->data_checksum_transition_lsn;
+ XLogCtl->data_checksum_state_is_local = ControlFile->data_checksum_state_is_local;
SetLocalDataChecksumState(XLogCtl->data_checksum_version);
SpinLockInit(&XLogCtl->Insert.insertpos_lck);
@@ -5512,6 +5528,8 @@ BootStrapXLOG(uint32 data_checksum_version)
checkPoint.time = (pg_time_t) time(NULL);
checkPoint.oldestActiveXid = InvalidTransactionId;
checkPoint.dataChecksumState = data_checksum_version;
+ checkPoint.dataChecksumTransitionLSN = InvalidXLogRecPtr;
+ checkPoint.dataChecksumStateIsLocal = false;
TransamVariables->nextXid = checkPoint.nextXid;
TransamVariables->nextOid = checkPoint.nextOid;
@@ -5850,6 +5868,10 @@ StartupXLOG(void)
bool didCrash;
bool haveTblspcMap;
bool haveBackupLabel;
+ bool backupFromStandby;
+ uint32 localDataChecksumState;
+ XLogRecPtr localDataChecksumTransitionLSN;
+ bool localDataChecksumStateIsLocal;
XLogRecPtr EndOfLog;
TimeLineID EndOfLogTLI;
TimeLineID newTLI;
@@ -5987,10 +6009,46 @@ StartupXLOG(void)
* starting checkpoint, and sets InRecovery and ArchiveRecoveryRequested.
* It also applies the tablespace map file, if any.
*/
+ localDataChecksumState = ControlFile->checkPointCopy.dataChecksumState;
+ localDataChecksumTransitionLSN = ControlFile->checkPointCopy.dataChecksumTransitionLSN;
+ localDataChecksumStateIsLocal = ControlFile->checkPointCopy.dataChecksumStateIsLocal;
InitWalRecovery(ControlFile, &wasShutdown,
- &haveBackupLabel, &haveTblspcMap);
+ &haveBackupLabel, &haveTblspcMap,
+ &backupFromStandby);
checkPoint = ControlFile->checkPointCopy;
+ /*
+ * The checkpoint copy in pg_control has been localized by this node.
+ * InitWalRecovery rereads the original WAL record, so restore the local
+ * fields unless a primary backup selected a different checkpoint. A
+ * standby backup contains pages matching its localized state.
+ */
+ if (!haveBackupLabel || backupFromStandby)
+ {
+ checkPoint.dataChecksumState = localDataChecksumState;
+ checkPoint.dataChecksumTransitionLSN = localDataChecksumTransitionLSN;
+ checkPoint.dataChecksumStateIsLocal = localDataChecksumStateIsLocal;
+ ControlFile->checkPointCopy.dataChecksumState = localDataChecksumState;
+ ControlFile->checkPointCopy.dataChecksumTransitionLSN = localDataChecksumTransitionLSN;
+ ControlFile->checkPointCopy.dataChecksumStateIsLocal = localDataChecksumStateIsLocal;
+ }
+
+ /*
+ * A local offline operation override the state in the selected
+ * checkpoint. Otherwise, restore that checkpoint's state before replay so
+ * later online checksum records are applied in the right order.
+ */
+ if (!ControlFile->data_checksum_state_is_local)
+ {
+ XLogCtl->data_checksum_version = checkPoint.dataChecksumState;
+ XLogCtl->data_checksum_transition_lsn = checkPoint.dataChecksumTransitionLSN;
+ XLogCtl->data_checksum_state_is_local = checkPoint.dataChecksumStateIsLocal;
+ SetLocalDataChecksumState(XLogCtl->data_checksum_version);
+ ControlFile->data_checksum_version = checkPoint.dataChecksumState;
+ ControlFile->data_checksum_transition_lsn = checkPoint.dataChecksumTransitionLSN;
+ ControlFile->data_checksum_state_is_local = checkPoint.dataChecksumStateIsLocal;
+ }
+
/* initialize shared memory variables from the checkpoint record */
TransamVariables->nextXid = checkPoint.nextXid;
TransamVariables->nextOid = checkPoint.nextOid;
@@ -6603,10 +6661,9 @@ StartupXLOG(void)
*/
if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON)
{
- XLogChecksums(PG_DATA_CHECKSUM_OFF);
+ (void) XLogChecksums(PG_DATA_CHECKSUM_OFF);
SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
SetLocalDataChecksumState(XLogCtl->data_checksum_version);
SpinLockRelease(&XLogCtl->info_lck);
@@ -6624,10 +6681,9 @@ StartupXLOG(void)
*/
else if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
{
- XLogChecksums(PG_DATA_CHECKSUM_OFF);
+ (void) XLogChecksums(PG_DATA_CHECKSUM_OFF);
SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
SetLocalDataChecksumState(XLogCtl->data_checksum_version);
SpinLockRelease(&XLogCtl->info_lck);
@@ -6654,6 +6710,8 @@ StartupXLOG(void)
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_transition_lsn = XLogCtl->data_checksum_transition_lsn;
+ ControlFile->data_checksum_state_is_local = XLogCtl->data_checksum_state_is_local;
XLogCtl->SharedRecoveryState = RECOVERY_STATE_DONE;
SpinLockRelease(&XLogCtl->info_lck);
@@ -7524,6 +7582,8 @@ CreateCheckPoint(int flags)
*/
SpinLockAcquire(&XLogCtl->info_lck);
checkPoint.dataChecksumState = XLogCtl->data_checksum_version;
+ checkPoint.dataChecksumTransitionLSN = XLogCtl->data_checksum_transition_lsn;
+ checkPoint.dataChecksumStateIsLocal = XLogCtl->data_checksum_state_is_local;
SpinLockRelease(&XLogCtl->info_lck);
if (shutdown)
@@ -7580,10 +7640,13 @@ CreateCheckPoint(int flags)
{
xl_checkpoint_redo redo_rec;
+ LWLockAcquire(DataChecksumsStateLock, LW_EXCLUSIVE);
WALInsertLockAcquire();
redo_rec.wal_level = wal_level;
SpinLockAcquire(&XLogCtl->info_lck);
redo_rec.data_checksum_version = XLogCtl->data_checksum_version;
+ redo_rec.data_checksum_transition_lsn = XLogCtl->data_checksum_transition_lsn;
+ redo_rec.data_checksum_state_is_local = XLogCtl->data_checksum_state_is_local;
SpinLockRelease(&XLogCtl->info_lck);
WALInsertLockRelease();
@@ -7591,6 +7654,7 @@ CreateCheckPoint(int flags)
XLogBeginInsert();
XLogRegisterData(&redo_rec, sizeof(xl_checkpoint_redo));
(void) XLogInsert(RM_XLOG_ID, XLOG_CHECKPOINT_REDO);
+ LWLockRelease(DataChecksumsStateLock);
/*
* XLogInsertRecord will have updated XLogCtl->Insert.RedoRecPtr in
@@ -7941,6 +8005,8 @@ CreateEndOfRecoveryRecord(void)
/* start with the latest checksum version (as of the end of recovery) */
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_transition_lsn = XLogCtl->data_checksum_transition_lsn;
+ ControlFile->data_checksum_state_is_local = XLogCtl->data_checksum_state_is_local;
SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
@@ -8292,8 +8358,15 @@ CreateRestartPoint(int flags)
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
}
- /* we shall start with the latest checksum version */
- ControlFile->data_checksum_version = lastCheckPoint.dataChecksumState;
+ /*
+ * Keep the current checksum state separate from the historical state
+ * in checkPointCopy, which is used when replay starts again.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_transition_lsn = XLogCtl->data_checksum_transition_lsn;
+ ControlFile->data_checksum_state_is_local = XLogCtl->data_checksum_state_is_local;
+ SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
}
@@ -8734,9 +8807,9 @@ XLogReportParameters(void)
}
/*
- * Log the new state of checksums
+ * Log and publish the new state of checksums
*/
-static void
+static XLogRecPtr
XLogChecksums(uint32 new_type)
{
xl_checksum_state xlrec;
@@ -8744,11 +8817,23 @@ XLogChecksums(uint32 new_type)
xlrec.new_checksum_state = new_type;
+ LWLockAcquire(DataChecksumsStateLock, LW_EXCLUSIVE);
+
XLogBeginInsert();
XLogRegisterData((char *) &xlrec, sizeof(xl_checksum_state));
recptr = XLogInsert(RM_XLOG2_ID, XLOG2_CHECKSUMS);
XLogFlush(recptr);
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ XLogCtl->data_checksum_version = new_type;
+ XLogCtl->data_checksum_transition_lsn = recptr;
+ XLogCtl->data_checksum_state_is_local = false;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockRelease(DataChecksumsStateLock);
+
+ return recptr;
}
/*
@@ -8933,10 +9018,24 @@ xlog_redo(XLogReaderState *record)
ProcArrayApplyRecoveryInfo(&running);
}
+ /*
+ * Checkpoint checksum state is local to the node that wrote it.
+ * Preserve this node's state before retaining the checkpoint as a
+ * possible restartpoint.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checkPoint.dataChecksumState = XLogCtl->data_checksum_version;
+ checkPoint.dataChecksumTransitionLSN = XLogCtl->data_checksum_transition_lsn;
+ checkPoint.dataChecksumStateIsLocal = XLogCtl->data_checksum_state_is_local;
+ SpinLockRelease(&XLogCtl->info_lck);
+ replayedCheckpointDataChecksumStateValid = false;
+
/* ControlFile->checkPointCopy always tracks the latest ckpt XID */
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->checkPointCopy.nextXid = checkPoint.nextXid;
ControlFile->data_checksum_version = checkPoint.dataChecksumState;
+ ControlFile->data_checksum_transition_lsn = checkPoint.dataChecksumTransitionLSN;
+ ControlFile->data_checksum_state_is_local = checkPoint.dataChecksumStateIsLocal;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -9000,6 +9099,18 @@ xlog_redo(XLogReaderState *record)
checkPoint.oldestXid))
SetTransactionIdLimit(checkPoint.oldestXid,
checkPoint.oldestXidDB);
+
+ /*
+ * Use this node's state from the matching redo record. Using the
+ * state at checkpoint completion would be wrong if an online
+ * transition occurred while the checkpoint was running.
+ */
+ Assert(replayedCheckpointDataChecksumStateValid);
+ checkPoint.dataChecksumState = replayedCheckpointDataChecksumState;
+ checkPoint.dataChecksumTransitionLSN = replayedCheckpointDataChecksumTransitionLSN;
+ checkPoint.dataChecksumStateIsLocal = replayedCheckpointDataChecksumStateIsLocal;
+ replayedCheckpointDataChecksumStateValid = false;
+
/* ControlFile->checkPointCopy always tracks the latest ckpt XID */
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->checkPointCopy.nextXid = checkPoint.nextXid;
@@ -9180,12 +9291,29 @@ xlog_redo(XLogReaderState *record)
memcpy(&redo_rec, XLogRecGetData(record), sizeof(xl_checkpoint_redo));
+ /*
+ * An offline state overrites transitions already replayed on this
+ * node. Apply a newer online transition from the redo record, but
+ * never propagate an offline state from the node that wrote it.
+ */
SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = redo_rec.data_checksum_version;
- SetLocalDataChecksumState(redo_rec.data_checksum_version);
- if (redo_rec.data_checksum_version != ControlFile->data_checksum_version)
- new_state = true;
+ if (!redo_rec.data_checksum_state_is_local &&
+ redo_rec.data_checksum_transition_lsn >
+ XLogCtl->data_checksum_transition_lsn)
+ {
+ if (XLogCtl->data_checksum_version != redo_rec.data_checksum_version)
+ new_state = true;
+ XLogCtl->data_checksum_version = redo_rec.data_checksum_version;
+ XLogCtl->data_checksum_transition_lsn = redo_rec.data_checksum_transition_lsn;
+ XLogCtl->data_checksum_state_is_local = false;
+ SetLocalDataChecksumState(redo_rec.data_checksum_version);
+ }
+
+ replayedCheckpointDataChecksumState = XLogCtl->data_checksum_version;
+ replayedCheckpointDataChecksumTransitionLSN = XLogCtl->data_checksum_transition_lsn;
+ replayedCheckpointDataChecksumStateIsLocal = XLogCtl->data_checksum_state_is_local;
SpinLockRelease(&XLogCtl->info_lck);
+ replayedCheckpointDataChecksumStateValid = true;
if (new_state)
EmitAndWaitDataChecksumsBarrier(redo_rec.data_checksum_version);
@@ -9249,15 +9377,28 @@ xlog2_redo(XLogReaderState *record)
if (info == XLOG2_CHECKSUMS)
{
xl_checksum_state state;
+ bool apply;
memcpy(&state, XLogRecGetData(record), sizeof(xl_checksum_state));
SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = state.new_checksum_state;
+ apply = record->EndRecPtr > XLogCtl->data_checksum_transition_lsn;
+
+ if (apply)
+ {
+ XLogCtl->data_checksum_version = state.new_checksum_state;
+ XLogCtl->data_checksum_transition_lsn = record->EndRecPtr;
+ XLogCtl->data_checksum_state_is_local = false;
+ }
SpinLockRelease(&XLogCtl->info_lck);
+ if (!apply)
+ return;
+
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = state.new_checksum_state;
+ ControlFile->data_checksum_transition_lsn = record->EndRecPtr;
+ ControlFile->data_checksum_state_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
diff --git a/src/backend/access/transam/xlogrecovery.c b/src/backend/access/transam/xlogrecovery.c
index 6de13b91748..96ff2749e1d 100644
--- a/src/backend/access/transam/xlogrecovery.c
+++ b/src/backend/access/transam/xlogrecovery.c
@@ -453,11 +453,13 @@ EnableStandbyMode(void)
* PerformWalRecovery().
*
* This initializes some global variables like ArchiveRecoveryRequested, and
- * StandbyModeRequested and InRecovery.
+ * StandbyModeRequested and InRecovery. The output parameters describe the
+ * selected checkpoint and any backup metadata.
*/
void
InitWalRecovery(ControlFileData *ControlFile, bool *wasShutdown_ptr,
- bool *haveBackupLabel_ptr, bool *haveTblspcMap_ptr)
+ bool *haveBackupLabel_ptr, bool *haveTblspcMap_ptr,
+ bool *backupFromStandby_ptr)
{
XLogPageReadPrivate *private;
struct stat st;
@@ -982,6 +984,7 @@ InitWalRecovery(ControlFileData *ControlFile, bool *wasShutdown_ptr,
*wasShutdown_ptr = wasShutdown;
*haveBackupLabel_ptr = haveBackupLabel;
*haveTblspcMap_ptr = haveTblspcMap;
+ *backupFromStandby_ptr = backupFromStandby;
}
/*
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 256b3a3c02e..c0a38640c46 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -371,6 +371,7 @@ WaitLSN "Waiting to read or update shared Wait-for-LSN state."
LogicalDecodingControl "Waiting to read or update logical decoding status information."
DataChecksumsWorker "Waiting for data checksums worker."
AioWorkerControl "Waiting to update AIO worker information."
+DataChecksumsState "Waiting to read or update data checksum state."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 3b3ae23f1a6..fa645d3af3a 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -647,6 +647,7 @@ main(int argc, char *argv[])
{
ControlFile->data_checksum_version =
(mode == PG_MODE_ENABLE) ? PG_DATA_CHECKSUM_VERSION : PG_DATA_CHECKSUM_OFF;
+ ControlFile->data_checksum_state_is_local = true;
if (do_sync)
{
diff --git a/src/bin/pg_resetwal/pg_resetwal.c b/src/bin/pg_resetwal/pg_resetwal.c
index 1542a56ca4b..fd012c10c46 100644
--- a/src/bin/pg_resetwal/pg_resetwal.c
+++ b/src/bin/pg_resetwal/pg_resetwal.c
@@ -914,6 +914,9 @@ RewriteControlFile(void)
XLogSegNoOffsetToRecPtr(newXlogSegNo, SizeOfXLogLongPHD, WalSegSz,
ControlFile.checkPointCopy.redo);
ControlFile.checkPointCopy.time = (pg_time_t) time(NULL);
+ ControlFile.checkPointCopy.dataChecksumState = ControlFile.data_checksum_version;
+ ControlFile.checkPointCopy.dataChecksumTransitionLSN = ControlFile.data_checksum_transition_lsn;
+ ControlFile.checkPointCopy.dataChecksumStateIsLocal = ControlFile.data_checksum_state_is_local;
ControlFile.state = DB_SHUTDOWNED;
ControlFile.checkPoint = ControlFile.checkPointCopy.redo;
diff --git a/src/include/access/xlog_internal.h b/src/include/access/xlog_internal.h
index be718993401..494748962f2 100644
--- a/src/include/access/xlog_internal.h
+++ b/src/include/access/xlog_internal.h
@@ -315,6 +315,8 @@ typedef struct xl_checkpoint_redo
{
int wal_level;
uint32 data_checksum_version;
+ XLogRecPtr data_checksum_transition_lsn;
+ bool data_checksum_state_is_local;
} xl_checkpoint_redo;
/*
diff --git a/src/include/access/xlogrecovery.h b/src/include/access/xlogrecovery.h
index a1d8a81dbc1..46d7cedd518 100644
--- a/src/include/access/xlogrecovery.h
+++ b/src/include/access/xlogrecovery.h
@@ -155,7 +155,8 @@ extern PGDLLIMPORT bool StandbyMode;
extern void InitWalRecovery(ControlFileData *ControlFile,
bool *wasShutdown_ptr, bool *haveBackupLabel_ptr,
- bool *haveTblspcMap_ptr);
+ bool *haveTblspcMap_ptr,
+ bool *backupFromStandby_ptr);
extern void PerformWalRecovery(void);
/*
diff --git a/src/include/catalog/pg_control.h b/src/include/catalog/pg_control.h
index 80b3a730e03..da5e2bb369d 100644
--- a/src/include/catalog/pg_control.h
+++ b/src/include/catalog/pg_control.h
@@ -64,8 +64,10 @@ typedef struct CheckPoint
*/
TransactionId oldestActiveXid;
- /* data checksums state at the time of the checkpoint */
+ /* data checksum state and latest online transition */
uint32 dataChecksumState;
+ XLogRecPtr dataChecksumTransitionLSN;
+ bool dataChecksumStateIsLocal; /* set by offline pg_checksums */
} CheckPoint;
/* XLOG info values for XLOG rmgr */
@@ -228,8 +230,10 @@ typedef struct ControlFileData
bool float8ByVal; /* float8, int8, etc pass-by-value? */
- /* Are data pages protected by checksums? Zero if no checksum version */
+ /* Current data checksum state and latest online transition */
uint32 data_checksum_version;
+ XLogRecPtr data_checksum_transition_lsn;
+ bool data_checksum_state_is_local; /* set by offline pg_checksums */
/*
* True if the default signedness of char is "signed" on a platform where
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index d7eb648bd27..0b18d6f4198 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -89,6 +89,7 @@ PG_LWLOCK(54, WaitLSN)
PG_LWLOCK(55, LogicalDecodingControl)
PG_LWLOCK(56, DataChecksumsWorker)
PG_LWLOCK(57, AioWorkerControl)
+PG_LWLOCK(58, DataChecksumsState)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index 9b1421a9b91..fb84f678a3a 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -33,6 +33,7 @@ tests += {
't/007_pgbench_standby.pl',
't/008_pitr.pl',
't/009_fpi.pl',
+ 't/010_offline_standby.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/010_offline_standby.pl b/src/test/modules/test_checksums/t/010_offline_standby.pl
new file mode 100644
index 00000000000..066541d84c4
--- /dev/null
+++ b/src/test/modules/test_checksums/t/010_offline_standby.pl
@@ -0,0 +1,87 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test interactions between offline checksum changes and streaming replication.
+# An offline change must remain local to its data directory: a standby must not
+# adopt the primary's offline change, and its own offline change must survive
+# WAL replay and restarts.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->start;
+
+$primary->backup('backup');
+
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_replay_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# An offline change on the standby must override online transitions it has
+# already replayed.
+enable_data_checksums($primary, wait => 'on');
+wait_for_checksum_state($standby, 'on');
+$primary->wait_for_replay_catchup($standby);
+
+$standby->stop;
+$standby->checksum_disable_offline;
+$standby->start;
+test_checksum_state($standby, 'off');
+
+# Return both nodes to the initial state for the next test.
+disable_data_checksums($primary, wait => 1);
+$primary->wait_for_replay_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+$standby->stop;
+$primary->stop;
+
+# This changes only the primary's pages and control file.
+$primary->checksum_enable_offline;
+$primary->start;
+test_checksum_state($primary, 'on');
+
+# Publish the primary's local state in a checkpoint, then make that checkpoint
+# the starting point for crash recovery on the same node.
+$primary->safe_psql('postgres', 'CHECKPOINT');
+$primary->stop('immediate');
+$primary->start;
+test_checksum_state($primary, 'on');
+
+# The standby must not adopt the primary's offline transition while replaying
+# the next checkpoint.
+$primary->safe_psql('postgres', 'CHECKPOINT');
+$standby->start;
+$primary->wait_for_replay_catchup($standby);
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', 'SELECT count(*) FROM pg_class') > 0,
+ 1, 'standby remains readable');
+
+# A clean stop creates a restartpoint. Its localized checksum state must
+# survive both a normal restart and subsequent crash recovery.
+$standby->stop;
+$standby->start;
+test_checksum_state($standby, 'off');
+
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+$standby->stop;
+$primary->stop;
+
+done_testing();
--
2.34.1
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-14 13:59 Daniel Gustafsson <daniel@yesql.se>
parent: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 1 reply; 43+ messages in thread
From: Daniel Gustafsson @ 2026-08-14 13:59 UTC (permalink / raw)
To: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; Zsolt Parragi <zsolt.parragi@percona.com>
> On 12 Aug 2026, at 09:55, Bertrand Drouvot <bertranddrouvot.pg@gmail.com> wrote:
> First one, is a case where a standby can enable checksum verification without its
> own pages having been checksummed, making the standby unreadable.
Thanks for the report. While I don't have a proposal ready at this time of
writing, I wanted to ACK having seen this and make it known (to RMT) that it is
being worked on.
> Therefore, I think a simpler fix is to preserve the pre-f19c0eccae96 behavior
> for offline checksum changes: they are not propagated through WAL. In the second
> repro, the standby therefore remains off, honoring its local offline change.
I wholeheartedly disagree, running a cluster with mismatched data_checksums
settings across the nodes is not a supported mode of operation, and is already
documented to not work (albeit it way too vague wording IMO). This doesn't
work as it is right now (in any version of postgres), pg_rewind or other file
based tools can break it, and we should not attempt to make it work.
Detecting a cluster with mismatched settings and safely erroring out as well as
improving the documentation is what I think we should do.
> If this looks like too much code changes so close to the v19 release, another
> option could be to remove pg_checksums --enable and --disable while keeping --check
> and require checksum state changes to be done online.
That's also not a good option, I think we need to make sure offline enabling of
checksums *if done correctly* works as intended, and if done incorrectly errors
out safely.
I have a patch proposal brewing, and I know Zsolt has been looking into it as
well. Hopefully there will be something to share very soon.
--
Daniel Gustafsson
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-14 15:27 Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
parent: Daniel Gustafsson <daniel@yesql.se>
0 siblings, 1 reply; 43+ messages in thread
From: Bertrand Drouvot @ 2026-08-14 15:27 UTC (permalink / raw)
To: Daniel Gustafsson <daniel@yesql.se>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; Zsolt Parragi <zsolt.parragi@percona.com>
Hi,
On Fri, Aug 14, 2026 at 03:59:06PM +0200, Daniel Gustafsson wrote:
> running a cluster with mismatched data_checksums
> settings across the nodes is not a supported mode of operation, and is already
> documented to not work
Thanks for feedback!
The pre f19c0eccae96 pg_checksums documentation said:
"
When using a replication setup with tools which perform direct copies
of relation file blocks (for example pg_rewind), enabling or disabling
checksums can lead to page corruptions in the shape of incorrect
checksums if the operation is not done consistently across all nodes.
"
I read that as a recommendation to stop and switch all nodes consistently, and
as a warning about direct block-copy tools. That interpretation, together with
the pre f19c0eccae96 behavior, is why v1 proposed keeping offline pg_checksums
changes local. The intention was to preserve the previous behavior, not to
introduce a new supported mode.
I just realized that f19c0eccae96 explicitly changed the "Off-line Enabling of
Checksums" documentation:
"
Data checksums are enabled or disabled at the full cluster level, and cannot
be specified individually for databases or tables.
"
by:
"
Data checksums are enabled or disabled at the full cluster level, and cannot
be specified individually for databases, tables or replicated cluster members.
"
while leaving the pg_checksums documentation quoted above unchanged.
Depending on how the issue will be addressed, it might be worth changing this
pg_checksums wording too?
Looking forward to seeing your and Zsolt's proposals.
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-26 20:37 Daniel Gustafsson <daniel@yesql.se>
parent: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 1 reply; 43+ messages in thread
From: Daniel Gustafsson @ 2026-08-26 20:37 UTC (permalink / raw)
To: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; Zsolt Parragi <zsolt.parragi@percona.com>
> On 14 Aug 2026, at 17:27, Bertrand Drouvot <bertranddrouvot.pg@gmail.com> wrote:
> Looking forward to seeing your and Zsolt's proposals.
This has now been worked on quite extensively by Zsolt, myself and Tomas Vondra
and a number of patchrevisions have been created and rewritten. There are two
separate issues in this report: a) a correctly done offline checksum change in
a replicated cluster doesn't work; b) mismatched checksum states across
replicated nodes is not detected and will make pg_rewind and similar tools
dangerous. The former is a regression due to the online checksums work, the
latter is an issue which exists in all supported versions due to how
pg_checksums was implemented. This is today mentioned very briefly in the docs
but clearly hasn't been looked into to fix. Below are each issue discussed in
more detail.
The regression in a correctly done offline change is due to combining
pg_checksums which rewrite data on risk without WAL logging (or any logging at
all) the transformation, with online checksums which WAL log the state change.
With online checksums, the local state on the standby was overwritten during
replay by the dataChecksumState in the checkpoint. By definition, the
checkpoint will be from before the offline change and thus enabling checksums
would replay checksums being disabled. The fix in 0001 is to not adopt the
state change from the replay of checkpoints, only from XLOG2_CHECKSUMS records,
and to alert the user with a log entry if the states mismatch.
Detecting a state mismatch, and refusing to start a standby which does not
match the primary is a lot harder than it may seem. We have had a few
different patches implementing this and they all have the flaw that by the time
the standby can be shut down due to mismatch, it needs to be rebuilt from a
base backup and cannot be recovered with the (presumably) missing pg_checksums
command. Due to this, the current approach is to log a WARNING for mismatched
states, which while not perfect improves upon what we have today in v14 through
v18 where it's silently ignored.
There are also few more commits in this patchset related to issue B):
0002 makes pg_checksums refuse to operate on standbys where the state hasn't
been resolved from an online checksums change. On a primary, the state will
heal itself upon startup and pg_checksums refuse a crashed primary already.
0003-0004 fixes pg_rewind and pg_combinebackup to error out on mismatched
states instead of risk damaging data. A variant of these fixes should be
backpatched into all supported versions are the issue is present with offline
checksums.
So why wasn't this regression caught before feature freeze? The main reason is
that I failed to add test cases for offline checksum changes in replicated
clusters when I wrote the online checksums patch. It contains tests for
offline change of a single primary which is the easier case to handle. The
pg_checksums test suite also doesn't test the replicated scenario at all. Even
if online checksums end up reverted, tests for replicated clusters should be
added to pg_checksums as it currently lacks test coverage.
The 0001 patch is the least invasive patch to solve the regression that either
of us has managed to come up with, but it's still far from trivial. The plan
going forward for this hinges on whether or not online checksums get reverted,
but here is at least a patchset addressing the open item for future reference.
Should this get committed we probably need to gate a few tests under
PG_TEST_EXTRA to keep things at a reasonable scale, but for now they are all
left in the main path.
--
Daniel Gustafsson
Attachments:
[application/octet-stream] v4-0001-Do-not-adopt-data-checksum-state-from-another-nod.patch (96.0K, ../../188A1307-1A92-45CA-9DBE-FB0962D3756E@yesql.se/2-v4-0001-Do-not-adopt-data-checksum-state-from-another-nod.patch)
download | inline diff:
From 2097faf164d5f57222187e5e644bd53e0cb9bbb1 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:02 +0000
Subject: [PATCH v4 1/4] Do not adopt data checksum state from another node
during replay
Offline pg_checksums changes are local to one node, but replay adopted
the data checksum state carried by checkpoint records unconditionally.
After an offline enable on the primary, a standby whose pages were
never rewritten started verifying checksums it does not have; after an
offline change on a standby, the next replayed checkpoint silently
reverted it.
Instead, cross-check the replayed state against the local one and warn
once per divergent value. When the states match again, say so in the
log and re-arm the warning.
The control file now tracks this node's state alone, which makes the
moment it is written out part of the design: it may only claim "on"
once every page on disk carries a checksum, or a crash-restart would
verify pages that were flushed before the transition rewrote them.
XLOG2_CHECKSUMS replay therefore persists every state but "on" as soon
as it is replayed, none of them verifying anything, and leaves "on" to
the flush of the next restartpoint. A restartpoint only persists a
state its flush ran under from beginning to end, a promotion writes a
full checkpoint instead of an end-of-recovery record while the state is
still unpersisted, and a shutdown restartpoint with no new checkpoint
record to work from flushes before catching the control file up rather
than leaving it behind. Enabling checksums on the primary defers the
same write to after the checkpoint that flushes the rewritten pages:
the ones the worker found in shared buffers do not go out through its
ring buffer. A checkpoint persists the state its own redo point ran
under, under the same rule a restartpoint follows, since the checkpoint
that licenses "on" also moves the redo point past the record announcing
it: without that, a crash before the deferred write would resume above
the record and resolve a finished transition as interrupted. Checkpoint
records replayed below the consistency point are not cross-checked, the
persisted state being legitimately newer than what they carry.
A base backup copies the control file at an arbitrary moment, so its
state can be newer than the redo point replay starts from. Adopt the
state carried by the starting checkpoint record under backup-label
recovery; for a shutdown checkpoint, which replay does not see, take
it in StartupXLOG. This also covers pg_rewind, which installs the
source's control file while replay begins at the last common
checkpoint. Backups taken from a standby are the exception: their
starting checkpoint was written by the upstream primary, so they keep
the state of the control file that was copied with them.
The documented procedure for offline changes in a replication setup
becomes the lockstep one: stop all nodes, run pg_checksums on each of
them, then restart. Tests cover the lockstep procedure, divergent
offline changes on either node and down a cascading chain,
crash-restarts and promotions around online transitions, checkpoints
racing an online enable, a base backup taken during an online enable,
and pg_rewind across one.
---
doc/src/sgml/ref/pg_checksums.sgml | 32 +-
doc/src/sgml/wal.sgml | 8 +
src/backend/access/transam/xlog.c | 358 ++++++++++++++++--
src/test/modules/test_checksums/Makefile | 2 +-
src/test/modules/test_checksums/meson.build | 14 +
.../test_checksums/t/012_offline_standby.pl | 213 +++++++++++
.../modules/test_checksums/t/013_rewind.pl | 198 ++++++++++
.../modules/test_checksums/t/014_lockstep.pl | 89 +++++
.../t/015_backup_online_enable.pl | 71 ++++
.../t/016_backup_from_standby.pl | 100 +++++
.../t/017_standby_crash_after_disable.pl | 125 ++++++
.../t/018_promote_enable_crash.pl | 134 +++++++
.../test_checksums/t/019_restartpoint_race.pl | 138 +++++++
.../t/020_primary_enable_crash.pl | 118 ++++++
.../t/021_standby_shutdown_catchup.pl | 93 +++++
.../t/022_resident_enable_crash.pl | 138 +++++++
.../t/023_concurrent_checkpoint_enable.pl | 114 ++++++
.../t/024_enable_crash_after_checkpoint.pl | 101 +++++
.../t/025_cascade_divergence.pl | 115 ++++++
19 files changed, 2122 insertions(+), 39 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/012_offline_standby.pl
create mode 100644 src/test/modules/test_checksums/t/013_rewind.pl
create mode 100644 src/test/modules/test_checksums/t/014_lockstep.pl
create mode 100644 src/test/modules/test_checksums/t/015_backup_online_enable.pl
create mode 100644 src/test/modules/test_checksums/t/016_backup_from_standby.pl
create mode 100644 src/test/modules/test_checksums/t/017_standby_crash_after_disable.pl
create mode 100644 src/test/modules/test_checksums/t/018_promote_enable_crash.pl
create mode 100644 src/test/modules/test_checksums/t/019_restartpoint_race.pl
create mode 100644 src/test/modules/test_checksums/t/020_primary_enable_crash.pl
create mode 100644 src/test/modules/test_checksums/t/021_standby_shutdown_catchup.pl
create mode 100644 src/test/modules/test_checksums/t/022_resident_enable_crash.pl
create mode 100644 src/test/modules/test_checksums/t/023_concurrent_checkpoint_enable.pl
create mode 100644 src/test/modules/test_checksums/t/024_enable_crash_after_checkpoint.pl
create mode 100644 src/test/modules/test_checksums/t/025_cascade_divergence.pl
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index ae66fad3f0f..bf07da09754 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -242,15 +242,29 @@ PostgreSQL documentation
data directory must not be started or else data loss may occur.
</para>
<para>
- When using a replication setup with tools which perform direct copies
- of relation file blocks (for example <xref linkend="app-pgrewind"/>),
- enabling or disabling checksums can lead to page corruptions in the
- shape of incorrect checksums if the operation is not done consistently
- across all nodes. When enabling or disabling checksums in a replication
- setup, it is thus recommended to stop all the clusters before switching
- them all consistently. Destroying all standbys, performing the operation
- on the primary and finally recreating the standbys from scratch is also
- safe.
+ Enabling or disabling checksums with
+ <application>pg_checksums</application> changes only the local data
+ directory; the new state does not propagate over replication. In a
+ replication setup the same change must be applied to every node: stop all
+ nodes, run <application>pg_checksums</application> on each of them, and
+ only then restart them. Tools that copy relation file blocks directly
+ between nodes, such as <xref linkend="app-pgrewind"/>, likewise require
+ both nodes to be in the same data checksum state.
+ </para>
+ <para>
+ If the change is applied inconsistently, each node keeps its own state,
+ and a standby logs a warning when the state recorded in the replayed WAL
+ differs from its own. The same warning can appear transiently while a
+ standby catches up over WAL written before a consistent change; it stops
+ once a checkpoint record carrying the new state has been replayed.
+ </para>
+ <para>
+ A standby whose data directory was never checksummed must not have
+ checksums enabled by catching up this way. Converge the cluster by
+ running <application>pg_checksums</application> on it while stopped, by
+ enabling checksums online, or by recreating it from a base backup. Note
+ that an online enable only starts from a primary whose checksums are off,
+ so if they are already enabled there, disable them online first.
</para>
<para>
If <application>pg_checksums</application> is aborted or killed while
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index ec62d17fbcc..db18a8c516e 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -317,6 +317,14 @@
verify checksums, on an offline cluster.
</para>
+ <para>
+ An offline change only affects the data directory it is run on; the
+ new state does not propagate over replication. In a replication setup
+ the same change must be applied to all nodes while all of them are
+ stopped, as described in <xref linkend="app-pgchecksums"/>. A standby
+ whose state diverges logs a warning but keeps its local setting.
+ </para>
+
</sect2>
<sect2 id="checksums-online-enable-disable" xreflabel="Online Enabling of Checksums">
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index de4c96e135f..fa24b00e7d4 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -556,7 +556,7 @@ typedef struct XLogCtlData
*/
XLogRecPtr lastFpwDisableRecPtr;
- /* last data_checksum_version we've seen */
+ /* current data checksum state of this node */
uint32 data_checksum_version;
slock_t info_lck; /* locks shared variables shown above */
@@ -690,6 +690,14 @@ static ChecksumStateType LocalDataChecksumState = 0;
*/
int data_checksums = 0;
+/*
+ * Whether replay of the next checkpoint-family record must adopt the data
+ * checksum state it carries. Set when recovery starts from a base backup,
+ * where the state at the redo point takes precedence over the control file
+ * copied with the backup later.
+ */
+static bool adoptChecksumStateFromNextCheckpoint = false;
+
/* For WALInsertLockAcquire/Release functions */
static int MyLockNo = 0;
static bool holdingAllLocks = false;
@@ -730,6 +738,8 @@ static void ValidateXLOGDirectoryStructure(void);
static void CleanupBackupHistory(void);
static void UpdateMinRecoveryPoint(XLogRecPtr lsn, bool force);
static bool PerformRecoveryXLogAction(void);
+static void CheckReplayedDataChecksumState(uint32 replayed_version);
+static void AdoptReplayedDataChecksumState(uint32 new_version);
static void InitControlFile(uint64 sysidentifier, uint32 data_checksum_version);
static void WriteControlFile(void);
static void ReadControlFile(void);
@@ -4827,6 +4837,7 @@ void
SetDataChecksumsOn(void)
{
uint64 barrier;
+ bool persist;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4856,15 +4867,6 @@ SetDataChecksumsOn(void)
XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
SpinLockRelease(&XLogCtl->info_lck);
- /*
- * Update the controlfile before waiting since if we have an immediate
- * shutdown while waiting we want to come back up with checksums enabled.
- */
- LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- UpdateControlFile();
- LWLockRelease(ControlFileLock);
-
barrier = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_CHECKSUM_ON);
MyProc->delayChkptFlags &= ~DELAY_CHKPT_START;
@@ -4873,6 +4875,35 @@ SetDataChecksumsOn(void)
INJECTION_POINT("datachecksums-on-before-checkpoint", NULL);
RequestCheckpoint(CHECKPOINT_FORCE | CHECKPOINT_WAIT | CHECKPOINT_FAST);
+
+ INJECTION_POINT("datachecksums-on-after-checkpoint", NULL);
+
+ /*
+ * Persist "on" only now that the checkpoint has flushed the pages the
+ * transition rewrote. Crash recovery initializes verification from the
+ * control file but resumes from a checkpoint that can predate the
+ * transition, so an "on" persisted earlier would have replay verify pages
+ * whose rewrite never reached disk: the pages the worker found in shared
+ * buffers are not written back by its ring buffer.
+ *
+ * The checkpoint above normally persists the state itself; this covers the
+ * case where it started before the record written above and left the field
+ * alone. Crashing before this point is safe, as replay then re-establishes
+ * "on" from the full page images of the rewrite. Skip the write if the
+ * state moved on meanwhile, since whatever moved it persists its own.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ persist = (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_VERSION);
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (persist)
+ {
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+ }
+
WaitForProcSignalBarrier(barrier);
}
@@ -5001,6 +5032,124 @@ SetLocalDataChecksumState(uint32 data_checksum_version)
data_checksums = data_checksum_version;
}
+/*
+ * Cross-check the data checksum state carried by a replayed checkpoint record
+ * against the state of this node.
+ *
+ * The state in the record belongs to whichever node wrote the WAL, and must
+ * not be adopted: an offline state change made with pg_checksums on one node
+ * of a replication set generates no WAL, and adopting would leak it into the
+ * other nodes through replay. Backup label recovery is the one exception,
+ * see AdoptReplayedDataChecksumState(). XLOG_CHECKPOINT_ONLINE needs no call
+ * here, as an online checkpoint's state already traveled in the preceding
+ * XLOG_CHECKPOINT_REDO record.
+ *
+ * Only archive recovery can see a lasting mismatch, as only there can the WAL
+ * and the control file come from different nodes or different times. In
+ * crash recovery a mismatch means replay resumed from a restartpoint
+ * predating an already-applied XLOG2_CHECKSUMS record, and replaying forward
+ * re-establishes the same state.
+ */
+static void
+CheckReplayedDataChecksumState(uint32 replayed_version)
+{
+ /*
+ * Warn once per remote value, so a lasting mismatch does not flood the
+ * log. Matching states re-arm the warning. Backend-local state is
+ * enough: replay only runs in the startup process, and restarting it
+ * re-arms as well.
+ */
+ static uint32 last_warned_version = PG_UINT32_MAX;
+ uint32 local_version;
+
+ if (!ArchiveRecoveryRequested)
+ return;
+
+ /*
+ * Re-replayed WAL below the consistency point was already cross-checked
+ * before minRecoveryPoint was last persisted, and the persisted state can
+ * legitimately be newer than what checkpoint records there carry:
+ * XLOG2_CHECKSUMS replay persists most states ahead of the restartpoint
+ * horizon. In particular the checkpoint record recovery restarts from is
+ * such a re-replay.
+ */
+ if (!reachedConsistency)
+ return;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ local_version = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (replayed_version == local_version)
+ {
+ /*
+ * Report convergence if this process warned before. Nothing else
+ * tells the operator that running pg_checksums on the other nodes, or
+ * a rebuild, took effect. A restart in between loses the context,
+ * but it re-arms the warning too.
+ */
+ if (last_warned_version != PG_UINT32_MAX)
+ ereport(LOG,
+ errmsg("data checksum state \"%s\" of this node now agrees with the replayed WAL",
+ get_checksum_state_string(local_version)));
+
+ last_warned_version = PG_UINT32_MAX;
+ return;
+ }
+
+ /* the nodes legitimately differ while an online transition runs */
+ if (replayed_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ replayed_version == PG_DATA_CHECKSUM_INPROGRESS_OFF ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ return;
+
+ if (replayed_version == last_warned_version)
+ return;
+ last_warned_version = replayed_version;
+
+ ereport(WARNING,
+ errmsg("data checksum state \"%s\" of this node does not match the state \"%s\" in the replayed WAL",
+ get_checksum_state_string(local_version),
+ get_checksum_state_string(replayed_version)),
+ errdetail("The data checksum state was most likely changed with pg_checksums on another node."),
+ errhint("Apply the same change with pg_checksums on the primary and all standby servers, or rebuild this server from a base backup."));
+}
+
+/*
+ * Adopt the data checksum state found at the redo point of backup label
+ * recovery. Persist it immediately so that a crash before the first
+ * restartpoint does not resurrect the state copied with the backup; a crash
+ * at this point restarts from the same redo point, so the control file does
+ * not run ahead of the replay position. If the value is unchanged the
+ * control file already carries it, so both the barrier and the persist are
+ * skipped.
+ */
+static void
+AdoptReplayedDataChecksumState(uint32 new_version)
+{
+ bool changed = false;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (XLogCtl->data_checksum_version != new_version)
+ {
+ XLogCtl->data_checksum_version = new_version;
+ SetLocalDataChecksumState(new_version);
+ changed = true;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (!changed)
+ return;
+
+ EmitAndWaitDataChecksumsBarrier(new_version);
+
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = new_version;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+}
+
/* guc hook */
const char *
show_data_checksums(void)
@@ -6031,6 +6180,39 @@ StartupXLOG(void)
SetCommitTsLimit(checkPoint.oldestCommitTsXid,
checkPoint.newestCommitTsXid);
+ /*
+ * When recovery starts from a base backup, the control file was copied at
+ * an arbitrary moment and its data checksum state may differ from the
+ * state at the redo point, which is what the WAL from there on was
+ * written under. Adopt the state of the starting checkpoint: a shutdown
+ * checkpoint is not replayed, so take it from the record read above; the
+ * redo point of an online checkpoint is its CHECKPOINT_REDO record, so
+ * let the replay of that record adopt it. Check backupStartPoint in
+ * addition to the label: on a crash restart during backup recovery the
+ * label file is already renamed away, but the start point persists until
+ * the backup end record.
+ *
+ * Not for a base backup taken from a standby, though. Its starting
+ * checkpoint is the standby's last restartpoint, a record written by the
+ * upstream primary, whose state is not the one the copied files were
+ * written under. The copied control file is right there, as a standby
+ * persists its state only at restartpoint horizons and so never claims
+ * more than what reached disk. Such backups are recognized by
+ * backupEndPoint together with backupEndRequired; backupEndPoint is only
+ * set for "BACKUP FROM: standby" labels and persists across a crash
+ * restart. pg_rewind writes a standby label as well, but no
+ * backupEndPoint, and its recovery must keep adopting.
+ */
+ if ((haveBackupLabel || XLogRecPtrIsValid(ControlFile->backupStartPoint)) &&
+ !(XLogRecPtrIsValid(ControlFile->backupEndPoint) &&
+ ControlFile->backupEndRequired))
+ {
+ if (wasShutdown)
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState);
+ else
+ adoptChecksumStateFromNextCheckpoint = true;
+ }
+
/*
* Clear out any old relcache cache files. This is *necessary* if we do
* any WAL replay, since that would probably result in the cache files
@@ -6814,6 +6996,23 @@ static bool
PerformRecoveryXLogAction(void)
{
bool promoted = false;
+ bool flushForChecksums;
+ uint32 checksum_state;
+
+ /*
+ * The end-of-recovery record persists the data checksum state without
+ * flushing the buffer pool, but the control file may only claim "on" once
+ * every page on disk carries a checksum. If replay entered that state
+ * without a restartpoint following it, the pages rewritten by the
+ * transition are still only in the buffer pool, so take the full
+ * checkpoint below instead of the lightweight record.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ flushForChecksums = (checksum_state == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile->data_checksum_version != checksum_state);
/*
* Perform a checkpoint to update all our recovery activity to disk.
@@ -6829,7 +7028,7 @@ PerformRecoveryXLogAction(void)
* fully out of recovery mode and already accepting queries.
*/
if (ArchiveRecoveryRequested && IsUnderPostmaster &&
- PromoteIsTriggered())
+ PromoteIsTriggered() && !flushForChecksums)
{
promoted = true;
@@ -7823,6 +8022,32 @@ CreateCheckPoint(int flags)
ControlFile->minRecoveryPoint = InvalidXLogRecPtr;
ControlFile->minRecoveryPointTLI = 0;
+ /*
+ * Persist the data checksum state this node runs under. Only the
+ * top-level field tracks this node; ControlFile->checkPointCopy above is
+ * a historical record used to resume replay.
+ *
+ * checkPoint.dataChecksumState was sampled while holding the WAL insert
+ * locks, so it is the state in effect at the redo point. If it was "on",
+ * the XLOG2_CHECKSUMS record announcing that precedes the redo point and
+ * every page the transition rewrote was dirtied before it, so
+ * CheckPointGuts() has just written all of them out. Recording the state
+ * here is what keeps a finished transition from being resolved as
+ * interrupted when this checkpoint is the one crash recovery resumes
+ * from: replay never sees the record announcing it.
+ *
+ * Persist it only if the state did not change while the flush was in
+ * progress. If it changed in between, the pages written out straddle two
+ * states, and the newer one could claim checksums that pages already on
+ * disk do not carry; leave the field to the next checkpoint then.
+ * SetDataChecksumsOff() persists the states that are safe to enter
+ * without a flush already.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checkPoint.dataChecksumState == XLogCtl->data_checksum_version)
+ ControlFile->data_checksum_version = checkPoint.dataChecksumState;
+ SpinLockRelease(&XLogCtl->info_lck);
+
/*
* Persist unloggedLSN value. It's reset on crash recovery, so this goes
* unused on non-shutdown checkpoints, but seems useful to store it always
@@ -7967,7 +8192,7 @@ CreateEndOfRecoveryRecord(void)
ControlFile->minRecoveryPoint = recptr;
ControlFile->minRecoveryPointTLI = xlrec.ThisTimeLineID;
- /* start with the latest checksum version (as of the end of recovery) */
+ /* persist the data checksum state this node ended recovery with */
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
SpinLockRelease(&XLogCtl->info_lck);
@@ -8174,6 +8399,7 @@ CreateRestartPoint(int flags)
XLogRecPtr endptr;
XLogSegNo _logSegNo;
TimestampTz xtime;
+ uint32 checksum_state;
/* Concurrent checkpoint/restartpoint cannot happen */
Assert(!IsUnderPostmaster || MyBackendType == B_CHECKPOINTER);
@@ -8220,8 +8446,39 @@ CreateRestartPoint(int flags)
UpdateMinRecoveryPoint(InvalidXLogRecPtr, true);
if (flags & CHECKPOINT_IS_SHUTDOWN)
{
+ bool catchUpChecksums;
+
+ /*
+ * There is no new restartpoint to persist the data checksum state
+ * with, but a cleanly stopped node should not leave the control
+ * file behind the state replay reached: pg_checksums and
+ * pg_rewind read it, and an in-progress state there makes them
+ * refuse to run. Catching it up needs the same guarantee a
+ * restartpoint gives: that every page on disk carries a checksum,
+ * so flush the buffer pool first. Only a transition to
+ * "on" that no restartpoint followed can get here; the other
+ * states are already persisted by XLOG2_CHECKSUMS replay. Replay
+ * has ended by now, so the state cannot change under us.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ catchUpChecksums =
+ (checksum_state != ControlFile->data_checksum_version &&
+ XLogRecPtrIsValid(lastCheckPointRecPtr));
+
+ if (catchUpChecksums)
+ {
+ MemSet(&CheckpointStats, 0, sizeof(CheckpointStats));
+ CheckpointStats.ckpt_start_t = GetCurrentTimestamp();
+ CheckPointGuts(lastCheckPoint.redo, flags);
+ }
+
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
+ if (catchUpChecksums)
+ ControlFile->data_checksum_version = checksum_state;
UpdateControlFile();
LWLockRelease(ControlFileLock);
}
@@ -8262,6 +8519,16 @@ CreateRestartPoint(int flags)
/* Update the process title */
update_checkpoint_display(flags, true, false);
+ /*
+ * Note the data checksum state the flush below starts under. Replay runs
+ * concurrently and can change the state while the flush is in progress,
+ * in which case the flush covers pages written under both states; see
+ * where the state is persisted further down.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
CheckPointGuts(lastCheckPoint.redo, flags);
/*
@@ -8321,8 +8588,21 @@ CreateRestartPoint(int flags)
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
}
- /* we shall start with the latest checksum version */
- ControlFile->data_checksum_version = lastCheckPoint.dataChecksumState;
+ /*
+ * Persist the data checksum state of this node. Not the state of the
+ * replayed checkpoint: that one belongs to the node that wrote it and
+ * may differ after an offline change on either side.
+ * ControlFile->checkPointCopy above keeps the replayed value on
+ * purpose, being a historical record used to resume replay rather
+ * than a tracker of node state.
+ *
+ * Persist only if the flush above ran under one state throughout; see
+ * CreateCheckPoint() for why.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksum_state == XLogCtl->data_checksum_version)
+ ControlFile->data_checksum_version = checksum_state;
+ SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
}
@@ -8966,11 +9246,18 @@ xlog_redo(XLogReaderState *record)
/* ControlFile->checkPointCopy always tracks the latest ckpt XID */
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->checkPointCopy.nextXid = checkPoint.nextXid;
- ControlFile->data_checksum_version = checkPoint.dataChecksumState;
UpdateControlFile();
LWLockRelease(ControlFileLock);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState);
+ }
+ else
+ CheckReplayedDataChecksumState(checkPoint.dataChecksumState);
+
/*
* We should've already switched to the new TLI before replaying this
* record.
@@ -9206,19 +9493,16 @@ xlog_redo(XLogReaderState *record)
else if (info == XLOG_CHECKPOINT_REDO)
{
xl_checkpoint_redo redo_rec;
- bool new_state = false;
memcpy(&redo_rec, XLogRecGetData(record), sizeof(xl_checkpoint_redo));
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = redo_rec.data_checksum_version;
- SetLocalDataChecksumState(redo_rec.data_checksum_version);
- if (redo_rec.data_checksum_version != ControlFile->data_checksum_version)
- new_state = true;
- SpinLockRelease(&XLogCtl->info_lck);
-
- if (new_state)
- EmitAndWaitDataChecksumsBarrier(redo_rec.data_checksum_version);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(redo_rec.data_checksum_version);
+ }
+ else
+ CheckReplayedDataChecksumState(redo_rec.data_checksum_version);
}
else if (info == XLOG_LOGICAL_DECODING_STATUS_CHANGE)
{
@@ -9288,17 +9572,33 @@ xlog2_redo(XLogReaderState *record)
SpinLockAcquire(&XLogCtl->info_lck);
XLogCtl->data_checksum_version = state.new_checksum_state;
+ SetLocalDataChecksumState(state.new_checksum_state);
SpinLockRelease(&XLogCtl->info_lck);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = state.new_checksum_state;
+
+ /*
+ * Persist the new state, except when it is "on". Only "on" verifies
+ * checksums during reads, and between the last restartpoint and this
+ * record there may be pages on disk flushed under the old state; a
+ * crash-restart initializes verification from the control file and
+ * replay reads those pages back, so the control file may only say
+ * "on" once everything written under the transition has been flushed,
+ * as restartpoints and the end of recovery do.
+ * The opposite direction cannot wait for the restartpoint: once this
+ * record is replayed, evicted pages are written without checksums,
+ * and a control file still saying "on" would fail verification on
+ * exactly those pages after a crash.
+ */
+ if (state.new_checksum_state != PG_DATA_CHECKSUM_VERSION)
+ ControlFile->data_checksum_version = state.new_checksum_state;
/*
* Update minRecoveryPoint to ensure that if recovery is aborted, we
* recover back up to this point before allowing hot standby again.
- * The new state is durable in pg_control while its location is only
- * tracked in shared memory; a standby becoming consistent below this
- * record would let base backups resume checksum verification with the
+ * The change location is only tracked in shared memory and is lost
+ * over a restart; a standby becoming consistent below this record
+ * would let base backups resume checksum verification with the
* location unknown. The local copies cannot be updated as long as
* crash recovery is happening and we expect all the WAL to be
* replayed.
diff --git a/src/test/modules/test_checksums/Makefile b/src/test/modules/test_checksums/Makefile
index 71455cd5577..80f54bf6d8c 100644
--- a/src/test/modules/test_checksums/Makefile
+++ b/src/test/modules/test_checksums/Makefile
@@ -9,7 +9,7 @@
#
#-------------------------------------------------------------------------
-EXTRA_INSTALL = src/test/modules/injection_points
+EXTRA_INSTALL = contrib/pg_buffercache src/test/modules/injection_points
export enable_injection_points
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index fb7129d796f..0269d4c6e38 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -35,6 +35,20 @@ tests += {
't/009_fpi.pl',
't/010_backup_straddle.pl',
't/011_standby_straddle.pl',
+ 't/012_offline_standby.pl',
+ 't/013_rewind.pl',
+ 't/014_lockstep.pl',
+ 't/015_backup_online_enable.pl',
+ 't/016_backup_from_standby.pl',
+ 't/017_standby_crash_after_disable.pl',
+ 't/018_promote_enable_crash.pl',
+ 't/019_restartpoint_race.pl',
+ 't/020_primary_enable_crash.pl',
+ 't/021_standby_shutdown_catchup.pl',
+ 't/022_resident_enable_crash.pl',
+ 't/023_concurrent_checkpoint_enable.pl',
+ 't/024_enable_crash_after_checkpoint.pl',
+ 't/025_cascade_divergence.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
new file mode 100644
index 00000000000..94c476293b7
--- /dev/null
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -0,0 +1,213 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Offline checksum changes with pg_checksums are local to one node. A
+# standby must neither adopt the state of the primary from replayed
+# checkpoint records, nor lose its own offline change to them.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', 'autovacuum = off');
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Scenario 1: enable offline on the primary only. The standby must
+# stay off, warn about the mismatch, and remain readable.
+$standby->stop;
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+test_checksum_state($primary, 'on');
+test_checksum_state($standby, 'off');
+
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after offline enable on the primary');
+
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+my @warnings = $log =~ /(does not match the state)/g;
+is(scalar(@warnings), 1, 'mismatch warned once per remote value');
+
+# Matching states re-arm the warning: undo the divergence on the primary,
+# then diverge again to the same value, all without restarting the standby.
+$primary->stop;
+$primary->checksum_disable_offline;
+$logstart = -s $standby->logfile;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no warning while the states match again');
+like($log, qr/now agrees with the replayed WAL/,
+ 'convergence is reported once the states match again');
+
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+test_checksum_state($standby, 'off');
+
+# The local state survives both clean and immediate restarts.
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+# Converge the cluster: enable offline on the standby too.
+$standby->stop;
+command_checks_all(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ 0,
+ [qr/appears to be a standby/],
+ [],
+ 'standby-role notice on offline enable');
+$standby->start;
+test_checksum_state($standby, 'on');
+$primary->wait_for_catchup($standby);
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after converging');
+
+# Scenario 2: disable offline on the standby only. The replayed
+# checkpoint records of the still-enabled primary must not override it.
+$standby->stop;
+$standby->checksum_disable_offline;
+$standby->start;
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'off');
+
+# Restartpoints must persist the local state, not the replayed copy.
+$standby->safe_psql('postgres', "CHECKPOINT;");
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable with checksums disabled locally');
+
+# Scenario 3: crash-restart right after an online transition, before the
+# next restartpoint. Replay then resumes from an older restartpoint whose
+# checkpoint records still carry the pre-transition state. Those must
+# still match the state seeded from the control file, and the transition
+# itself must be re-established by re-replaying the XLOG2_CHECKSUMS record,
+# without a spurious mismatch warning along the way.
+
+# Converge first: bring the standby back to "on" offline.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+disable_data_checksums($primary, wait => 'off');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+# Crash-restart the standby immediately, before any restartpoint has had a
+# chance to persist the new state to its control file.
+$logstart = -s $standby->logfile;
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001',
+ 'standby readable after crash-restart across an online transition');
+
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no spurious mismatch warning after crash-restart across an online transition'
+);
+
+# Scenario 4: a standby stopped while replaying an interrupted online
+# transition keeps the interrupted state in its own control file, and
+# pg_checksums must refuse to touch it. A primary is never caught this
+# way, as its checksums launcher resolves inprogress-on back to off from
+# its exit cleanup. A standby has no launcher; it carries forward
+# whatever the last replayed record left it in.
+
+# Block an online enable on the primary at inprogress-on with a
+# blocking temp table, same trick as in 004_offline.pl.
+my $bsession = $primary->background_psql('postgres');
+$bsession->query_safe('CREATE TEMPORARY TABLE tt (a integer);');
+enable_data_checksums($primary, wait => 'inprogress-on');
+
+# The standby picks up the in-progress state from the XLOG2_CHECKSUMS record.
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# Stop the standby cleanly; its restartpoint persists inprogress-on to
+# its own control file, since nothing on a standby resolves it away.
+$standby->stop;
+
+command_fails_like(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --enable refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --check refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--disable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --disable refuses a standby stopped mid-transition');
+
+$standby->start;
+$bsession->quit;
+wait_for_checksum_state($primary, 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable once the transition completes');
+
+$standby->stop;
+$primary->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/013_rewind.pl b/src/test/modules/test_checksums/t/013_rewind.pl
new file mode 100644
index 00000000000..42fee544f62
--- /dev/null
+++ b/src/test/modules/test_checksums/t/013_rewind.pl
@@ -0,0 +1,198 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test pg_rewind across an online data checksum enable.
+#
+# A clean switchover leaves the shutdown checkpoint of the old primary
+# as the last common checkpoint between the two nodes. When data
+# checksums are enabled online on the new primary before the old one is
+# rewound, pg_rewind installs the control file of the new primary, which
+# already claims checksums are fully enabled, while replay begins at the
+# shutdown checkpoint whose record still carries the old state. Replay
+# of the WAL stretch from before the enable must run with checksums off,
+# as recorded in the checkpoint, else it would verify pages which never
+# had checksums written and fail recovery.
+#
+# The new primary requests a checkpoint right after promotion, and
+# replaying its CHECKPOINT_REDO record would repair the state before
+# any interesting WAL is reached. Hold the checkpointer on the standby
+# in a restartpoint over the promotion, like in the recovery test
+# 041_checkpoint_at_promote, so that the post-promotion writes end up
+# in WAL before the first checkpoint of the new timeline.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Old primary. full_page_writes is off so that the updates done on the
+# promoted node do not carry full page images, forcing replay on the
+# rewound node to read the pages from disk. wal_log_hints is required
+# by pg_rewind on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+full_page_writes = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+log_checkpoints = on
+]);
+$node_a->start;
+
+if (!$node_a->check_extension('injection_points'))
+{
+ plan skip_all => 'Extension injection_points not installed';
+}
+
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+$node_a->safe_psql('postgres', "CREATE EXTENSION injection_points;");
+
+# Set the hint bits on t before taking the backup, so that reads on the
+# promoted node do not emit full page images for its pages later.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->safe_psql('postgres', "SELECT count(*) FROM t;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_a, 'off');
+test_checksum_state($node_b, 'off');
+
+# Hold the next restartpoint on the standby.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point', 'wait');");
+
+# Give the restartpoint a checkpoint record to work on, then start it
+# in a background session; it will block on the injection point with
+# the checkpointer busy until released.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+my $bg_psql = $node_b->background_psql('postgres', on_error_stop => 0);
+$bg_psql->query_until(
+ qr/starting_restartpoint/, q(
+ \echo starting_restartpoint
+ CHECKPOINT;
+));
+$node_b->wait_for_event('checkpointer', 'create-restart-point');
+
+# Clean switchover: the shutdown checkpoint of A streams to B and
+# becomes the last common checkpoint.
+$node_a->stop('fast');
+
+my ($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+$stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+my $shutdown_ckpt = $1;
+$node_b->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_ckpt'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+
+my $logstart = -s $node_b->logfile;
+$node_b->promote;
+
+# Accidental restart of the old primary, diverging its timeline.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->stop('fast');
+
+# Updates on the new primary before its first checkpoint; replay of
+# these on the rewound node has to read the pages from disk.
+$node_b->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 25 = 0;");
+ok( !$node_b->log_contains("checkpoint complete", $logstart),
+ "no checkpoint on the new timeline before the updates");
+
+# Release the checkpointer; the queued post-promotion checkpoint runs
+# after the updates.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+$bg_psql->quit;
+
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+# pg_rewind refuses to run with full_page_writes disabled on the
+# source; the updates it was disabled for are already in WAL.
+$node_b->safe_psql('postgres', "ALTER SYSTEM SET full_page_writes = on;");
+$node_b->safe_psql('postgres', "SELECT pg_reload_conf();");
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the new primary');
+
+# Replay on the rewound node must start at the shutdown checkpoint of
+# the switchover, with the control file of the new primary.
+my $backup_label = slurp_file($node_a->data_dir . '/backup_label');
+$backup_label =~ /^CHECKPOINT LOCATION: ([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from backup_label";
+is($1, $shutdown_ckpt, 'replay starts at the switchover checkpoint');
+
+($stdout, $stderr) = run_command(
+ [
+ 'pg_waldump',
+ '-p' => $node_a->data_dir . '/pg_wal',
+ '-t' => 1,
+ '-s' => $shutdown_ckpt,
+ '-n' => 1,
+ ]);
+like($stdout, qr/CHECKPOINT_SHUTDOWN/,
+ 'last common checkpoint is a shutdown checkpoint');
+
+($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+like(
+ $stdout,
+ qr/^Data page checksum version:\s*1$/m,
+ 'rewound node received a control file with checksums on');
+
+# Start the rewound node as a standby of the new primary. Replay runs
+# through the pre-enable WAL stretch and the online enable.
+#
+# The rewind replaced the configuration files with those of the new
+# primary, so put the port back.
+my $connstr = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_b->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/014_lockstep.pl b/src/test/modules/test_checksums/t/014_lockstep.pl
new file mode 100644
index 00000000000..89c3b9da2cc
--- /dev/null
+++ b/src/test/modules/test_checksums/t/014_lockstep.pl
@@ -0,0 +1,89 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# The lockstep procedure for offline checksum changes in a replication
+# setup: stop all nodes, run pg_checksums on all of them, restart.
+# Replay of checkpoint records written before the change must not
+# revert the state of the standby.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Stop the standby first: the WAL written after this point is replayed
+# only after the offline switch, and every checkpoint record in it
+# still carries the old state.
+$standby->stop;
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 0;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 1;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+# The lockstep procedure.
+$primary->checksum_enable_offline;
+$standby->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to off.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "off" in the replayed WAL/,
+ 0);
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep enable');
+
+# Crash the standby and replay the same stretch again.
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'on');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after crash restart');
+
+# Once a post-switch checkpoint has been replayed the states match and
+# no warning may be logged.
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no mismatch warning once the states match');
+
+# Every page the standby wrote in this window must carry a checksum.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$standby->stop;
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'checksums valid on the standby');
+
+$primary->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/015_backup_online_enable.pl b/src/test/modules/test_checksums/t/015_backup_online_enable.pl
new file mode 100644
index 00000000000..8ea6679dc63
--- /dev/null
+++ b/src/test/modules/test_checksums/t/015_backup_online_enable.pl
@@ -0,0 +1,71 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A base backup taken while an online checksum enable is running. The
+# control file copied with the backup, and the redo point of the backup
+# checkpoint, both carry the in-progress state; replay of the
+# XLOG2_CHECKSUMS records completes the transition on the new standby.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', 'autovacuum = off');
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+# Hold the enable at inprogress-on with a pre-existing temp table the
+# checksum worker has to wait out, same trick as in 004_offline.pl and
+# scenario 4 of 012_offline_standby.pl.
+my $bsession = $primary->background_psql('postgres');
+$bsession->query_safe('CREATE TEMPORARY TABLE tt (a integer);');
+
+enable_data_checksums($primary, wait => 'inprogress-on');
+
+# The backup's checkpoint is taken while the primary sits at
+# inprogress-on, so both the control file and the redo point's
+# XLOG_CHECKPOINT_REDO record carry that state.
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+
+# Backup label recovery adopts the redo point's state immediately, before
+# any further WAL is replayed.
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# Release the barrier; the transition completes on the primary and
+# replicates. Nothing before this point can legitimately disagree: the
+# standby starts out at the same inprogress-on state as the primary, so
+# there is nothing yet to warn about.
+$bsession->quit;
+wait_for_checksum_state($primary, 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after the transition completed');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no mismatch warning at any point during a backup taken mid-transition');
+
+# No extra CHECKPOINT needed: the transition's completion checkpoint wrote
+# every page with a checksum, and the shutdown restartpoint flushes the rest.
+$standby->stop;
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'checksums valid on the standby');
+
+$primary->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/016_backup_from_standby.pl b/src/test/modules/test_checksums/t/016_backup_from_standby.pl
new file mode 100644
index 00000000000..3d01259e424
--- /dev/null
+++ b/src/test/modules/test_checksums/t/016_backup_from_standby.pl
@@ -0,0 +1,100 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that a base backup taken from a standby keeps the standby's data
+# checksum state.
+#
+# A backup taken on a standby uses the last restartpoint as its starting
+# checkpoint (do_pg_backup_start()), so the record at the redo point was
+# written by the upstream primary and carries the primary's state. Recovery
+# from such a backup must not adopt that state: the files were copied from
+# the standby, and with the standby diverged to "off" under an "on" primary
+# the new node would come up verifying checksums its files do not have.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', 'autovacuum = off');
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Offline enable on the primary only. Per 012_offline_standby.pl this is a
+# divergence the standby is expected to survive: it keeps its own "off"
+# state, warns once, and stays readable.
+$standby->stop;
+$primary->stop;
+system_or_bail('pg_checksums', '--enable', '--pgdata', $primary->data_dir);
+$primary->start;
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'on');
+test_checksum_state($standby, 'off');
+
+# Make sure the standby has written pages under its own "off" state, so its
+# files really do lack checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,50000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', "SELECT count(*) FROM t2;");
+
+# Now take a base backup *from the standby* and bring the copy up.
+$standby->backup('from_standby');
+my $newnode = PostgreSQL::Test::Cluster->new('newnode');
+$newnode->init_from_backup($standby, 'from_standby', has_streaming => 1);
+$newnode->append_conf('postgresql.conf',
+ "primary_conninfo = '" . $primary->connstr . "'");
+$newnode->start;
+
+# The new node's files all came from a cluster running with checksums off.
+# Anything but "off" here means it adopted the primary's state through the
+# checkpoint record at the redo point.
+my ($rc, $stdout, $stderr) = $newnode->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'the node copied from the standby accepts connections')
+ or diag("stderr: $stderr");
+is($stdout, 'off', 'backup of an "off" standby comes up with checksums off');
+
+# And it must be able to read the pages the standby wrote without checksums.
+($rc, $stdout, $stderr) =
+ $newnode->psql('postgres', "SELECT count(*) FROM t2;");
+is($rc, 0, 'pages copied from the standby are readable')
+ or diag("stderr: $stderr");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($newnode->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the node copied from the standby');
+
+$newnode->stop('immediate');
+
+($stdout, $stderr) = run_command([ 'pg_controldata', $newnode->data_dir ]);
+my ($ctl_state) = $stdout =~ /Data page checksum version:\s+(\d+)/;
+note("newnode control file data checksum version: $ctl_state");
+
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/017_standby_crash_after_disable.pl b/src/test/modules/test_checksums/t/017_standby_crash_after_disable.pl
new file mode 100644
index 00000000000..315b27b93d4
--- /dev/null
+++ b/src/test/modules/test_checksums/t/017_standby_crash_after_disable.pl
@@ -0,0 +1,125 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that a standby crashing after a replayed online disable, but before
+# its next restartpoint, restarts cleanly.
+#
+# Pages dirtied before the XLOG2_CHECKSUMS record and evicted after it are
+# written out under the new "off" state, without checksums. The control
+# file must follow the record immediately in this direction: a crash-restart
+# initializes checksum verification from the control file, and one still
+# saying "on" would fail verification on exactly those pages while replaying
+# records older than the state change.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$primary->start;
+
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,1000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+
+# Small buffer pool so that replaying a bulk load evicts the pages we care
+# about, and no restartpoints so the control file stays where it is.
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 1MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_delay = 10000
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'on');
+
+# No full page images, so that replay has to read the pages of "t" back from
+# disk instead of overwriting them from the WAL.
+$primary->append_conf('postgresql.conf', 'full_page_writes = off');
+$primary->reload;
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" on the standby, *before* the state change. They are
+# not flushed: the standby has no restartpoint from here on.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+# Online disable. The standby replays it and switches to "off", but its
+# control file is not updated.
+$primary->safe_psql('postgres', 'SELECT pg_disable_data_checksums();');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+my ($ctl_before) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($ctl_state) = $ctl_before =~ /Data page checksum version:\s+(\d+)/;
+note(
+ "standby control file data_checksum_version after the disable: "
+ . "$ctl_state (0 = off, 1 = on)");
+
+# Force the standby to evict the dirty pages of "t" now that it is "off":
+# they are written back without checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE filler AS SELECT generate_series(1,300000) AS a;");
+$primary->wait_for_catchup($standby);
+
+# Crash the standby before it gets a chance to run a restartpoint.
+$standby->stop('immediate');
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'standby restarts after crashing between a checksum state '
+ . 'change and the next restartpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'the pages written while "off" are readable')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("standby log tail:\n" . join("\n", @lines[ -12 .. -1 ]))
+ if @lines >= 12;
+ fail('the pages written while "off" are readable');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/018_promote_enable_crash.pl b/src/test/modules/test_checksums/t/018_promote_enable_crash.pl
new file mode 100644
index 00000000000..97f84da5685
--- /dev/null
+++ b/src/test/modules/test_checksums/t/018_promote_enable_crash.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a standby promoted right after replaying an online enable, before any
+# restartpoint has flushed the rewritten pages.
+#
+# The end-of-recovery record persists the checksum state without flushing the
+# buffer pool, while the control file still points at a restartpoint older than
+# the transition. Persisting "on" there would make a crash before the
+# post-promotion checkpoint verify checksums over the pages that were flushed
+# while the state was still "off"; promotion must take the full checkpoint
+# instead.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Dirty the pages of "t" without full page images, then push them out to disk
+# on the standby while checksums are still off.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Online enable on the primary; the standby replays it into its own buffers,
+# where the rewritten pages stay dirty (no restartpoint, no bgwriter).
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+note("standby control file before the promotion: $before");
+
+# Promote. The end-of-recovery record persists the live state without any
+# buffer flush; the checkpoint requested afterwards is not immediate.
+$standby->promote;
+$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery();')
+ or die 'timed out waiting for the promotion';
+$standby->stop('immediate');
+
+($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($ckpt) = $ctl =~ /Latest checkpoint location:\s+(\S+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note(
+ "standby control file after the promotion: $after, checkpoint $ckpt, redo $redo"
+);
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'promoted node restarts after crashing before its first checkpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/019_restartpoint_race.pl b/src/test/modules/test_checksums/t/019_restartpoint_race.pl
new file mode 100644
index 00000000000..65aa834f366
--- /dev/null
+++ b/src/test/modules/test_checksums/t/019_restartpoint_race.pl
@@ -0,0 +1,138 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a restartpoint whose flush races the replay of an online enable.
+#
+# Replay keeps running while CheckPointGuts() writes out the buffer pool, so
+# the state at the end of the flush can be newer than the one the flush ran
+# under. Persisting it would claim checksums for pages the flush wrote out
+# before the transition, against a redo pointer that predates it.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# C1: the checkpoint the pending restartpoint will target.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" without full page images and push them to disk on
+# the standby while it is still "off".
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Start a restartpoint for C1 and hold it right after CheckPointGuts().
+$standby->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point','wait');");
+my $bg = $standby->background_psql('postgres');
+$bg->query_until(qr//, "\\echo restartpoint\nCHECKPOINT;\n");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint injection point';
+
+# The startup process keeps replaying while the checkpointer is held: the
+# whole online enable lands, and the rewritten pages stay dirty.
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+# Let the restartpoint finish; it persists the state it samples now.
+$standby->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint to finish';
+$standby->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+
+$standby->stop('immediate');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note("standby control file after the restartpoint: $after, redo $redo");
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started, 'standby restarts after a restartpoint that raced the enable');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/020_primary_enable_crash.pl b/src/test/modules/test_checksums/t/020_primary_enable_crash.pl
new file mode 100644
index 00000000000..196914bc40c
--- /dev/null
+++ b/src/test/modules/test_checksums/t/020_primary_enable_crash.pl
@@ -0,0 +1,118 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a primary crashing inside SetDataChecksumsOn(), just before the forced
+# checkpoint that flushes the rewritten pages, with crash recovery resuming
+# from a checkpoint older than the transition.
+#
+# The control file must still say "inprogress-on" there. Replay does not
+# adopt the state of the checkpoint record it resumes from, so an "on" written
+# before the flush would stay in effect while replay reads pages whose rewrite
+# never reached disk. Here the pages went out through the rewriting worker's
+# ring buffer, so only the control file state discriminates; see
+# 022_resident_enable_crash.pl for the case where they did not.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('primary_enable_crash');
+$node->init(no_data_checksums => 1);
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+shared_buffers = 512MB
+bgwriter_lru_maxpages = 0
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$node->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+# Establish the checkpoint that crash recovery will resume from.
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+# Dirty the pages of "t" without full page images, then push them out to disk
+# while checksums are still off.
+$node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+my $relpath =
+ $node->safe_psql('postgres', "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $node->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on t: $evicted, relpath $relpath");
+
+# Stop the enable right after the control file has been updated to "on" but
+# before the checkpoint that flushes the rewritten pages.
+$node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+$node->safe_psql('postgres', 'SELECT pg_enable_data_checksums();');
+
+$node->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'datachecksums-on-before-checkpoint';"
+) or die 'timed out waiting for the injection point';
+
+my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+note("control file data_checksum_version before the crash: $ctl_state");
+is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+$node->stop('immediate');
+
+# Show the on-disk checksum field of the first page of "t".
+my $page;
+open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $node->start(fail_ok => 1);
+ok($started, 'primary restarts after crashing inside the online enable');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $node->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/021_standby_shutdown_catchup.pl b/src/test/modules/test_checksums/t/021_standby_shutdown_catchup.pl
new file mode 100644
index 00000000000..699dae4542a
--- /dev/null
+++ b/src/test/modules/test_checksums/t/021_standby_shutdown_catchup.pl
@@ -0,0 +1,93 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a standby stopped after replaying an online enable, but before the
+# checkpoint record that follows it on the primary.
+#
+# The shutdown restartpoint has no new checkpoint record to work from and is
+# skipped, so the control file would keep the in-progress state the transition
+# passed through, even though replay left the node at "on" and rewrote every
+# page. pg_checksums reads that field and refuses to run on an in-progress
+# state, so the shutdown has to flush and catch it up instead.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Anchor the standby's last restartpoint here, so that nothing replayed from
+# now on gives the shutdown restartpoint a newer checkpoint record to use.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Hold the enable right after the state change record, before the checkpoint
+# it requests once the transition is complete.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+enable_data_checksums($primary);
+$primary->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'datachecksums-on-before-checkpoint';"
+) or die 'timed out waiting for the injection point';
+wait_for_checksum_state($primary, 'on');
+
+# Flush the state change record out to the standby without writing a
+# checkpoint record of any kind.
+$primary->safe_psql('postgres', 'CREATE TABLE flush_marker (a int);');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+$standby->stop;
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($state, '1',
+ 'shutdown catches the control file up with the replayed state');
+
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'pg_checksums verifies the stopped standby');
+
+$primary->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-checkpoint');");
+$primary->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-checkpoint');");
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/022_resident_enable_crash.pl b/src/test/modules/test_checksums/t/022_resident_enable_crash.pl
new file mode 100644
index 00000000000..95f230ba39e
--- /dev/null
+++ b/src/test/modules/test_checksums/t/022_resident_enable_crash.pl
@@ -0,0 +1,138 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a primary crashing inside SetDataChecksumsOn(), just before the forced
+# checkpoint that flushes the rewritten pages, when those pages are resident
+# in shared buffers.
+#
+# The rewriting worker reads through a BAS_VACUUM ring, which writes the pages
+# back as the ring recycles, but a page already resident in shared buffers is
+# not read through the ring: ReadBufferExtended() hands back the existing
+# buffer, and it stays dirty until a checkpoint. The control file may
+# therefore not say "on" before that checkpoint has run, or crash recovery
+# would resume from a checkpoint older than the transition with verification
+# already enabled, and replay records that read those still-unchecksummed
+# pages. full_page_writes is off so that the records do not simply overwrite
+# the pages with a full page image.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('resident_enable_crash');
+$node->init(no_data_checksums => 1);
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+wal_level = replica
+full_page_writes = off
+wal_log_hints = off
+shared_buffers = 512MB
+bgwriter_lru_maxpages = 0
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$node->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,100000) AS a;');
+my $relpath =
+ $node->safe_psql('postgres', "SELECT pg_relation_filepath('t'::regclass);");
+
+# Establish the checkpoint that crash recovery will resume from, with the
+# pages of "t" written out while checksums are still off.
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+# Dirty those pages again without emitting full page images, and leave them in
+# shared buffers. Replay of these records has to read the pages from disk.
+$node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+
+my $dirty_before = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+cmp_ok($dirty_before, '>', 0,
+ 'pages of t are resident and dirty before enabling checksums');
+
+# Hold the enabling right before the checkpoint that flushes the rewritten
+# pages.
+$node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+
+enable_data_checksums($node);
+$node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+note("control file data_checksum_version before the crash: $ctl_state");
+is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+# The rewritten pages must still be sitting dirty in shared buffers, or the
+# window this test is about does not exist.
+my $dirty_after = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+cmp_ok($dirty_after, '>', 0,
+ 'rewritten pages of t are still dirty before the checkpoint');
+
+$node->stop('immediate');
+
+# The on-disk copy is the one written before the transition, without a
+# checksum.
+my $page;
+open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+is($pd_checksum, 0, 'block 0 of t on disk carries no checksum');
+
+my $started = $node->start(fail_ok => 1);
+ok($started, 'primary restarts after crashing inside the online enable');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $node->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/023_concurrent_checkpoint_enable.pl b/src/test/modules/test_checksums/t/023_concurrent_checkpoint_enable.pl
new file mode 100644
index 00000000000..2d1b78fc5cd
--- /dev/null
+++ b/src/test/modules/test_checksums/t/023_concurrent_checkpoint_enable.pl
@@ -0,0 +1,114 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A checkpoint that runs between the XLOG2_CHECKSUMS("on") record and the
+# control file write at the end of SetDataChecksumsOn() must not make a crash
+# throw the completed transition away.
+#
+# SetDataChecksumsOn() writes the record, flips shared memory to "on", emits
+# the barrier and only then requests the checkpoint that flushes the rewritten
+# pages; the control file is written after that checkpoint returns. A crash in
+# that window is harmless only as long as recovery still starts before the
+# record. Any checkpoint completing in the window moves the redo point past
+# the record, so recovery would never see it, would come up with the control
+# file's "inprogress-on" and StartupXLOG() would demote that to "off", even
+# though every page on disk carries a checksum by then.
+#
+# CreateCheckPoint() therefore persists the state the checkpoint ran under, the
+# same way CreateRestartPoint() does. The window is naturally reachable:
+# checkpoint_timeout, max_wal_size, an explicit CHECKPOINT, pg_basebackup or
+# pg_backup_start can all fire there. Here it is made deterministic by holding
+# the launcher at the datachecksums-on-before-checkpoint injection point and
+# checkpointing from another session.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('concurrent_checkpoint');
+$node->init(no_data_checksums => 1);
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+my $relpath =
+ $node->safe_psql('postgres', "SELECT pg_relation_filepath('t'::regclass);");
+
+# Hold the launcher after the record, the shared memory flip and the barrier,
+# but before the checkpoint SetDataChecksumsOn() requests itself.
+$node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+
+enable_data_checksums($node);
+$node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+# Every backend already sees "on" and writes checksums.
+test_checksum_state($node, 'on');
+
+# ... while the control file still says "inprogress-on".
+my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($before, '3', 'control file says "inprogress-on" inside the window');
+
+# A concurrent checkpoint. It flushes the rewritten pages and moves the redo
+# point past the XLOG2_CHECKSUMS("on") record, so it has to record the "on"
+# state in the control file as well.
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($after, '1',
+ 'the concurrent checkpoint records "on" in the control file');
+
+$node->stop('immediate');
+
+# The checkpoint flushed the rewritten pages, so they carry a checksum on
+# disk: the transition really did complete.
+my $page;
+open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+isnt($pd_checksum, 0, 'block 0 of t on disk carries a checksum');
+
+my $log_offset = -s $node->logfile;
+$node->start;
+
+# The transition is complete on disk, so the cluster has to come back "on".
+test_checksum_state($node, 'on');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+$node->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/024_enable_crash_after_checkpoint.pl b/src/test/modules/test_checksums/t/024_enable_crash_after_checkpoint.pl
new file mode 100644
index 00000000000..e75f17394cb
--- /dev/null
+++ b/src/test/modules/test_checksums/t/024_enable_crash_after_checkpoint.pl
@@ -0,0 +1,101 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A crash between the checkpoint an online enable requests and the control
+# file write that follows it must not lose the transition.
+#
+# The last steps of SetDataChecksumsOn() are
+#
+# WAL record -> shmem -> barrier -> checkpoint -> persist
+#
+# The checkpoint is what makes "on" safe to persist, but it also moves the
+# redo point above the XLOG2_CHECKSUMS record that carries the new state.
+# Crash recovery started from that checkpoint therefore never replays the
+# record, so without the state the checkpoint itself persists, a crash in the
+# remaining window would bring the cluster back at "inprogress-on", which
+# StartupXLOG() resolves to "off", discarding a transition whose pages are all
+# on disk with a checksum.
+#
+# t/023 exercises the same window through a checkpoint requested by another
+# session. This test closes it from the other side: the transition's own
+# checkpoint is the one that moves the redo point, so persisting the state
+# from CreateCheckPoint() is what has to save it, not the write below the
+# injection point.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('crash_after_checkpoint');
+$node->init(no_data_checksums => 1);
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+# Hold the launcher after the checkpoint that licenses "on" has completed and
+# before the state reaches the control file.
+$node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-after-checkpoint','wait');"
+);
+
+enable_data_checksums($node);
+$node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-after-checkpoint');
+
+# The transition is complete as far as the running cluster is concerned.
+test_checksum_state($node, 'on');
+
+# The checkpoint has already recorded it, so the pending write below the
+# injection point has nothing left to do.
+my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+my ($version) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($version, '1',
+ 'the requested checkpoint recorded "on" in the control file');
+
+# Crash before the control file write that follows the checkpoint.
+$node->stop('immediate');
+
+$node->start;
+
+# Every page on disk carries a checksum and the checkpoint that flushed them
+# completed, so the cluster has to come back verifying them.
+test_checksum_state($node, 'on');
+
+is($node->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '10000', 'relation readable after the crash');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+# The data directory must be one the offline tools accept.
+$node->stop;
+$node->command_ok([ 'pg_checksums', '--check', '-D', $node->data_dir ],
+ 'pg_checksums accepts the data directory');
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/025_cascade_divergence.pl b/src/test/modules/test_checksums/t/025_cascade_divergence.pl
new file mode 100644
index 00000000000..b4b474fa583
--- /dev/null
+++ b/src/test/modules/test_checksums/t/025_cascade_divergence.pl
@@ -0,0 +1,115 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# An offline data checksum state change made on one node of a cascading setup
+# is reported all the way down the chain.
+#
+# CheckReplayedDataChecksumState() is reached from xlog_redo() when a
+# checkpoint record (XLOG_CHECKPOINT_SHUTDOWN or XLOG_CHECKPOINT_REDO) is
+# replayed after consistency has been reached. Those records are written by
+# the root primary only and are relayed verbatim by every intermediate
+# standby, so a cascaded standby that was changed offline learns about the
+# divergence too. An intermediate standby's own restartpoints are not
+# WAL-logged and cannot surface it.
+#
+# Note that this makes detection checkpoint-driven: an idle or read-only
+# primary surfaces the divergence no sooner than its next checkpoint, so up to
+# checkpoint_timeout may pass before the operator sees anything. Closing that
+# window would require comparing the state when a walreceiver connects.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# checkpoint_timeout is deliberately long so that the only checkpoint record in
+# this test is the one requested explicitly below.
+my $conf = qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+);
+
+my $primary = PostgreSQL::Test::Cluster->new('cascade_primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', $conf);
+$primary->start;
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,1000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('cascade_standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf('postgresql.conf', $conf);
+$standby->start;
+
+$standby->backup('backup');
+my $cascade = PostgreSQL::Test::Cluster->new('cascade_cascade');
+$cascade->init_from_backup($standby, 'backup', has_streaming => 1);
+$cascade->append_conf('postgresql.conf', $conf);
+$cascade->start;
+
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+test_checksum_state($cascade, 'off');
+
+# Change the cascaded standby offline, and only it.
+$cascade->stop;
+$cascade->command_ok([ 'pg_checksums', '--enable', '-D', $cascade->data_dir ],
+ 'pg_checksums enables checksums on the cascaded standby only');
+
+my $logstart = -s $cascade->logfile;
+$cascade->start;
+
+test_checksum_state($cascade, 'on');
+
+# Ordinary WAL traffic carries no state, so the cascaded standby keeps
+# streaming from a chain whose state is "off" without noticing.
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (1);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+is($cascade->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '1001', 'cascaded standby keeps streaming from a divergent chain');
+
+# Restartpoints on the intermediate standby are not WAL-logged, so this does
+# not surface anything on the cascaded standby either.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (2);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+my $log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+unlike(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'an intermediate restartpoint does not surface the divergence');
+
+# A checkpoint on the root primary does: the record is relayed down the whole
+# chain, so both the direct and the cascaded standby compare it against their
+# own state. wait_for_catchup() below waits for replay, not just receipt, so
+# the record has been through xlog_redo() on both nodes once it returns.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+$log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+like(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'the cascaded standby reports the divergence on the primary checkpoint');
+
+$cascade->stop;
+$standby->stop;
+$primary->stop;
+
+done_testing();
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v4-0002-pg_checksums-Refuse-interrupted-transitions-note-.patch (5.9K, ../../188A1307-1A92-45CA-9DBE-FB0962D3756E@yesql.se/3-v4-0002-pg_checksums-Refuse-interrupted-transitions-note-.patch)
download | inline diff:
From 9395c0584dbe0c6b800ff33198f90adffe547b0e Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:19 +0000
Subject: [PATCH v4 2/4] pg_checksums: Refuse interrupted transitions, note
change is local
An inprogress-on or inprogress-off control file means an online
transition was cut short; running the offline tool on top of it mixes
two procedures. Refuse it in every mode, with a hint pointing at the
start-stop cycle (or, on a standby, at letting replication finish)
that resets the state. Also print that the change applies to one
data directory only, with standby-aware wording, since in a
replication setup the same change must be made on every node.
On a primary, inprogress-on cannot survive a graceful stop: the
checksums launcher resolves it from its exit cleanup, and a crashed
primary is already rejected by the existing "cluster must be shut
down" check. inprogress-off can, however: it is set by the backend
running pg_disable_data_checksums(), and a fast shutdown arriving
between its two barriers leaves it behind in a cleanly shut down
control file, where the next start-stop cycle resolves it at end of
recovery. A standby can be stopped with either state, since it has
no launcher and only carries forward whatever state the last replayed
WAL record left it in, with a restartpoint persisting that as-is.
The standby is also the deterministic way to reach the new guard,
which is why the test coverage uses one; the primary window would
need an injection point between the two barriers.
The pg_checksums page said the tool still processes all relation files
regardless of an interrupted online transition; describe the refusal
instead.
---
doc/src/sgml/ref/pg_checksums.sgml | 11 +++++-----
src/bin/pg_checksums/pg_checksums.c | 21 +++++++++++++++++++
.../test_checksums/t/012_offline_standby.pl | 8 +++----
3 files changed, 30 insertions(+), 10 deletions(-)
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index bf07da09754..000940e7cb7 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -49,11 +49,12 @@ PostgreSQL documentation
</para>
<para>
- When enabling checksums with <application>pg_checksums</application>, if
- checksums were in the process of being enabled using
- <xref linkend="checksums-online-enable-disable"/> when the cluster was shut
- down, <application>pg_checksums</application> will still process all
- relation files regardless of the progress of online checksum processing.
+ If checksums were in the process of being enabled or disabled using
+ <xref linkend="checksums-online-enable-disable"/> when the cluster was
+ shut down, the control file still records that interrupted state, and
+ <application>pg_checksums</application> refuses to run in any mode.
+ Start the cluster and shut it down cleanly to reset the state, then
+ retry; on a standby, let replication complete the transition first.
</para>
<para>
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 3b3ae23f1a6..e9fccbe9f71 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -586,6 +586,21 @@ main(int argc, char *argv[])
ControlFile->state != DB_SHUTDOWNED_IN_RECOVERY)
pg_fatal("cluster must be shut down");
+ /*
+ * An inprogress state means an online transition was cut short. A
+ * standby stopped mid-transition carries either state; a cleanly shut
+ * down primary can still carry inprogress-off, which a fast shutdown
+ * during pg_disable_data_checksums() leaves behind, while inprogress-on
+ * is always resolved by the launcher's exit cleanup.
+ */
+ if (ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted");
+ pg_log_error_hint("Start and cleanly shut down the cluster once to reset the data checksum state, then retry. On a standby, let replication complete the transition first.");
+ exit(1);
+ }
+
if (ControlFile->data_checksum_version != PG_DATA_CHECKSUM_VERSION &&
mode == PG_MODE_CHECK)
pg_fatal("data checksums are not enabled in cluster");
@@ -663,6 +678,12 @@ main(int argc, char *argv[])
printf(_("Checksums enabled in cluster\n"));
else
printf(_("Checksums disabled in cluster\n"));
+
+ printf(_("This change applies to this data directory only.\n"));
+ if (ControlFile->state == DB_SHUTDOWNED_IN_RECOVERY)
+ printf(_("This node appears to be a standby; apply the same change to the primary and all other standbys.\n"));
+ else
+ printf(_("In a replication setup, apply the same change to every node while all are stopped, before restarting any of them.\n"));
}
return 0;
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
index 94c476293b7..33ee6802168 100644
--- a/src/test/modules/test_checksums/t/012_offline_standby.pl
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -96,10 +96,8 @@ test_checksum_state($standby, 'off');
$standby->stop;
command_checks_all(
[ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
- 0,
- [qr/appears to be a standby/],
- [],
- 'standby-role notice on offline enable');
+ 0, [qr/appears to be a standby/],
+ [], 'standby-role notice on offline enable');
$standby->start;
test_checksum_state($standby, 'on');
$primary->wait_for_catchup($standby);
@@ -205,7 +203,7 @@ wait_for_checksum_state($primary, 'on');
$primary->wait_for_catchup($standby);
wait_for_checksum_state($standby, 'on');
-is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
'10001', 'standby readable once the transition completes');
$standby->stop;
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v4-0003-pg_rewind-Check-the-data-checksum-states-of-sourc.patch (21.6K, ../../188A1307-1A92-45CA-9DBE-FB0962D3756E@yesql.se/4-v4-0003-pg_rewind-Check-the-data-checksum-states-of-sourc.patch)
download | inline diff:
From a25bb13a9911ebd9d5496c5ca311fa6d3d9d5176 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 18:40:22 +0000
Subject: [PATCH v4 3/4] pg_rewind: Check the data checksum states of source
and target
pg_rewind installs the source's control file on the target, but every
block it does not copy keeps the target's content. With checksums
enabled on the target and disabled on the source, the rewound server
claims enabled checksums while the blocks copied from the source have
none, and fails checksum verification as soon as it reads them; in the
worst case every connection attempt dies on an unverifiable catalog
page.
Comparing the control files is not enough. Replay on the rewound
server resumes from the last common checkpoint and adopts the data
checksum state recorded there, so an offline disable on the source
after the divergence leaves both control files saying "off" while the
rewound server still resumes with verification enabled. Check the
state carried by the divergence checkpoint as well, which pg_rewind
already reads.
Refuse both cases, and an interrupted online transition on either
side, same as pg_checksums does. The opposite mismatch, checksums on
the source only, is allowed with a warning: recovery under the backup
label adopts the divergence state, so the rewound server keeps
checksums disabled (or converges through the replayed WAL if the
source enabled them online), and no page can fail verification.
---
doc/src/sgml/ref/pg_rewind.sgml | 14 ++
src/bin/pg_rewind/parsexlog.c | 4 +-
src/bin/pg_rewind/pg_rewind.c | 93 +++++++++++-
src/bin/pg_rewind/pg_rewind.h | 1 +
src/test/modules/test_checksums/meson.build | 2 +
.../test_checksums/t/026_rewind_state.pl | 133 +++++++++++++++++
.../t/027_rewind_standby_target.pl | 141 ++++++++++++++++++
7 files changed, 386 insertions(+), 2 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/026_rewind_state.pl
create mode 100644 src/test/modules/test_checksums/t/027_rewind_standby_target.pl
diff --git a/doc/src/sgml/ref/pg_rewind.sgml b/doc/src/sgml/ref/pg_rewind.sgml
index b95e3868e9d..ac3d0c9328f 100644
--- a/doc/src/sgml/ref/pg_rewind.sgml
+++ b/doc/src/sgml/ref/pg_rewind.sgml
@@ -124,6 +124,20 @@ PostgreSQL documentation
<literal>on</literal>, but is enabled by default.
</para>
+ <para>
+ The data checksum states of the source and target server must be
+ compatible. <application>pg_rewind</application> refuses to run if an
+ online checksum state transition was interrupted on either server, or
+ if data checksums are enabled on the target server, or were enabled at
+ the point of divergence, while the source server runs without them: the
+ rewound server would fail checksum verification on the blocks copied
+ from the source. The opposite combination is allowed with a warning;
+ the rewound server keeps data checksums disabled unless the source
+ server enabled them online after the point of divergence. See
+ <xref linkend="app-pgchecksums"/> for how to change the state
+ consistently across a replication setup.
+ </para>
+
<warning>
<title>Warning: Failures While Rewinding</title>
<para>
diff --git a/src/bin/pg_rewind/parsexlog.c b/src/bin/pg_rewind/parsexlog.c
index 023e23b063c..6e87b00f8c2 100644
--- a/src/bin/pg_rewind/parsexlog.c
+++ b/src/bin/pg_rewind/parsexlog.c
@@ -167,7 +167,8 @@ readOneRecord(const char *datadir, XLogRecPtr ptr, int tliIndex,
void
findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
- XLogRecPtr *lastchkptredo, const char *restoreCommand)
+ XLogRecPtr *lastchkptredo, uint32 *lastchkptdatachecksums,
+ const char *restoreCommand)
{
/* Walk backwards, starting from the given record */
XLogRecord *record;
@@ -255,6 +256,7 @@ findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
*lastchkptrec = searchptr;
*lastchkpttli = checkPoint.ThisTimeLineID;
*lastchkptredo = checkPoint.redo;
+ *lastchkptdatachecksums = checkPoint.dataChecksumState;
break;
}
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 2e86fd158d0..6c0f11e01ba 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -144,6 +144,7 @@ main(int argc, char **argv)
XLogRecPtr chkptrec;
TimeLineID chkpttli;
XLogRecPtr chkptredo;
+ uint32 chkptdatachecksums;
TimeLineID source_tli;
TimeLineID target_tli;
XLogRecPtr target_wal_endrec;
@@ -471,10 +472,51 @@ main(int argc, char **argv)
keepwal_init();
findLastCheckpoint(datadir_target, divergerec, lastcommontliIndex,
- &chkptrec, &chkpttli, &chkptredo, restore_command);
+ &chkptrec, &chkpttli, &chkptredo, &chkptdatachecksums,
+ restore_command);
pg_log_info("rewinding from last common checkpoint at %X/%08X on timeline %u",
LSN_FORMAT_ARGS(chkptrec), chkpttli);
+ /*
+ * Replay on the rewound server resumes from the last common checkpoint
+ * and adopts the data checksum state recorded there, not the state in the
+ * control file installed from the source. If checksums were enabled at
+ * the divergence point but the source runs without them, the rewound
+ * server would verify checksums while the blocks copied from the source
+ * have none. The control files cannot reveal this: an offline disable on
+ * the source after the divergence leaves both of them saying "off".
+ *
+ * Only the fully enabled state needs checking. An in-progress state at
+ * the divergence point means a transition was still running there, and
+ * all of its page rewrites are logged after that point, so replay brings
+ * the rewound server to whatever state the copied WAL ends in.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the source server or recreate the target server from a base backup.");
+ exit(1);
+ }
+
+ /*
+ * The same comparison is needed against the target: every block the
+ * rewind does not copy keeps the target's content. The control files
+ * cannot reveal this case either, in the other direction: a standby's
+ * checkpoints are written by its upstream primary, so a standby whose
+ * checksums were disabled offline still has "on" checkpoints in its WAL,
+ * while both control files may agree.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_target.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the target server");
+ pg_log_error_detail("Blocks kept from the target would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the target server with pg_checksums or recreate the target server from a base backup.");
+ exit(1);
+ }
+
/* Initialize the hash table to track the status of each file */
filehash_init();
@@ -770,6 +812,55 @@ sanityChecks(void)
pg_fatal("target server needs to use either data checksums or \"wal_log_hints = on\"");
}
+ /*
+ * The rewound target keeps its control file fields from the source, but
+ * every block the rewind does not copy keeps the target's content. If
+ * checksums are enabled on the target and disabled on the source, the
+ * result would claim enabled checksums while the blocks copied from the
+ * source have none, and the target would fail checksum verification as
+ * soon as it reads them. Refuse that combination, and refuse an
+ * interrupted online transition on either side, same as pg_checksums.
+ *
+ * The opposite mismatch is allowed: recovery under the backup label
+ * written by pg_rewind adopts the checksum state as of the divergence
+ * point, so a target whose own checkpoints carry no checksums stays
+ * without them, or converges through the replayed WAL if the source
+ * enabled them online. Only warn about it, so that an offline change on
+ * the source is not overlooked.
+ *
+ * That reasoning does not hold when the target is a standby, whose
+ * checkpoints were written by its upstream primary; see the checks after
+ * findLastCheckpoint().
+ */
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted on the target server");
+ pg_log_error_hint("Start the server and shut it down cleanly to reset the state, then retry; on a standby, let replication complete the transition first.");
+ exit(1);
+ }
+ if (ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition is incomplete on the source server");
+ pg_log_error_hint("Let the transition complete, or reset the state with a clean restart, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_OFF)
+ {
+ pg_log_error("data checksums are enabled on the target server but disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, and the target would fail checksum verification after the rewind.");
+ pg_log_error_hint("Either disable data checksums on the target server or enable them on the source server, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_OFF &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_warning("data checksums are disabled on the target server but enabled on the source server");
+ pg_log_warning_detail("The rewound server will keep data checksums disabled unless the source server enabled them online after the point of divergence.");
+ }
+
/*
* Target cluster better not be running. This doesn't guard against
* someone starting the cluster concurrently. Also, this is probably more
diff --git a/src/bin/pg_rewind/pg_rewind.h b/src/bin/pg_rewind/pg_rewind.h
index 9a981f7f246..9187c3bf72f 100644
--- a/src/bin/pg_rewind/pg_rewind.h
+++ b/src/bin/pg_rewind/pg_rewind.h
@@ -39,6 +39,7 @@ extern void findLastCheckpoint(const char *datadir, XLogRecPtr forkptr,
int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
XLogRecPtr *lastchkptredo,
+ uint32 *lastchkptdatachecksums,
const char *restoreCommand);
extern XLogRecPtr readOneRecord(const char *datadir, XLogRecPtr ptr,
int tliIndex, const char *restoreCommand);
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index 0269d4c6e38..406cc946b0d 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -49,6 +49,8 @@ tests += {
't/023_concurrent_checkpoint_enable.pl',
't/024_enable_crash_after_checkpoint.pl',
't/025_cascade_divergence.pl',
+ 't/026_rewind_state.pl',
+ 't/027_rewind_standby_target.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/026_rewind_state.pl b/src/test/modules/test_checksums/t/026_rewind_state.pl
new file mode 100644
index 00000000000..b7adc968c4d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/026_rewind_state.pl
@@ -0,0 +1,133 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind checks the data checksum states of source and target.
+# Checksums enabled on the target, or at the point of divergence, with
+# a source running without them are refused: the rewound server would
+# verify checksums on blocks copied from a source that has none. The
+# opposite mismatch only warns; replay keeps the target's own state.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# Scenarios 1 and 2: checksums on from initdb. wal_log_hints keeps the
+# target eligible for pg_rewind after checksums are disabled on it.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b);
+
+# Failover to B, and divergence on A.
+$node_b->promote;
+$node_b->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_a->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_a->stop('fast');
+
+# Offline disable on the source only: enabled target, disabled source.
+$node_b->stop;
+$node_b->checksum_disable_offline;
+$node_b->start;
+test_checksum_state($node_b, 'off');
+
+my @rewind_cmd = (
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'));
+
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums are enabled on the target server but disabled on the source server/,
+ 'refuses an enabled target with a disabled source');
+
+# The refusal happens before any modification: the target still starts
+# on its own timeline.
+$node_a->start;
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'target untouched by the refused rewind');
+$node_a->stop('fast');
+
+# Scenario 2: disabling the target too makes the control files match,
+# but checksums were still enabled at the point of divergence, which is
+# the state replay on the rewound server would resume with.
+$node_a->checksum_disable_offline;
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums were enabled at the point of divergence but are disabled on the source server/,
+ 'refuses when checksums were enabled at the point of divergence');
+
+# Scenario 3: fresh pair without checksums; offline enable on the
+# source only. Allowed with a warning, and the rewound server keeps
+# checksums disabled.
+my $node_c = PostgreSQL::Test::Cluster->new('node_c');
+$node_c->init(allows_streaming => 1, no_data_checksums => 1);
+$node_c->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_c->start;
+$node_c->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_c->backup('backup');
+my $node_d = PostgreSQL::Test::Cluster->new('node_d');
+$node_d->init_from_backup($node_c, 'backup', has_streaming => 1);
+$node_d->start;
+$node_c->wait_for_catchup($node_d);
+
+$node_d->promote;
+$node_d->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_c->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_c->stop('fast');
+
+$node_d->stop;
+$node_d->checksum_enable_offline;
+$node_d->start;
+test_checksum_state($node_d, 'on');
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_c->data_dir,
+ '--source-server' => $node_d->connstr('postgres'),
+ ]);
+like(
+ $stderr,
+ qr/data checksums are disabled on the target server but enabled on the source server/,
+ 'warns for a disabled target with an enabled source');
+like($stderr, qr/Done!/, 'rewind completed despite the warning');
+
+# The rewound server follows D and keeps its own state.
+$node_c->append_conf('postgresql.conf', 'port = ' . $node_c->port);
+$node_c->enable_streaming($node_d);
+$node_c->set_standby_mode;
+$node_c->start;
+$node_d->wait_for_catchup($node_c);
+is($node_c->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+test_checksum_state($node_c, 'off');
+
+$node_c->stop;
+$node_d->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/027_rewind_standby_target.pl b/src/test/modules/test_checksums/t/027_rewind_standby_target.pl
new file mode 100644
index 00000000000..3f5a3c6be8b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/027_rewind_standby_target.pl
@@ -0,0 +1,141 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_rewind compares the data checksum state at the point of
+# divergence against the target as well as the source.
+#
+# A standby's checkpoints are written by its upstream primary, so a standby
+# whose checksums were disabled offline still has "on" checkpoints in its
+# WAL. Rewinding it onto a promoted sibling passes every control file
+# check (target off + source on only warns), yet recovery after the rewind
+# adopts the divergence checkpoint's state and the rewound server would
+# verify checksums over the pages it wrote while it was locally "off".
+# pg_rewind must refuse; after an offline enable on the target, which
+# rewrites all of its pages, the rewind goes through.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t0 AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+
+my $lagging = PostgreSQL::Test::Cluster->new('lagging');
+$lagging->init_from_backup($primary, 'backup', has_streaming => 1);
+$lagging->start;
+
+my $failover = PostgreSQL::Test::Cluster->new('failover');
+$failover->init_from_backup($primary, 'backup', has_streaming => 1);
+$failover->start;
+
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+
+test_checksum_state($primary, 'on');
+test_checksum_state($lagging, 'on');
+test_checksum_state($failover, 'on');
+
+# Offline-disable checksums on the lagging standby only. This is a
+# divergence 012_offline_standby.pl declares survivable: the standby keeps
+# its own state, warns, and stays readable.
+$lagging->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $lagging->data_dir);
+$lagging->start;
+test_checksum_state($lagging, 'off');
+
+# Give the lagging standby plenty of pages to write while it is locally
+# "off", and force a restartpoint so they reach disk without checksums.
+# The other standby replays the same WAL with checksums on.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+$lagging->safe_psql('postgres', "CHECKPOINT;");
+
+# Take the "failover" standby out of the picture. It is promoted from
+# this position later, so everything replayed from here on is the part of
+# the WAL that diverges and that pg_rewind will roll back. t1 is already
+# replayed everywhere, so its blocks are *not* rolled back.
+$failover->stop;
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$primary->wait_for_catchup($lagging);
+
+# Failover: promote the other standby, which forks the timeline behind the
+# replay position of the lagging standby.
+$primary->stop('fast');
+$failover->start;
+$failover->promote;
+$failover->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$failover->safe_psql('postgres', "CHECKPOINT;");
+
+# Re-attaching the lagging standby with pg_rewind must be refused: the
+# divergence checkpoint says "on" while the target is "off", so the rewound
+# server would verify checksums over the blocks it keeps.
+$lagging->stop;
+
+command_fails_like(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ],
+ qr/data checksums were enabled at the point of divergence but are disabled on the target server/,
+ 'pg_rewind refuses a target whose divergence checkpoint has checksums');
+
+# Enabling checksums offline on the target rewrites all of its pages, after
+# which the rewind is safe.
+system_or_bail('pg_checksums', '--enable', '--pgdata', $lagging->data_dir);
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ]);
+like($stderr, qr/Done!/, 'pg_rewind completes after the offline enable');
+
+$lagging->append_conf('postgresql.conf', 'port = ' . $lagging->port);
+$lagging->enable_streaming($failover);
+$lagging->set_standby_mode;
+$lagging->start;
+
+my ($rc, $out, $err) =
+ $lagging->psql('postgres',
+ "SELECT setting FROM pg_settings " . "WHERE name = 'data_checksums';");
+is($rc, 0, 'rewound standby accepts connections') or diag("stderr: $err");
+is($out, 'on', 'rewound standby resumes with checksums on');
+
+($rc, $out, $err) = $lagging->psql('postgres', "SELECT count(*) FROM t1;");
+is($rc, 0, 'blocks kept from the target stay readable')
+ or diag("stderr: $err");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($lagging->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the rewound standby');
+
+$lagging->stop('immediate');
+$failover->stop;
+done_testing();
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v4-0004-pg_combinebackup-Refuse-mixed-data-checksum-state.patch (9.5K, ../../188A1307-1A92-45CA-9DBE-FB0962D3756E@yesql.se/5-v4-0004-pg_combinebackup-Refuse-mixed-data-checksum-state.patch)
download | inline diff:
From 28bfea6760fcc044c76b2dc4df29278841f0d2f7 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Tue, 25 Aug 2026 15:06:58 +0000
Subject: [PATCH v4 4/4] pg_combinebackup: Refuse mixed data checksum states in
a backup chain
check_control_files() detected a chain whose backups were taken under
different data checksum states, warned, and proceeded. The output
directory keeps the last backup's control file, so a full backup taken
with checksums off combined with an incremental taken after an offline
enable produces a cluster whose control file says "on" while most of
its blocks carry no checksums; it starts, and then every connection
dies on the first unchecksummed catalog page. An offline enable
between two backups of a chain is all it takes, since it rewrites
every page without logging anything, so the incremental backup does
not re-ship the pages.
Turn the warning into an error, matching what pg_rewind does for the
equivalent combinations. The check stays asymmetric on purpose: when
the last backup was taken with checksums off, stale checksums from an
earlier backup are never verified, and that chain remains usable.
---
doc/src/sgml/ref/pg_combinebackup.sgml | 19 +--
src/bin/pg_combinebackup/pg_combinebackup.c | 14 +-
src/test/modules/test_checksums/meson.build | 1 +
.../t/028_combinebackup_mixed.pl | 134 ++++++++++++++++++
4 files changed, 155 insertions(+), 13 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/028_combinebackup_mixed.pl
diff --git a/doc/src/sgml/ref/pg_combinebackup.sgml b/doc/src/sgml/ref/pg_combinebackup.sgml
index 9a6d201e0b8..f5d7e177b36 100644
--- a/doc/src/sgml/ref/pg_combinebackup.sgml
+++ b/doc/src/sgml/ref/pg_combinebackup.sgml
@@ -306,18 +306,19 @@ PostgreSQL documentation
<para>
<literal>pg_combinebackup</literal> does not recompute page checksums when
- writing the output directory. Therefore, if any of the backups used for
- reconstruction were taken with checksums disabled, but the final backup was
- taken with checksums enabled, the resulting directory may contain pages
- with invalid checksums.
+ writing the output directory. It therefore refuses a chain in which some
+ of the backups used for reconstruction were taken with checksums disabled
+ but the final backup was taken with checksums enabled: the resulting
+ directory would contain pages with invalid checksums, and the cluster
+ would fail checksum verification as soon as it reads them. Take a new
+ full backup after enabling data checksums with
+ <xref linkend="app-pgchecksums"/>.
</para>
<para>
- To avoid this problem, taking a new full backup after changing the checksum
- state of the cluster using <xref linkend="app-pgchecksums"/> is
- recommended. Otherwise, you can disable and then optionally reenable
- checksums on the directory produced by <literal>pg_combinebackup</literal>
- in order to correct the problem.
+ The reverse case is accepted: when the final backup was taken with
+ checksums disabled, stale checksums copied from an earlier backup are
+ never verified.
</para>
</refsect1>
diff --git a/src/bin/pg_combinebackup/pg_combinebackup.c b/src/bin/pg_combinebackup/pg_combinebackup.c
index 86e2ee37c40..254a27b125b 100644
--- a/src/bin/pg_combinebackup/pg_combinebackup.c
+++ b/src/bin/pg_combinebackup/pg_combinebackup.c
@@ -670,13 +670,19 @@ check_control_files(int n_backups, char **backup_dirs)
pg_log_debug("system identifier is %" PRIu64, system_identifier);
/*
- * Warn the user if not all backups are in the same state with regards to
- * checksums.
+ * Reject a chain whose backups were not all in the same state with
+ * regards to checksums. The loop above only flags that when the last
+ * backup has checksums enabled: the blocks taken from the older backups
+ * have no checksums, and pg_combinebackup does not recompute them, so the
+ * combined cluster would fail verification as soon as it reads them. The
+ * other direction is harmless, since stale checksums are never verified
+ * while checksums are disabled.
*/
if (data_checksum_mismatch)
{
- pg_log_warning("only some backups have checksums enabled");
- pg_log_warning_hint("Disable, and optionally reenable, checksums on the output directory to avoid failures.");
+ pg_log_error("only some backups have checksums enabled");
+ pg_log_error_hint("Take a new full backup after changing the data checksum state with pg_checksums.");
+ exit(1);
}
return system_identifier;
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index 406cc946b0d..3db744859fc 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -51,6 +51,7 @@ tests += {
't/025_cascade_divergence.pl',
't/026_rewind_state.pl',
't/027_rewind_standby_target.pl',
+ 't/028_combinebackup_mixed.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/028_combinebackup_mixed.pl b/src/test/modules/test_checksums/t/028_combinebackup_mixed.pl
new file mode 100644
index 00000000000..ce9dee9413b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/028_combinebackup_mixed.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_combinebackup refuses a chain whose backups were taken under
+# different data checksum states when the final state has them enabled.
+#
+# The output directory keeps the last backup's control file, so a full
+# backup taken with checksums off plus an incremental taken after an
+# offline enable would produce a directory whose control file says "on"
+# while most of its blocks have no checksums. An offline enable between
+# the two backups is enough to get there: it rewrites every page but logs
+# nothing, so the incremental backup does not re-ship the pages.
+#
+# The reverse order stays allowed: with the final backup taken with
+# checksums off, stale checksums from an earlier backup are never verified.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $node = PostgreSQL::Test::Cluster->new('node');
+$node->init(
+ no_data_checksums => 1,
+ has_archiving => 1,
+ allows_streaming => 1);
+$node->append_conf('postgresql.conf', 'summarize_wal = on');
+$node->append_conf('postgresql.conf', 'autovacuum = off');
+$node->start;
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+test_checksum_state($node, 'off');
+
+$node->backup('full');
+
+# Offline enable: rewrites every page, logs nothing.
+$node->stop;
+system_or_bail('pg_checksums', '--enable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'on');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full/backup_manifest',
+ ],
+ 'incremental backup with checksums on');
+
+# Combining the chain must be refused: the result would say "on" while the
+# blocks inherited from the full backup have no checksums.
+command_fails_like(
+ [
+ 'pg_combinebackup',
+ $node->backup_dir . '/full',
+ $node->backup_dir . '/incr',
+ '--output' => $node->backup_dir . '/combined',
+ ],
+ qr/only some backups have checksums enabled/,
+ 'pg_combinebackup refuses a chain crossing an offline enable');
+
+# The other direction: full backup with checksums on, offline disable, then
+# an incremental. The last backup wins, so the combined cluster comes up off.
+$node->backup('full2');
+
+$node->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr2',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full2/backup_manifest',
+ ],
+ 'incremental backup with checksums off');
+
+my $restored = PostgreSQL::Test::Cluster->new('restored');
+$restored->init_from_backup(
+ $node, 'incr2',
+ combine_with_prior => ['full2'],
+ has_restoring => 1,
+ standby => 0);
+
+$restored->start;
+
+my ($rc, $stdout, $stderr) = $restored->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'combined backup accepts connections') or diag("stderr: $stderr");
+is($stdout, 'off', 'combined backup follows the last backup state');
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t3;');
+is($rc, 0, 'blocks from the incremental backup are readable')
+ or diag("stderr: $stderr");
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t1;');
+is($rc, 0, 'blocks inherited from the full backup are readable')
+ or diag("stderr: $stderr");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($restored->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures in the combined backup');
+
+$restored->stop('immediate');
+$node->stop;
+
+done_testing();
--
2.39.3 (Apple Git-146)
=
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-28 05:33 Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
parent: Daniel Gustafsson <daniel@yesql.se>
0 siblings, 1 reply; 43+ messages in thread
From: Bertrand Drouvot @ 2026-08-28 05:33 UTC (permalink / raw)
To: Daniel Gustafsson <daniel@yesql.se>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; Zsolt Parragi <zsolt.parragi@percona.com>
Hi,
On Wed, Aug 26, 2026 at 10:37:09PM +0200, Daniel Gustafsson wrote:
> > On 14 Aug 2026, at 17:27, Bertrand Drouvot <bertranddrouvot.pg@gmail.com> wrote:
>
> > Looking forward to seeing your and Zsolt's proposals.
>
> This has now been worked on quite extensively by Zsolt, myself and Tomas Vondra
> and a number of patchrevisions have been created and rewritten.
Thanks for looking at it!
> The regression in a correctly done offline change is due to combining
> pg_checksums which rewrite data on risk without WAL logging (or any logging at
> all) the transformation, with online checksums which WAL log the state change.
> With online checksums, the local state on the standby was overwritten during
> replay by the dataChecksumState in the checkpoint.
Agreed.
> The fix in 0001 is to not adopt the
> state change from the replay of checkpoints, only from XLOG2_CHECKSUMS records,
> and to alert the user with a log entry if the states mismatch.
That makes sense to me and matches the intent of my v1 for this regression: keep
offline changes local while still applying WAL online transitions, with the useful
addition of a warning on mismatch.
> The 0001 patch is the least invasive patch to solve the regression that either
> of us has managed to come up with, but it's still far from trivial.
Only looking at 0001 here, I've a few comments:
=== 1
@@ -9288,17 +9572,33 @@ xlog2_redo(XLogReaderState *record)
SpinLockAcquire(&XLogCtl->info_lck);
XLogCtl->data_checksum_version = state.new_checksum_state;
+ SetLocalDataChecksumState(state.new_checksum_state);
SpinLockRelease(&XLogCtl->info_lck);
This applies every XLOG2_CHECKSUMS record encountered during recovery, even when
the same record was applied before.
For example, a standby can replay the final "on" record and then stop cleanly
without advancing its restartpoint beyond that record. If checksums are subsequently
disabled offline, the next startup begins from the older restartpoint and replays
the same on record again, overriding the offline disable.
=== 2
+ * checkPoint.dataChecksumState was sampled while holding the WAL insert
+ * locks, so it is the state in effect at the redo point.
.
.
.
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checkPoint.dataChecksumState == XLogCtl->data_checksum_version)
+ ControlFile->data_checksum_version = checkPoint.dataChecksumState;
+ SpinLockRelease(&XLogCtl->info_lck);
I’m not sure the "state in effect at the redo point" is always correct. There is
a window between inserting the checksum transition record and updating
XLogCtl->data_checksum_version.
XLOG2_CHECKSUMS(on) records the target state. The transition can therefore
proceed as follows:
1. The current shared state is inprogress-on.
2. XLogChecksums() inserts XLOG2_CHECKSUMS(on) and releases its WAL insertion
lock.
3. Before the shared state is updated to on, the checkpoint still reads
inprogress-on and inserts XLOG_CHECKPOINT_REDO.
4. The transition then updates the shared state to on.
The WAL order is then:
XLOG2_CHECKSUMS(on)
XLOG_CHECKPOINT_REDO(inprogress-on)
If the server crashes after the concurrent checkpoint from step 3 completes,
but before the enabling operation’s later checkpoint completes, recovery starts
from that redo point and does not replay the preceding on record. It can
therefore resolve inprogress-on back to off. The equality check above does
not repair this ordering.
=== 3
+ if ((haveBackupLabel || XLogRecPtrIsValid(ControlFile->backupStartPoint)) &&
+ !(XLogRecPtrIsValid(ControlFile->backupEndPoint) &&
+ ControlFile->backupEndRequired))
+ {
+ if (wasShutdown)
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState);
+ else
+ adoptChecksumStateFromNextCheckpoint = true;
+ }
IIUC, this can overwrite the checksum state copied from the source during
pg_rewind with the state from the last common checkpoint.
For example, if the common checkpoint says off, but both source and target
were enabled offline after divergence, recovery adopts off. Since the
offline enable has no XLOG2_CHECKSUMS record, nothing restores on.
I have only looked at 0001 for this point, so I don't know whether one of the
following patches handles this case.
=== 4
+ /*
+ * Note the data checksum state the flush below starts under. Replay runs
+ * concurrently and can change the state while the flush is in progress,
+ * in which case the flush covers pages written under both states; see
+ * where the state is persisted further down.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
CheckPointGuts(lastCheckPoint.redo, flags);
+ * Persist only if the flush above ran under one state throughout; see
+ * CreateCheckPoint() for why.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksum_state == XLogCtl->data_checksum_version)
+ ControlFile->data_checksum_version = checksum_state;
+ SpinLockRelease(&XLogCtl->info_lck);
I think comparing only data_checksum_version cannot detect a complete
on->off->on transition during CheckPointGuts(). The initial and final values
match even though the flush ran under multiple states.
FWIW, while v4-0001 may address other issues present in v1, v1 would avoid the
specific cases described in === 1 through === 3. Some parts of it may therefore
be worth considering here.
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-28 06:53 Daniel Gustafsson <daniel@yesql.se>
parent: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 1 reply; 43+ messages in thread
From: Daniel Gustafsson @ 2026-08-28 06:53 UTC (permalink / raw)
To: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; Zsolt Parragi <zsolt.parragi@percona.com>
> On 28 Aug 2026, at 07:33, Bertrand Drouvot <bertranddrouvot.pg@gmail.com> wrote:
> Only looking at 0001 here, I've a few comments:
Thanks, I've yet to dig into it completely but below are a few quick questions
to help me along the way.
> === 1
>
> @@ -9288,17 +9572,33 @@ xlog2_redo(XLogReaderState *record)
>
> SpinLockAcquire(&XLogCtl->info_lck);
> XLogCtl->data_checksum_version = state.new_checksum_state;
> + SetLocalDataChecksumState(state.new_checksum_state);
> SpinLockRelease(&XLogCtl->info_lck);
>
> This applies every XLOG2_CHECKSUMS record encountered during recovery, even when
> the same record was applied before.
>
> For example, a standby can replay the final "on" record and then stop cleanly
> without advancing its restartpoint beyond that record. If checksums are subsequently
> disabled offline, the next startup begins from the older restartpoint and replays
> the same on record again, overriding the offline disable.
Do you mean that checksums are disabled offline across the cluster on all
nodes, or just on the standby?
> === 2
>
> + * checkPoint.dataChecksumState was sampled while holding the WAL insert
> + * locks, so it is the state in effect at the redo point.
> .
> .
> .
> + SpinLockAcquire(&XLogCtl->info_lck);
> + if (checkPoint.dataChecksumState == XLogCtl->data_checksum_version)
> + ControlFile->data_checksum_version = checkPoint.dataChecksumState;
> + SpinLockRelease(&XLogCtl->info_lck);
>
> I’m not sure the "state in effect at the redo point" is always correct. There is
> a window between inserting the checksum transition record and updating
> XLogCtl->data_checksum_version.
>
> XLOG2_CHECKSUMS(on) records the target state. The transition can therefore
> proceed as follows:
>
> 1. The current shared state is inprogress-on.
> 2. XLogChecksums() inserts XLOG2_CHECKSUMS(on) and releases its WAL insertion
> lock.
> 3. Before the shared state is updated to on, the checkpoint still reads
> inprogress-on and inserts XLOG_CHECKPOINT_REDO.
> 4. The transition then updates the shared state to on.
>
> The WAL order is then:
>
> XLOG2_CHECKSUMS(on)
> XLOG_CHECKPOINT_REDO(inprogress-on)
>
> If the server crashes after the concurrent checkpoint from step 3 completes,
> but before the enabling operation’s later checkpoint completes, recovery starts
> from that redo point and does not replay the preceding on record. It can
> therefore resolve inprogress-on back to off. The equality check above does
> not repair this ordering.
If this can happen then online checksums wouldn't work at all right? This
window is happening inside a critical section while DELAY_CHKPT_START is set to
prevent a checkpoint from storing the state and completing to protect against
this. Have you been able to construct a repro (with injection points) where a
REDO record after a CHECKSUM record carries the wrong state?
--
Daniel Gustafsson
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-28 09:24 Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
parent: Daniel Gustafsson <daniel@yesql.se>
0 siblings, 1 reply; 43+ messages in thread
From: Bertrand Drouvot @ 2026-08-28 09:24 UTC (permalink / raw)
To: Daniel Gustafsson <daniel@yesql.se>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; Zsolt Parragi <zsolt.parragi@percona.com>
Hi,
On Fri, Aug 28, 2026 at 08:53:38AM +0200, Daniel Gustafsson wrote:
> > On 28 Aug 2026, at 07:33, Bertrand Drouvot <bertranddrouvot.pg@gmail.com> wrote:
>
> > Only looking at 0001 here, I've a few comments:
>
> Thanks, I've yet to dig into it completely but below are a few quick questions
> to help me along the way.
>
> > === 1
> >
> > @@ -9288,17 +9572,33 @@ xlog2_redo(XLogReaderState *record)
> >
> > SpinLockAcquire(&XLogCtl->info_lck);
> > XLogCtl->data_checksum_version = state.new_checksum_state;
> > + SetLocalDataChecksumState(state.new_checksum_state);
> > SpinLockRelease(&XLogCtl->info_lck);
> >
> > This applies every XLOG2_CHECKSUMS record encountered during recovery, even when
> > the same record was applied before.
> >
> > For example, a standby can replay the final "on" record and then stop cleanly
> > without advancing its restartpoint beyond that record. If checksums are subsequently
> > disabled offline, the next startup begins from the older restartpoint and replays
> > the same on record again, overriding the offline disable.
>
> Do you mean that checksums are disabled offline across the cluster on all
> nodes, or just on the standby?
Disabling checksums offline on the standby is sufficient although that is not the
intended procedure.
Disabling offline on both the primary and standby also produce the issue.
> > === 2
> >
>
> If this can happen then online checksums wouldn't work at all right?
You’re right, my previous explanation was not fully accurate.
The 0001-specific concern is that a checkpoint can capture
checkPoint.dataChecksumState as inprogress-on, then insert XLOG_CHECKPOINT_REDO
correctly carrying on. The delay protects the flush, but the earlier value remains
stale. The equality check then does not persist on, and recovery no longer adopts
it from the REDO record, so a crash before the following checkpoint completes can
resolve the state back to off.
> Have you been able to construct a repro (with injection points) where a
> REDO record after a CHECKSUM record carries the wrong state?
Not with an injection point, but you can repro that way:
In xlog.c add 3 sleeps (see repro.txt attached):
- In SetDataChecksumsOn() to hold the launcher at inprogress-on.
- In SetDataChecksumsOn() to park it at on before its own checkpoint.
- In CreateCheckPoint() sleep/spin until XLogCtl->data_checksum_version == on.
Then:
start a cluster with initdb --no-data-checksums
Run SELECT pg_enable_data_checksums()
Then within 60s run CHECKPOINT
Once the checkpoint completes (SHOW data_checksums = on but pg_controldata still shows version 3)
pkill -9 the cluster
restart
check SHOW data_checksums: it comes back off.
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index fa24b00e7d4..672ea66cb0a 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -4858,6 +4858,10 @@ SetDataChecksumsOn(void)
SpinLockRelease(&XLogCtl->info_lck);
INJECTION_POINT("datachecksums-enable-checksums-delay", NULL);
+ ereport(LOG, (errmsg("BDTTESTHACK: launcher holding at inprogress-on for 60s")));
+ pg_usleep(60 * 1000000L);
+ ereport(LOG, (errmsg("BDTTESTHACK: launcher releasing, about to write XLOG2_CHECKSUMS(on) and flip")));
+
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
@@ -4874,6 +4878,10 @@ SetDataChecksumsOn(void)
INJECTION_POINT("datachecksums-on-before-checkpoint", NULL);
+ ereport(LOG, (errmsg("BDTTESTHACK: launcher holding at inprogress-on for 60s")));
+ pg_usleep(60 * 1000000L);
+ ereport(LOG, (errmsg("BDTTESTHACK: launcher releasing, about to write XLOG2_CHECKSUMS(on) and flip")));
+
RequestCheckpoint(CHECKPOINT_FORCE | CHECKPOINT_WAIT | CHECKPOINT_FAST);
INJECTION_POINT("datachecksums-on-after-checkpoint", NULL);
@@ -7795,6 +7803,22 @@ CreateCheckPoint(int flags)
*/
WALInsertLockRelease();
+ if (!shutdown && checkPoint.dataChecksumState == PG_DATA_CHECKSUM_INPROGRESS_ON)
+ {
+ int i;
+ ereport(LOG, (errmsg("BDTTESTHACK: checkpoint sampled dataChecksumState=inprogress-on, waiting for flip to on")));
+ for (i = 0; i < 600; i++) /* up to ~60s, then give up */
+ {
+ uint32 v;
+ SpinLockAcquire(&XLogCtl->info_lck);
+ v = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+ if (v == PG_DATA_CHECKSUM_VERSION) /* "on" */
+ break;
+ pg_usleep(100 * 1000L); /* 100ms */
+ }
+ ereport(LOG, (errmsg("TESTHACK: checkpoint resuming after %d iterations; will sample redo_rec and write XLOG_CHECKPOINT_REDO", i)));
+ }
/*
* If this is an online checkpoint, we have not yet determined the redo
* point. We do so now by inserting the special XLOG_CHECKPOINT_REDO
Attachments:
[text/plain] repro.txt (2.0K, ../../apFTwa68Qn0S0dIX@bdtpg/2-repro.txt)
download | inline diff:
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index fa24b00e7d4..672ea66cb0a 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -4858,6 +4858,10 @@ SetDataChecksumsOn(void)
SpinLockRelease(&XLogCtl->info_lck);
INJECTION_POINT("datachecksums-enable-checksums-delay", NULL);
+ ereport(LOG, (errmsg("BDTTESTHACK: launcher holding at inprogress-on for 60s")));
+ pg_usleep(60 * 1000000L);
+ ereport(LOG, (errmsg("BDTTESTHACK: launcher releasing, about to write XLOG2_CHECKSUMS(on) and flip")));
+
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
@@ -4874,6 +4878,10 @@ SetDataChecksumsOn(void)
INJECTION_POINT("datachecksums-on-before-checkpoint", NULL);
+ ereport(LOG, (errmsg("BDTTESTHACK: launcher holding at inprogress-on for 60s")));
+ pg_usleep(60 * 1000000L);
+ ereport(LOG, (errmsg("BDTTESTHACK: launcher releasing, about to write XLOG2_CHECKSUMS(on) and flip")));
+
RequestCheckpoint(CHECKPOINT_FORCE | CHECKPOINT_WAIT | CHECKPOINT_FAST);
INJECTION_POINT("datachecksums-on-after-checkpoint", NULL);
@@ -7795,6 +7803,22 @@ CreateCheckPoint(int flags)
*/
WALInsertLockRelease();
+ if (!shutdown && checkPoint.dataChecksumState == PG_DATA_CHECKSUM_INPROGRESS_ON)
+ {
+ int i;
+ ereport(LOG, (errmsg("BDTTESTHACK: checkpoint sampled dataChecksumState=inprogress-on, waiting for flip to on")));
+ for (i = 0; i < 600; i++) /* up to ~60s, then give up */
+ {
+ uint32 v;
+ SpinLockAcquire(&XLogCtl->info_lck);
+ v = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+ if (v == PG_DATA_CHECKSUM_VERSION) /* "on" */
+ break;
+ pg_usleep(100 * 1000L); /* 100ms */
+ }
+ ereport(LOG, (errmsg("TESTHACK: checkpoint resuming after %d iterations; will sample redo_rec and write XLOG_CHECKPOINT_REDO", i)));
+ }
/*
* If this is an online checkpoint, we have not yet determined the redo
* point. We do so now by inserting the special XLOG_CHECKPOINT_REDO
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-28 11:16 Zsolt Parragi <zsolt.parragi@percona.com>
parent: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 1 reply; 43+ messages in thread
From: Zsolt Parragi @ 2026-08-28 11:16 UTC (permalink / raw)
To: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: pgsql-hackers@lists.postgresql.org, Daniel Gustafsson <daniel@yesql.se>
Thanks!
I agree that v4 is not a complete fix, and we could make it better.
The question is the balance, as every change also makes it more
complex. The main point of it is to try to minimize how invasive of a
patch it is, and making sure that offline changes work and don't
result in completely breaking a standby.
1: this is a valid issue, but requires somebody doing an offline
change immediately after an online change. The effect is that in this
case, the standby might roll back the offline change, but it will
print out a warning about this into the log, so this is visible, and
everything will continue working.
2: my understanding is that if there's a concurrent checkpoint and a
crash shortly after it, we might throw away an otherwise completed
online checksum after restarting. We properly log that checksums were
interrupted, and the state remains "off" on all nodes. While this is
not ideal, I think this is an unlikely scenario and not the only such
issue, for example a failing DROP DATABASE foo FORCE similarly can
interrupt checksums in an unlikely case, as I reported in another
thread.
3: also valid, but in my repro of this the warning fired, so it's not
silent, and things seem to work fine after the warning, and the user
can issue either an online or an offline change.
4: I couldn't construct a repro for this case, I think this can only
happen in theory in very specific engineered scenarios
> FWIW, while v4-0001 may address other issues present in v1, v1 would avoid the
> specific cases described in === 1 through === 3. Some parts of it may therefore
> be worth considering here.
I agree that combining the two patches would be the best solution in
the warning direction, e.g. solving 1+3 requires the pg_control
changes from v1. The reason I left that out is what I started with in
this reply: simplicity. I was mainly considering combining the two
because of v4 can emit spurious warnings in some cases (and then the
additional log state stating the correction), but even with these I am
not sure if we should make it more complex, as none of these result in
crashes/data corruption, only in state rolling back in some
engineering situations. I'll try to look into what adding the two
patches together looks like, but it most likely combines their size,
as they improve the current master code in different ways.
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-28 15:13 Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
parent: Zsolt Parragi <zsolt.parragi@percona.com>
0 siblings, 1 reply; 43+ messages in thread
From: Bertrand Drouvot @ 2026-08-28 15:13 UTC (permalink / raw)
To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: pgsql-hackers@lists.postgresql.org, Daniel Gustafsson <daniel@yesql.se>
Hi,
On Fri, Aug 28, 2026 at 04:16:55AM -0700, Zsolt Parragi wrote:
> Thanks!
>
> I agree that v4 is not a complete fix, and we could make it better.
> The question is the balance, as every change also makes it more
> complex. The main point of it is to try to minimize how invasive of a
> patch it is, and making sure that offline changes work and don't
> result in completely breaking a standby.
>
> 1: this is a valid issue, but requires somebody doing an offline
> change immediately after an online change.
I don't think it has to be immediate. The window lasts until the standby records
a restartpoint after the XLOG2_CHECKSUMS(on) record. That may happen considerably
later, depending on checkpoint replay and restartpoint creation.
> The effect is that in this
> print out a warning about this into the log, so this is visible, and
> everything will continue working.
Yes, but that leaves mismatched states and that's what we try to avoid.
> 2: my understanding is that if there's a concurrent checkpoint and a
> crash shortly after it, we might throw away an otherwise completed
> online checksum after restarting. We properly log that checksums were
> interrupted, and the state remains "off" on all nodes.
Right, but the online transition had reached on before the crash and is rolled
back because of 0001’s handling of the stale inprogress-on value. Even if all
nodes return to off, that still looks like an incorrect state rollback introduced
by 0001.
> While this is
> not ideal, I think this is an unlikely scenario
yeah, probably.
> and not the only such
> issue, for example a failing DROP DATABASE foo FORCE similarly can
> interrupt checksums in an unlikely case, as I reported in another
> thread.
I think our case is different.
The checksum state has reached on, and both XLOG2_CHECKSUMS and the later REDO
record carry on. It returns to off due to 0001.
The sleeps in the repro only make the possible interleaving deterministic.
> 3: also valid, but in my repro of this the warning fired, so it's not
> silent, and things seem to work fine after the warning, and the user
> can issue either an online or an offline change.
The warning is useful, but I don't think it makes the resulting state correct.
I think that leaves precisely the mismatched state we are trying to avoid.
> 4: I couldn't construct a repro for this case, I think this can only
> happen in theory in very specific engineered scenarios
Yes, it's doable. While that does not lead to correctness issue, it still
questions the logic here.
> I agree that combining the two patches would be the best solution in
> the warning direction, e.g. solving 1+3 requires the pg_control
> changes from v1.
Yeah, but if the resulting patch ends up being significantly more complex, that
would not be reassuring either.
> The reason I left that out is what I started with in
> this reply: simplicity. I was mainly considering combining the two
> because of v4 can emit spurious warnings in some cases
Does that refer to the cases I reported for v4, or did you have additional cases
in mind?
> additional log state stating the correction), but even with these I am
> not sure if we should make it more complex, as none of these result in
> crashes/data corruption, only in state rolling back in some
> engineering situations.
Yeah, I understand each of these tradeoffs in isolation. What concerns me is their
cumulative effect: we moved from wanting to reject mismatched states to warning
about them, and we are now considering leaving some known rollback cases unhandled
to keep the patch manageable.
Also, I’m not sure 1 and 3 are limited to engineered situations. Even without
immediate crashes, they can leave the nodes with mismatched states, which is
what the patchset is trying to address.
> I'll try to look into what adding the two
> patches together looks like, but it most likely combines their size,
> as they improve the current master code in different ways.
Thanks!
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-28 16:04 Zsolt Parragi <zsolt.parragi@percona.com>
parent: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 1 reply; 43+ messages in thread
From: Zsolt Parragi @ 2026-08-28 16:04 UTC (permalink / raw)
To: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: pgsql-hackers@lists.postgresql.org
> Yeah, I understand each of these tradeoffs in isolation. What concerns me is their
> cumulative effect: we moved from wanting to reject mismatched states to warning
> about them, and we are now considering leaving some known rollback cases unhandled
> to keep the patch manageable.
>
> Also, I’m not sure 1 and 3 are limited to engineered situations. Even without
> immediate crashes, they can leave the nodes with mismatched states, which is
> what the patchset is trying to address.
1, for example requires executing an offline change quickly after an
online change. I'm not saying that it shouldn't work better, just
questioning how realistic that scenario is.
> Yeah, but if the resulting patch ends up being significantly more complex, that
> would not be reassuring either.
and
> > I'll try to look into what adding the two
> > patches together looks like, but it most likely combines their size,
> > as they improve the current master code in different ways.
>
> Thanks!
I looked into this, and I was right that if I add the pg_control
changes to v4 it nearly doubles the size of the actual code changes
from ~300 to ~550 lines, and fixes all the issues you reported while
also keeping the existing suite of tests passing. The diff compared to
v4 is relatively simple, so I don't think that would be an issue by
itself, but it's another control version change.
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-29 02:41 Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
parent: Zsolt Parragi <zsolt.parragi@percona.com>
0 siblings, 1 reply; 43+ messages in thread
From: Bertrand Drouvot @ 2026-08-29 02:41 UTC (permalink / raw)
To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: pgsql-hackers@lists.postgresql.org
Hi,
On Fri, Aug 28, 2026 at 11:04:28AM -0500, Zsolt Parragi wrote:
> > > I'll try to look into what adding the two
> > > patches together looks like, but it most likely combines their size,
> > > as they improve the current master code in different ways.
> >
> > Thanks!
>
> I looked into this,
Thanks!
> and I was right that if I add the pg_control
> changes to v4 it nearly doubles the size of the actual code changes
> from ~300 to ~550 lines, and fixes all the issues you reported
Yeah, that's what I expected, as v1 included those changes to handle these cases.
> while
> also keeping the existing suite of tests passing. The diff compared to
> v4 is relatively simple, so I don't think that would be an issue by
> itself,
Thanks! I don't see the patch attached. Would you mind sharing it?
> but it's another control version change.
Do you see the control version change as a concern?
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-29 21:36 Zsolt Parragi <zsolt.parragi@percona.com>
parent: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 1 reply; 43+ messages in thread
From: Zsolt Parragi @ 2026-08-29 21:36 UTC (permalink / raw)
To: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: pgsql-hackers@lists.postgresql.org
> Thanks! I don't see the patch attached. Would you mind sharing it?
Sorry, I forgot to attach it to the previous email.
> Do you see the control version change as a concern?
Yes, it is another non-trivial change in an already complex patch,
really close to RC1. It's also not an area where we could easily
implement bug fixes in a minor version, if we discover something
later.
Attachments:
[application/octet-stream] v5-0003-pg_rewind-Check-the-data-checksum-states-of-sourc.patch (21.7K, ../../CAN4CZFMoDj7mbY231r-pmPywwJpneksUiNXDwkNokkveGtobQg@mail.gmail.com/2-v5-0003-pg_rewind-Check-the-data-checksum-states-of-sourc.patch)
download | inline diff:
From c9a802f31515bee82ae067d77b5159d1a57bbfeb Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 18:40:22 +0000
Subject: [PATCH v5 3/4] pg_rewind: Check the data checksum states of source
and target
pg_rewind installs the source's control file on the target, but every
block it does not copy keeps the target's content. With checksums
enabled on the target and disabled on the source, the rewound server
claims enabled checksums while the blocks copied from the source have
none, and fails checksum verification as soon as it reads them; in the
worst case every connection attempt dies on an unverifiable catalog
page.
Comparing the control files is not enough. Replay on the rewound
server resumes from the last common checkpoint and adopts the data
checksum state recorded there, so an offline disable on the source
after the divergence leaves both control files saying "off" while the
rewound server still resumes with verification enabled. Check the
state carried by the divergence checkpoint as well, which pg_rewind
already reads.
Refuse both cases, and an interrupted online transition on either
side, same as pg_checksums does. The opposite mismatch, checksums on
the source only, is allowed with a warning: recovery under the backup
label adopts the divergence state, so the rewound server keeps
checksums disabled (or converges through the replayed WAL if the
source enabled them online), and no page can fail verification.
---
doc/src/sgml/ref/pg_rewind.sgml | 14 ++
src/bin/pg_rewind/parsexlog.c | 4 +-
src/bin/pg_rewind/pg_rewind.c | 93 +++++++++++-
src/bin/pg_rewind/pg_rewind.h | 1 +
src/test/modules/test_checksums/meson.build | 2 +
.../test_checksums/t/026_rewind_state.pl | 133 +++++++++++++++++
.../t/027_rewind_standby_target.pl | 141 ++++++++++++++++++
7 files changed, 386 insertions(+), 2 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/026_rewind_state.pl
create mode 100644 src/test/modules/test_checksums/t/027_rewind_standby_target.pl
diff --git a/doc/src/sgml/ref/pg_rewind.sgml b/doc/src/sgml/ref/pg_rewind.sgml
index b95e3868e9d..ac3d0c9328f 100644
--- a/doc/src/sgml/ref/pg_rewind.sgml
+++ b/doc/src/sgml/ref/pg_rewind.sgml
@@ -124,6 +124,20 @@ PostgreSQL documentation
<literal>on</literal>, but is enabled by default.
</para>
+ <para>
+ The data checksum states of the source and target server must be
+ compatible. <application>pg_rewind</application> refuses to run if an
+ online checksum state transition was interrupted on either server, or
+ if data checksums are enabled on the target server, or were enabled at
+ the point of divergence, while the source server runs without them: the
+ rewound server would fail checksum verification on the blocks copied
+ from the source. The opposite combination is allowed with a warning;
+ the rewound server keeps data checksums disabled unless the source
+ server enabled them online after the point of divergence. See
+ <xref linkend="app-pgchecksums"/> for how to change the state
+ consistently across a replication setup.
+ </para>
+
<warning>
<title>Warning: Failures While Rewinding</title>
<para>
diff --git a/src/bin/pg_rewind/parsexlog.c b/src/bin/pg_rewind/parsexlog.c
index 023e23b063c..6e87b00f8c2 100644
--- a/src/bin/pg_rewind/parsexlog.c
+++ b/src/bin/pg_rewind/parsexlog.c
@@ -167,7 +167,8 @@ readOneRecord(const char *datadir, XLogRecPtr ptr, int tliIndex,
void
findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
- XLogRecPtr *lastchkptredo, const char *restoreCommand)
+ XLogRecPtr *lastchkptredo, uint32 *lastchkptdatachecksums,
+ const char *restoreCommand)
{
/* Walk backwards, starting from the given record */
XLogRecord *record;
@@ -255,6 +256,7 @@ findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
*lastchkptrec = searchptr;
*lastchkpttli = checkPoint.ThisTimeLineID;
*lastchkptredo = checkPoint.redo;
+ *lastchkptdatachecksums = checkPoint.dataChecksumState;
break;
}
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 936469ea5f8..8c1357c718f 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -144,6 +144,7 @@ main(int argc, char **argv)
XLogRecPtr chkptrec;
TimeLineID chkpttli;
XLogRecPtr chkptredo;
+ uint32 chkptdatachecksums;
TimeLineID source_tli;
TimeLineID target_tli;
XLogRecPtr target_wal_endrec;
@@ -471,10 +472,51 @@ main(int argc, char **argv)
keepwal_init();
findLastCheckpoint(datadir_target, divergerec, lastcommontliIndex,
- &chkptrec, &chkpttli, &chkptredo, restore_command);
+ &chkptrec, &chkpttli, &chkptredo, &chkptdatachecksums,
+ restore_command);
pg_log_info("rewinding from last common checkpoint at %X/%08X on timeline %u",
LSN_FORMAT_ARGS(chkptrec), chkpttli);
+ /*
+ * Replay on the rewound server resumes from the last common checkpoint
+ * and adopts the data checksum state recorded there, not the state in the
+ * control file installed from the source. If checksums were enabled at
+ * the divergence point but the source runs without them, the rewound
+ * server would verify checksums while the blocks copied from the source
+ * have none. The control files cannot reveal this: an offline disable on
+ * the source after the divergence leaves both of them saying "off".
+ *
+ * Only the fully enabled state needs checking. An in-progress state at
+ * the divergence point means a transition was still running there, and
+ * all of its page rewrites are logged after that point, so replay brings
+ * the rewound server to whatever state the copied WAL ends in.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the source server or recreate the target server from a base backup.");
+ exit(1);
+ }
+
+ /*
+ * The same comparison is needed against the target: every block the
+ * rewind does not copy keeps the target's content. The control files
+ * cannot reveal this case either, in the other direction: a standby's
+ * checkpoints are written by its upstream primary, so a standby whose
+ * checksums were disabled offline still has "on" checkpoints in its WAL,
+ * while both control files may agree.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_target.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the target server");
+ pg_log_error_detail("Blocks kept from the target would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the target server with pg_checksums or recreate the target server from a base backup.");
+ exit(1);
+ }
+
/* Initialize the hash table to track the status of each file */
filehash_init();
@@ -784,6 +826,55 @@ sanityChecks(void)
pg_fatal("target server needs to use either data checksums or \"wal_log_hints = on\"");
}
+ /*
+ * The rewound target keeps its control file fields from the source, but
+ * every block the rewind does not copy keeps the target's content. If
+ * checksums are enabled on the target and disabled on the source, the
+ * result would claim enabled checksums while the blocks copied from the
+ * source have none, and the target would fail checksum verification as
+ * soon as it reads them. Refuse that combination, and refuse an
+ * interrupted online transition on either side, same as pg_checksums.
+ *
+ * The opposite mismatch is allowed: recovery under the backup label
+ * written by pg_rewind adopts the checksum state as of the divergence
+ * point, so a target whose own checkpoints carry no checksums stays
+ * without them, or converges through the replayed WAL if the source
+ * enabled them online. Only warn about it, so that an offline change on
+ * the source is not overlooked.
+ *
+ * That reasoning does not hold when the target is a standby, whose
+ * checkpoints were written by its upstream primary; see the checks after
+ * findLastCheckpoint().
+ */
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted on the target server");
+ pg_log_error_hint("Start the server and shut it down cleanly to reset the state, then retry; on a standby, let replication complete the transition first.");
+ exit(1);
+ }
+ if (ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition is incomplete on the source server");
+ pg_log_error_hint("Let the transition complete, or reset the state with a clean restart, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_OFF)
+ {
+ pg_log_error("data checksums are enabled on the target server but disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, and the target would fail checksum verification after the rewind.");
+ pg_log_error_hint("Either disable data checksums on the target server or enable them on the source server, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_OFF &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_warning("data checksums are disabled on the target server but enabled on the source server");
+ pg_log_warning_detail("The rewound server will keep data checksums disabled unless the source server enabled them online after the point of divergence.");
+ }
+
/*
* Target cluster better not be running. This doesn't guard against
* someone starting the cluster concurrently. Also, this is probably more
diff --git a/src/bin/pg_rewind/pg_rewind.h b/src/bin/pg_rewind/pg_rewind.h
index 9a981f7f246..9187c3bf72f 100644
--- a/src/bin/pg_rewind/pg_rewind.h
+++ b/src/bin/pg_rewind/pg_rewind.h
@@ -39,6 +39,7 @@ extern void findLastCheckpoint(const char *datadir, XLogRecPtr forkptr,
int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
XLogRecPtr *lastchkptredo,
+ uint32 *lastchkptdatachecksums,
const char *restoreCommand);
extern XLogRecPtr readOneRecord(const char *datadir, XLogRecPtr ptr,
int tliIndex, const char *restoreCommand);
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index fb33149b2db..1f13aa1f3dc 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -49,6 +49,8 @@ tests += {
't/023_concurrent_checkpoint_enable.pl',
't/024_enable_crash_after_checkpoint.pl',
't/025_cascade_divergence.pl',
+ 't/026_rewind_state.pl',
+ 't/027_rewind_standby_target.pl',
't/029_checkpoint_transition_race.pl',
't/030_offline_survives_rereplay.pl',
't/031_rewind_offline_enable.pl',
diff --git a/src/test/modules/test_checksums/t/026_rewind_state.pl b/src/test/modules/test_checksums/t/026_rewind_state.pl
new file mode 100644
index 00000000000..b7adc968c4d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/026_rewind_state.pl
@@ -0,0 +1,133 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind checks the data checksum states of source and target.
+# Checksums enabled on the target, or at the point of divergence, with
+# a source running without them are refused: the rewound server would
+# verify checksums on blocks copied from a source that has none. The
+# opposite mismatch only warns; replay keeps the target's own state.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# Scenarios 1 and 2: checksums on from initdb. wal_log_hints keeps the
+# target eligible for pg_rewind after checksums are disabled on it.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b);
+
+# Failover to B, and divergence on A.
+$node_b->promote;
+$node_b->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_a->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_a->stop('fast');
+
+# Offline disable on the source only: enabled target, disabled source.
+$node_b->stop;
+$node_b->checksum_disable_offline;
+$node_b->start;
+test_checksum_state($node_b, 'off');
+
+my @rewind_cmd = (
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'));
+
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums are enabled on the target server but disabled on the source server/,
+ 'refuses an enabled target with a disabled source');
+
+# The refusal happens before any modification: the target still starts
+# on its own timeline.
+$node_a->start;
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'target untouched by the refused rewind');
+$node_a->stop('fast');
+
+# Scenario 2: disabling the target too makes the control files match,
+# but checksums were still enabled at the point of divergence, which is
+# the state replay on the rewound server would resume with.
+$node_a->checksum_disable_offline;
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums were enabled at the point of divergence but are disabled on the source server/,
+ 'refuses when checksums were enabled at the point of divergence');
+
+# Scenario 3: fresh pair without checksums; offline enable on the
+# source only. Allowed with a warning, and the rewound server keeps
+# checksums disabled.
+my $node_c = PostgreSQL::Test::Cluster->new('node_c');
+$node_c->init(allows_streaming => 1, no_data_checksums => 1);
+$node_c->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_c->start;
+$node_c->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_c->backup('backup');
+my $node_d = PostgreSQL::Test::Cluster->new('node_d');
+$node_d->init_from_backup($node_c, 'backup', has_streaming => 1);
+$node_d->start;
+$node_c->wait_for_catchup($node_d);
+
+$node_d->promote;
+$node_d->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_c->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_c->stop('fast');
+
+$node_d->stop;
+$node_d->checksum_enable_offline;
+$node_d->start;
+test_checksum_state($node_d, 'on');
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_c->data_dir,
+ '--source-server' => $node_d->connstr('postgres'),
+ ]);
+like(
+ $stderr,
+ qr/data checksums are disabled on the target server but enabled on the source server/,
+ 'warns for a disabled target with an enabled source');
+like($stderr, qr/Done!/, 'rewind completed despite the warning');
+
+# The rewound server follows D and keeps its own state.
+$node_c->append_conf('postgresql.conf', 'port = ' . $node_c->port);
+$node_c->enable_streaming($node_d);
+$node_c->set_standby_mode;
+$node_c->start;
+$node_d->wait_for_catchup($node_c);
+is($node_c->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+test_checksum_state($node_c, 'off');
+
+$node_c->stop;
+$node_d->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/027_rewind_standby_target.pl b/src/test/modules/test_checksums/t/027_rewind_standby_target.pl
new file mode 100644
index 00000000000..3f5a3c6be8b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/027_rewind_standby_target.pl
@@ -0,0 +1,141 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_rewind compares the data checksum state at the point of
+# divergence against the target as well as the source.
+#
+# A standby's checkpoints are written by its upstream primary, so a standby
+# whose checksums were disabled offline still has "on" checkpoints in its
+# WAL. Rewinding it onto a promoted sibling passes every control file
+# check (target off + source on only warns), yet recovery after the rewind
+# adopts the divergence checkpoint's state and the rewound server would
+# verify checksums over the pages it wrote while it was locally "off".
+# pg_rewind must refuse; after an offline enable on the target, which
+# rewrites all of its pages, the rewind goes through.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t0 AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+
+my $lagging = PostgreSQL::Test::Cluster->new('lagging');
+$lagging->init_from_backup($primary, 'backup', has_streaming => 1);
+$lagging->start;
+
+my $failover = PostgreSQL::Test::Cluster->new('failover');
+$failover->init_from_backup($primary, 'backup', has_streaming => 1);
+$failover->start;
+
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+
+test_checksum_state($primary, 'on');
+test_checksum_state($lagging, 'on');
+test_checksum_state($failover, 'on');
+
+# Offline-disable checksums on the lagging standby only. This is a
+# divergence 012_offline_standby.pl declares survivable: the standby keeps
+# its own state, warns, and stays readable.
+$lagging->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $lagging->data_dir);
+$lagging->start;
+test_checksum_state($lagging, 'off');
+
+# Give the lagging standby plenty of pages to write while it is locally
+# "off", and force a restartpoint so they reach disk without checksums.
+# The other standby replays the same WAL with checksums on.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+$lagging->safe_psql('postgres', "CHECKPOINT;");
+
+# Take the "failover" standby out of the picture. It is promoted from
+# this position later, so everything replayed from here on is the part of
+# the WAL that diverges and that pg_rewind will roll back. t1 is already
+# replayed everywhere, so its blocks are *not* rolled back.
+$failover->stop;
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$primary->wait_for_catchup($lagging);
+
+# Failover: promote the other standby, which forks the timeline behind the
+# replay position of the lagging standby.
+$primary->stop('fast');
+$failover->start;
+$failover->promote;
+$failover->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$failover->safe_psql('postgres', "CHECKPOINT;");
+
+# Re-attaching the lagging standby with pg_rewind must be refused: the
+# divergence checkpoint says "on" while the target is "off", so the rewound
+# server would verify checksums over the blocks it keeps.
+$lagging->stop;
+
+command_fails_like(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ],
+ qr/data checksums were enabled at the point of divergence but are disabled on the target server/,
+ 'pg_rewind refuses a target whose divergence checkpoint has checksums');
+
+# Enabling checksums offline on the target rewrites all of its pages, after
+# which the rewind is safe.
+system_or_bail('pg_checksums', '--enable', '--pgdata', $lagging->data_dir);
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ]);
+like($stderr, qr/Done!/, 'pg_rewind completes after the offline enable');
+
+$lagging->append_conf('postgresql.conf', 'port = ' . $lagging->port);
+$lagging->enable_streaming($failover);
+$lagging->set_standby_mode;
+$lagging->start;
+
+my ($rc, $out, $err) =
+ $lagging->psql('postgres',
+ "SELECT setting FROM pg_settings " . "WHERE name = 'data_checksums';");
+is($rc, 0, 'rewound standby accepts connections') or diag("stderr: $err");
+is($out, 'on', 'rewound standby resumes with checksums on');
+
+($rc, $out, $err) = $lagging->psql('postgres', "SELECT count(*) FROM t1;");
+is($rc, 0, 'blocks kept from the target stay readable')
+ or diag("stderr: $err");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($lagging->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the rewound standby');
+
+$lagging->stop('immediate');
+$failover->stop;
+done_testing();
--
2.55.0
[application/octet-stream] v5-0004-pg_combinebackup-Refuse-mixed-data-checksum-state.patch (9.6K, ../../CAN4CZFMoDj7mbY231r-pmPywwJpneksUiNXDwkNokkveGtobQg@mail.gmail.com/3-v5-0004-pg_combinebackup-Refuse-mixed-data-checksum-state.patch)
download | inline diff:
From 9b17d1ab357ca9610bb2ed065e6e4b11e45a4ce0 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Tue, 25 Aug 2026 15:06:58 +0000
Subject: [PATCH v5 4/4] pg_combinebackup: Refuse mixed data checksum states in
a backup chain
check_control_files() detected a chain whose backups were taken under
different data checksum states, warned, and proceeded. The output
directory keeps the last backup's control file, so a full backup taken
with checksums off combined with an incremental taken after an offline
enable produces a cluster whose control file says "on" while most of
its blocks carry no checksums; it starts, and then every connection
dies on the first unchecksummed catalog page. An offline enable
between two backups of a chain is all it takes, since it rewrites
every page without logging anything, so the incremental backup does
not re-ship the pages.
Turn the warning into an error, matching what pg_rewind does for the
equivalent combinations. The check stays asymmetric on purpose: when
the last backup was taken with checksums off, stale checksums from an
earlier backup are never verified, and that chain remains usable.
---
doc/src/sgml/ref/pg_combinebackup.sgml | 19 +--
src/bin/pg_combinebackup/pg_combinebackup.c | 14 +-
src/test/modules/test_checksums/meson.build | 1 +
.../t/028_combinebackup_mixed.pl | 134 ++++++++++++++++++
4 files changed, 155 insertions(+), 13 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/028_combinebackup_mixed.pl
diff --git a/doc/src/sgml/ref/pg_combinebackup.sgml b/doc/src/sgml/ref/pg_combinebackup.sgml
index 9a6d201e0b8..f5d7e177b36 100644
--- a/doc/src/sgml/ref/pg_combinebackup.sgml
+++ b/doc/src/sgml/ref/pg_combinebackup.sgml
@@ -306,18 +306,19 @@ PostgreSQL documentation
<para>
<literal>pg_combinebackup</literal> does not recompute page checksums when
- writing the output directory. Therefore, if any of the backups used for
- reconstruction were taken with checksums disabled, but the final backup was
- taken with checksums enabled, the resulting directory may contain pages
- with invalid checksums.
+ writing the output directory. It therefore refuses a chain in which some
+ of the backups used for reconstruction were taken with checksums disabled
+ but the final backup was taken with checksums enabled: the resulting
+ directory would contain pages with invalid checksums, and the cluster
+ would fail checksum verification as soon as it reads them. Take a new
+ full backup after enabling data checksums with
+ <xref linkend="app-pgchecksums"/>.
</para>
<para>
- To avoid this problem, taking a new full backup after changing the checksum
- state of the cluster using <xref linkend="app-pgchecksums"/> is
- recommended. Otherwise, you can disable and then optionally reenable
- checksums on the directory produced by <literal>pg_combinebackup</literal>
- in order to correct the problem.
+ The reverse case is accepted: when the final backup was taken with
+ checksums disabled, stale checksums copied from an earlier backup are
+ never verified.
</para>
</refsect1>
diff --git a/src/bin/pg_combinebackup/pg_combinebackup.c b/src/bin/pg_combinebackup/pg_combinebackup.c
index 86e2ee37c40..254a27b125b 100644
--- a/src/bin/pg_combinebackup/pg_combinebackup.c
+++ b/src/bin/pg_combinebackup/pg_combinebackup.c
@@ -670,13 +670,19 @@ check_control_files(int n_backups, char **backup_dirs)
pg_log_debug("system identifier is %" PRIu64, system_identifier);
/*
- * Warn the user if not all backups are in the same state with regards to
- * checksums.
+ * Reject a chain whose backups were not all in the same state with
+ * regards to checksums. The loop above only flags that when the last
+ * backup has checksums enabled: the blocks taken from the older backups
+ * have no checksums, and pg_combinebackup does not recompute them, so the
+ * combined cluster would fail verification as soon as it reads them. The
+ * other direction is harmless, since stale checksums are never verified
+ * while checksums are disabled.
*/
if (data_checksum_mismatch)
{
- pg_log_warning("only some backups have checksums enabled");
- pg_log_warning_hint("Disable, and optionally reenable, checksums on the output directory to avoid failures.");
+ pg_log_error("only some backups have checksums enabled");
+ pg_log_error_hint("Take a new full backup after changing the data checksum state with pg_checksums.");
+ exit(1);
}
return system_identifier;
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index 1f13aa1f3dc..927c968c054 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -51,6 +51,7 @@ tests += {
't/025_cascade_divergence.pl',
't/026_rewind_state.pl',
't/027_rewind_standby_target.pl',
+ 't/028_combinebackup_mixed.pl',
't/029_checkpoint_transition_race.pl',
't/030_offline_survives_rereplay.pl',
't/031_rewind_offline_enable.pl',
diff --git a/src/test/modules/test_checksums/t/028_combinebackup_mixed.pl b/src/test/modules/test_checksums/t/028_combinebackup_mixed.pl
new file mode 100644
index 00000000000..ce9dee9413b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/028_combinebackup_mixed.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_combinebackup refuses a chain whose backups were taken under
+# different data checksum states when the final state has them enabled.
+#
+# The output directory keeps the last backup's control file, so a full
+# backup taken with checksums off plus an incremental taken after an
+# offline enable would produce a directory whose control file says "on"
+# while most of its blocks have no checksums. An offline enable between
+# the two backups is enough to get there: it rewrites every page but logs
+# nothing, so the incremental backup does not re-ship the pages.
+#
+# The reverse order stays allowed: with the final backup taken with
+# checksums off, stale checksums from an earlier backup are never verified.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $node = PostgreSQL::Test::Cluster->new('node');
+$node->init(
+ no_data_checksums => 1,
+ has_archiving => 1,
+ allows_streaming => 1);
+$node->append_conf('postgresql.conf', 'summarize_wal = on');
+$node->append_conf('postgresql.conf', 'autovacuum = off');
+$node->start;
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+test_checksum_state($node, 'off');
+
+$node->backup('full');
+
+# Offline enable: rewrites every page, logs nothing.
+$node->stop;
+system_or_bail('pg_checksums', '--enable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'on');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full/backup_manifest',
+ ],
+ 'incremental backup with checksums on');
+
+# Combining the chain must be refused: the result would say "on" while the
+# blocks inherited from the full backup have no checksums.
+command_fails_like(
+ [
+ 'pg_combinebackup',
+ $node->backup_dir . '/full',
+ $node->backup_dir . '/incr',
+ '--output' => $node->backup_dir . '/combined',
+ ],
+ qr/only some backups have checksums enabled/,
+ 'pg_combinebackup refuses a chain crossing an offline enable');
+
+# The other direction: full backup with checksums on, offline disable, then
+# an incremental. The last backup wins, so the combined cluster comes up off.
+$node->backup('full2');
+
+$node->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr2',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full2/backup_manifest',
+ ],
+ 'incremental backup with checksums off');
+
+my $restored = PostgreSQL::Test::Cluster->new('restored');
+$restored->init_from_backup(
+ $node, 'incr2',
+ combine_with_prior => ['full2'],
+ has_restoring => 1,
+ standby => 0);
+
+$restored->start;
+
+my ($rc, $stdout, $stderr) = $restored->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'combined backup accepts connections') or diag("stderr: $stderr");
+is($stdout, 'off', 'combined backup follows the last backup state');
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t3;');
+is($rc, 0, 'blocks from the incremental backup are readable')
+ or diag("stderr: $stderr");
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t1;');
+is($rc, 0, 'blocks inherited from the full backup are readable')
+ or diag("stderr: $stderr");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($restored->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures in the combined backup');
+
+$restored->stop('immediate');
+$node->stop;
+
+done_testing();
--
2.55.0
[application/octet-stream] v5-0002-pg_checksums-Refuse-interrupted-transitions-note-.patch (7.6K, ../../CAN4CZFMoDj7mbY231r-pmPywwJpneksUiNXDwkNokkveGtobQg@mail.gmail.com/4-v5-0002-pg_checksums-Refuse-interrupted-transitions-note-.patch)
download | inline diff:
From aff2080b0baa2cb4429c946af699839041d22234 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:19 +0000
Subject: [PATCH v5 2/4] pg_checksums: Refuse interrupted transitions, note
change is local
An inprogress-on or inprogress-off control file means an online
transition was cut short; running the offline tool on top of it mixes
two procedures. Refuse it in every mode, with a hint pointing at the
start-stop cycle (or, on a standby, at letting replication finish)
that resets the state. Also print that the change applies to one
data directory only, with standby-aware wording, since in a
replication setup the same change must be made on every node.
On a primary, inprogress-on cannot survive a graceful stop: the
checksums launcher resolves it from its exit cleanup, and a crashed
primary is already rejected by the existing "cluster must be shut
down" check. inprogress-off can, however: it is set by the backend
running pg_disable_data_checksums(), and a fast shutdown arriving
between its two barriers leaves it behind in a cleanly shut down
control file, where the next start-stop cycle resolves it at end of
recovery. A standby can be stopped with either state, since it has
no launcher and only carries forward whatever state the last replayed
WAL record left it in, with a restartpoint persisting that as-is.
The standby is also the deterministic way to reach the new guard,
which is why the test coverage uses one; the primary window would
need an injection point between the two barriers.
The pg_checksums page said the tool still processes all relation files
regardless of an interrupted online transition; describe the refusal
instead.
---
doc/src/sgml/ref/pg_checksums.sgml | 11 ++++---
src/bin/pg_checksums/pg_checksums.c | 21 +++++++++++++
.../test_checksums/t/012_offline_standby.pl | 31 ++++++++++++++-----
3 files changed, 51 insertions(+), 12 deletions(-)
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index bf07da09754..000940e7cb7 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -49,11 +49,12 @@ PostgreSQL documentation
</para>
<para>
- When enabling checksums with <application>pg_checksums</application>, if
- checksums were in the process of being enabled using
- <xref linkend="checksums-online-enable-disable"/> when the cluster was shut
- down, <application>pg_checksums</application> will still process all
- relation files regardless of the progress of online checksum processing.
+ If checksums were in the process of being enabled or disabled using
+ <xref linkend="checksums-online-enable-disable"/> when the cluster was
+ shut down, the control file still records that interrupted state, and
+ <application>pg_checksums</application> refuses to run in any mode.
+ Start the cluster and shut it down cleanly to reset the state, then
+ retry; on a standby, let replication complete the transition first.
</para>
<para>
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 20a99112838..a49b6f368b7 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -586,6 +586,21 @@ main(int argc, char *argv[])
ControlFile->state != DB_SHUTDOWNED_IN_RECOVERY)
pg_fatal("cluster must be shut down");
+ /*
+ * An inprogress state means an online transition was cut short. A
+ * standby stopped mid-transition carries either state; a cleanly shut
+ * down primary can still carry inprogress-off, which a fast shutdown
+ * during pg_disable_data_checksums() leaves behind, while inprogress-on
+ * is always resolved by the launcher's exit cleanup.
+ */
+ if (ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted");
+ pg_log_error_hint("Start and cleanly shut down the cluster once to reset the data checksum state, then retry. On a standby, let replication complete the transition first.");
+ exit(1);
+ }
+
if (ControlFile->data_checksum_version != PG_DATA_CHECKSUM_VERSION &&
mode == PG_MODE_CHECK)
pg_fatal("data checksums are not enabled in cluster");
@@ -673,6 +688,12 @@ main(int argc, char *argv[])
printf(_("Checksums enabled in cluster\n"));
else
printf(_("Checksums disabled in cluster\n"));
+
+ printf(_("This change applies to this data directory only.\n"));
+ if (ControlFile->state == DB_SHUTDOWNED_IN_RECOVERY)
+ printf(_("This node appears to be a standby; apply the same change to the primary and all other standbys.\n"));
+ else
+ printf(_("In a replication setup, apply the same change to every node while all are stopped, before restarting any of them.\n"));
}
return 0;
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
index 1655ccded94..33ee6802168 100644
--- a/src/test/modules/test_checksums/t/012_offline_standby.pl
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -94,7 +94,10 @@ test_checksum_state($standby, 'off');
# Converge the cluster: enable offline on the standby too.
$standby->stop;
-$standby->checksum_enable_offline;
+command_checks_all(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ 0, [qr/appears to be a standby/],
+ [], 'standby-role notice on offline enable');
$standby->start;
test_checksum_state($standby, 'on');
$primary->wait_for_catchup($standby);
@@ -162,11 +165,11 @@ unlike(
);
# Scenario 4: a standby stopped while replaying an interrupted online
-# transition keeps the interrupted state in its own control file. A
-# primary is never caught this way, as its checksums launcher resolves
-# inprogress-on back to off from its exit cleanup. A standby has no
-# launcher; it carries forward whatever the last replayed record left
-# it in.
+# transition keeps the interrupted state in its own control file, and
+# pg_checksums must refuse to touch it. A primary is never caught this
+# way, as its checksums launcher resolves inprogress-on back to off from
+# its exit cleanup. A standby has no launcher; it carries forward
+# whatever the last replayed record left it in.
# Block an online enable on the primary at inprogress-on with a
# blocking temp table, same trick as in 004_offline.pl.
@@ -180,13 +183,27 @@ wait_for_checksum_state($standby, 'inprogress-on');
# Stop the standby cleanly; its restartpoint persists inprogress-on to
# its own control file, since nothing on a standby resolves it away.
$standby->stop;
+
+command_fails_like(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --enable refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --check refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--disable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --disable refuses a standby stopped mid-transition');
+
$standby->start;
$bsession->quit;
wait_for_checksum_state($primary, 'on');
$primary->wait_for_catchup($standby);
wait_for_checksum_state($standby, 'on');
-is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
'10001', 'standby readable once the transition completes');
$standby->stop;
--
2.55.0
[application/octet-stream] v5-0001-Do-not-adopt-data-checksum-state-from-another-nod.patch (127.8K, ../../CAN4CZFMoDj7mbY231r-pmPywwJpneksUiNXDwkNokkveGtobQg@mail.gmail.com/5-v5-0001-Do-not-adopt-data-checksum-state-from-another-nod.patch)
download | inline diff:
From 4798e5d0ccd6ce7b4632f19529fa7f0024de848a Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:02 +0000
Subject: [PATCH v5 1/4] Do not adopt data checksum state from another node
during replay
Offline pg_checksums changes are local to one node, but replay adopted
the data checksum state carried by checkpoint records unconditionally.
After an offline enable on the primary, a standby whose pages were
never rewritten started verifying checksums it does not have; after an
offline change on a standby, the next replayed checkpoint silently
reverted it.
Instead, cross-check the replayed state against the local one and warn
once per divergent value. When the states match again, say so in the
log and re-arm the warning.
The control file now tracks this node's state alone, which makes the
moment it is written out part of the design: it may only claim "on"
once every page on disk carries a checksum, or a crash-restart would
verify pages that were flushed before the transition rewrote them.
XLOG2_CHECKSUMS replay therefore persists every state but "on" as soon
as it is replayed, none of them verifying anything, and leaves "on" to
the flush of the next restartpoint. A restartpoint only persists a
state its flush ran under from beginning to end, a promotion writes a
full checkpoint instead of an end-of-recovery record while the state is
still unpersisted, and a shutdown restartpoint with no new checkpoint
record to work from flushes before catching the control file up rather
than leaving it behind. Enabling checksums on the primary defers the
same write to after the checkpoint that flushes the rewritten pages:
the ones the worker found in shared buffers do not go out through its
ring buffer. A checkpoint persists the state its own redo point ran
under, under the same rule a restartpoint follows, since the checkpoint
that licenses "on" also moves the redo point past the record announcing
it: without that, a crash before the deferred write would resume above
the record and resolve a finished transition as interrupted. Checkpoint
records replayed below the consistency point are not cross-checked, the
persisted state being legitimately newer than what they carry.
For the redo point to be a reliable anchor, the states carried by WAL
records must match their WAL order. Transitions insert their record
and publish the new state under the new DataChecksumTransitionLock,
which a checkpoint holds across sampling the state and inserting its
XLOG_CHECKPOINT_REDO record; without that, a redo record could follow
a transition record in WAL while still carrying the pre-transition
state, and recovery resuming there would resolve the finished
transition as interrupted all the same.
pg_control also gains a watermark, the end LSN of the newest
XLOG2_CHECKSUMS record the node has written or applied, persisted
together with the state it produced, and replay skips records at or
below it. Without it, a standby that replayed a transition and stopped
cleanly before any restartpoint moved past the record would re-apply it
on the next startup, overriding a pg_checksums change made while it was
down; the offline change writes no WAL, so nothing would restore it. A
flag next to the watermark marks a state last written by pg_checksums
as local to this node, and recovery never adopts a checkpoint-borne
state over a local one, nor over a control file whose watermark already
covers the starting checkpoint. The watermark also stands in for the
state comparisons around the flushes above: record positions are
unique, so a full round trip back to the sampled state cannot alias.
A base backup copies the control file at an arbitrary moment, so its
state can be newer than the redo point replay starts from. Adopt the
state carried by the starting checkpoint record under backup-label
recovery; for a shutdown checkpoint, which replay does not see, take
it in StartupXLOG. Backups taken from a standby are the exception:
their starting checkpoint was written by the upstream primary, so they
keep the state of the control file that was copied with them.
pg_rewind keeps the target's own state, watermark and flag in the
control file it installs, since most of the data directory remains the
target's; replay from the last common checkpoint still adopts a state
its watermark does not cover, and applies any online transition the
target has not seen.
Bump PG_CONTROL_VERSION.
The documented procedure for offline changes in a replication setup
becomes the lockstep one: stop all nodes, run pg_checksums on each of
them, then restart. Tests cover the lockstep procedure, divergent
offline changes on either node and down a cascading chain,
crash-restarts and promotions around online transitions, checkpoints
racing an online enable and the transition record itself, a base backup
taken during an online enable, an offline disable on a standby
surviving the re-replay of the enable that preceded it, and pg_rewind
across an online enable as well as across offline enables on both nodes
after a divergence.
---
doc/src/sgml/ref/pg_checksums.sgml | 32 +-
doc/src/sgml/wal.sgml | 8 +
src/backend/access/transam/xlog.c | 549 ++++++++++++++++--
.../utils/activity/wait_event_names.txt | 1 +
src/bin/pg_checksums/pg_checksums.c | 10 +
src/bin/pg_controldata/pg_controldata.c | 4 +
src/bin/pg_resetwal/pg_resetwal.c | 7 +
src/bin/pg_rewind/pg_rewind.c | 14 +
src/include/catalog/pg_control.h | 21 +-
src/include/storage/lwlocklist.h | 1 +
src/test/modules/test_checksums/Makefile | 2 +-
src/test/modules/test_checksums/meson.build | 17 +
.../test_checksums/t/012_offline_standby.pl | 194 +++++++
.../modules/test_checksums/t/013_rewind.pl | 201 +++++++
.../modules/test_checksums/t/014_lockstep.pl | 89 +++
.../t/015_backup_online_enable.pl | 71 +++
.../t/016_backup_from_standby.pl | 100 ++++
.../t/017_standby_crash_after_disable.pl | 125 ++++
.../t/018_promote_enable_crash.pl | 134 +++++
.../test_checksums/t/019_restartpoint_race.pl | 138 +++++
.../t/020_primary_enable_crash.pl | 118 ++++
.../t/021_standby_shutdown_catchup.pl | 93 +++
.../t/022_resident_enable_crash.pl | 138 +++++
.../t/023_concurrent_checkpoint_enable.pl | 114 ++++
.../t/024_enable_crash_after_checkpoint.pl | 101 ++++
.../t/025_cascade_divergence.pl | 115 ++++
.../t/029_checkpoint_transition_race.pl | 126 ++++
.../t/030_offline_survives_rereplay.pl | 100 ++++
.../t/031_rewind_offline_enable.pl | 86 +++
29 files changed, 2634 insertions(+), 75 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/012_offline_standby.pl
create mode 100644 src/test/modules/test_checksums/t/013_rewind.pl
create mode 100644 src/test/modules/test_checksums/t/014_lockstep.pl
create mode 100644 src/test/modules/test_checksums/t/015_backup_online_enable.pl
create mode 100644 src/test/modules/test_checksums/t/016_backup_from_standby.pl
create mode 100644 src/test/modules/test_checksums/t/017_standby_crash_after_disable.pl
create mode 100644 src/test/modules/test_checksums/t/018_promote_enable_crash.pl
create mode 100644 src/test/modules/test_checksums/t/019_restartpoint_race.pl
create mode 100644 src/test/modules/test_checksums/t/020_primary_enable_crash.pl
create mode 100644 src/test/modules/test_checksums/t/021_standby_shutdown_catchup.pl
create mode 100644 src/test/modules/test_checksums/t/022_resident_enable_crash.pl
create mode 100644 src/test/modules/test_checksums/t/023_concurrent_checkpoint_enable.pl
create mode 100644 src/test/modules/test_checksums/t/024_enable_crash_after_checkpoint.pl
create mode 100644 src/test/modules/test_checksums/t/025_cascade_divergence.pl
create mode 100644 src/test/modules/test_checksums/t/029_checkpoint_transition_race.pl
create mode 100644 src/test/modules/test_checksums/t/030_offline_survives_rereplay.pl
create mode 100644 src/test/modules/test_checksums/t/031_rewind_offline_enable.pl
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index ae66fad3f0f..bf07da09754 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -242,15 +242,29 @@ PostgreSQL documentation
data directory must not be started or else data loss may occur.
</para>
<para>
- When using a replication setup with tools which perform direct copies
- of relation file blocks (for example <xref linkend="app-pgrewind"/>),
- enabling or disabling checksums can lead to page corruptions in the
- shape of incorrect checksums if the operation is not done consistently
- across all nodes. When enabling or disabling checksums in a replication
- setup, it is thus recommended to stop all the clusters before switching
- them all consistently. Destroying all standbys, performing the operation
- on the primary and finally recreating the standbys from scratch is also
- safe.
+ Enabling or disabling checksums with
+ <application>pg_checksums</application> changes only the local data
+ directory; the new state does not propagate over replication. In a
+ replication setup the same change must be applied to every node: stop all
+ nodes, run <application>pg_checksums</application> on each of them, and
+ only then restart them. Tools that copy relation file blocks directly
+ between nodes, such as <xref linkend="app-pgrewind"/>, likewise require
+ both nodes to be in the same data checksum state.
+ </para>
+ <para>
+ If the change is applied inconsistently, each node keeps its own state,
+ and a standby logs a warning when the state recorded in the replayed WAL
+ differs from its own. The same warning can appear transiently while a
+ standby catches up over WAL written before a consistent change; it stops
+ once a checkpoint record carrying the new state has been replayed.
+ </para>
+ <para>
+ A standby whose data directory was never checksummed must not have
+ checksums enabled by catching up this way. Converge the cluster by
+ running <application>pg_checksums</application> on it while stopped, by
+ enabling checksums online, or by recreating it from a base backup. Note
+ that an online enable only starts from a primary whose checksums are off,
+ so if they are already enabled there, disable them online first.
</para>
<para>
If <application>pg_checksums</application> is aborted or killed while
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index ec62d17fbcc..db18a8c516e 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -317,6 +317,14 @@
verify checksums, on an offline cluster.
</para>
+ <para>
+ An offline change only affects the data directory it is run on; the
+ new state does not propagate over replication. In a replication setup
+ the same change must be applied to all nodes while all of them are
+ stopped, as described in <xref linkend="app-pgchecksums"/>. A standby
+ whose state diverges logs a warning but keeps its local setting.
+ </para>
+
</sect2>
<sect2 id="checksums-online-enable-disable" xreflabel="Online Enabling of Checksums">
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index de4c96e135f..20428ad68cf 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -556,9 +556,17 @@ typedef struct XLogCtlData
*/
XLogRecPtr lastFpwDisableRecPtr;
- /* last data_checksum_version we've seen */
+ /* current data checksum state of this node */
uint32 data_checksum_version;
+ /*
+ * In-memory copies of ControlFile->data_checksum_lsn and
+ * ControlFile->data_checksum_is_local, see there. Updated together with
+ * data_checksum_version under info_lck.
+ */
+ XLogRecPtr data_checksum_lsn;
+ bool data_checksum_is_local;
+
slock_t info_lck; /* locks shared variables shown above */
/*
@@ -690,6 +698,14 @@ static ChecksumStateType LocalDataChecksumState = 0;
*/
int data_checksums = 0;
+/*
+ * Whether replay of the next checkpoint-family record must adopt the data
+ * checksum state it carries. Set when recovery starts from a base backup,
+ * where the state at the redo point takes precedence over the control file
+ * copied with the backup later.
+ */
+static bool adoptChecksumStateFromNextCheckpoint = false;
+
/* For WALInsertLockAcquire/Release functions */
static int MyLockNo = 0;
static bool holdingAllLocks = false;
@@ -730,6 +746,8 @@ static void ValidateXLOGDirectoryStructure(void);
static void CleanupBackupHistory(void);
static void UpdateMinRecoveryPoint(XLogRecPtr lsn, bool force);
static bool PerformRecoveryXLogAction(void);
+static void CheckReplayedDataChecksumState(uint32 replayed_version);
+static void AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn);
static void InitControlFile(uint64 sysidentifier, uint32 data_checksum_version);
static void WriteControlFile(void);
static void ReadControlFile(void);
@@ -757,7 +775,7 @@ static void WALInsertLockAcquireExclusive(void);
static void WALInsertLockRelease(void);
static void WALInsertLockUpdateInsertingAt(XLogRecPtr insertingAt);
-static void XLogChecksums(uint32 new_type);
+static XLogRecPtr XLogChecksums(uint32 new_type);
/*
* Insert an XLOG record represented by an already-constructed chain of data
@@ -4774,6 +4792,7 @@ void
SetDataChecksumsOnInProgress(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
/*
* The state transition is performed in a critical section with
@@ -4782,14 +4801,12 @@ SetDataChecksumsOnInProgress(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4827,6 +4844,8 @@ void
SetDataChecksumsOn(void)
{
uint64 barrier;
+ bool persist;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4847,23 +4866,11 @@ SetDataChecksumsOn(void)
SpinLockRelease(&XLogCtl->info_lck);
INJECTION_POINT("datachecksums-enable-checksums-delay", NULL);
+ INJECTION_POINT_LOAD("datachecksums-on-before-publish");
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_VERSION);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- SpinLockRelease(&XLogCtl->info_lck);
-
- /*
- * Update the controlfile before waiting since if we have an immediate
- * shutdown while waiting we want to come back up with checksums enabled.
- */
- LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- UpdateControlFile();
- LWLockRelease(ControlFileLock);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_VERSION);
barrier = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_CHECKSUM_ON);
@@ -4873,6 +4880,39 @@ SetDataChecksumsOn(void)
INJECTION_POINT("datachecksums-on-before-checkpoint", NULL);
RequestCheckpoint(CHECKPOINT_FORCE | CHECKPOINT_WAIT | CHECKPOINT_FAST);
+
+ INJECTION_POINT("datachecksums-on-after-checkpoint", NULL);
+
+ /*
+ * Persist "on" only now that the checkpoint has flushed the pages the
+ * transition rewrote. Crash recovery initializes verification from the
+ * control file but resumes from a checkpoint that can predate the
+ * transition, so an "on" persisted earlier would have replay verify pages
+ * whose rewrite never reached disk: the pages the worker found in shared
+ * buffers are not written back by its ring buffer.
+ *
+ * The checkpoint above normally persists the state itself; this covers the
+ * case where it started before the record written above and left the field
+ * alone. Crashing before this point is safe, as replay then re-establishes
+ * "on" from the full page images of the rewrite. Skip the write if the
+ * state moved on meanwhile, since whatever moved it persists its own.
+ * Compare the watermark rather than the state: a state comparison could
+ * not tell our transition from a later round trip back to "on".
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ persist = (XLogCtl->data_checksum_lsn == recptr);
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (persist)
+ {
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+ }
+
WaitForProcSignalBarrier(barrier);
}
@@ -4893,6 +4933,7 @@ void
SetDataChecksumsOff(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4918,14 +4959,12 @@ SetDataChecksumsOff(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4956,14 +4995,12 @@ SetDataChecksumsOff(void)
/* Ensure that we don't incur a checkpoint during disabling checksums */
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -5001,6 +5038,132 @@ SetLocalDataChecksumState(uint32 data_checksum_version)
data_checksums = data_checksum_version;
}
+/*
+ * Cross-check the data checksum state carried by a replayed checkpoint record
+ * against the state of this node.
+ *
+ * The state in the record belongs to whichever node wrote the WAL, and must
+ * not be adopted: an offline state change made with pg_checksums on one node
+ * of a replication set generates no WAL, and adopting would leak it into the
+ * other nodes through replay. Backup label recovery is the one exception,
+ * see AdoptReplayedDataChecksumState(). XLOG_CHECKPOINT_ONLINE needs no call
+ * here, as an online checkpoint's state already traveled in the preceding
+ * XLOG_CHECKPOINT_REDO record.
+ *
+ * Only archive recovery can see a lasting mismatch, as only there can the WAL
+ * and the control file come from different nodes or different times. In
+ * crash recovery a mismatch means replay resumed from a restartpoint
+ * predating an already-applied XLOG2_CHECKSUMS record, and replaying forward
+ * re-establishes the same state.
+ */
+static void
+CheckReplayedDataChecksumState(uint32 replayed_version)
+{
+ /*
+ * Warn once per remote value, so a lasting mismatch does not flood the
+ * log. Matching states re-arm the warning. Backend-local state is
+ * enough: replay only runs in the startup process, and restarting it
+ * re-arms as well.
+ */
+ static uint32 last_warned_version = PG_UINT32_MAX;
+ uint32 local_version;
+
+ if (!ArchiveRecoveryRequested)
+ return;
+
+ /*
+ * Re-replayed WAL below the consistency point was already cross-checked
+ * before minRecoveryPoint was last persisted, and the persisted state can
+ * legitimately be newer than what checkpoint records there carry:
+ * XLOG2_CHECKSUMS replay persists most states ahead of the restartpoint
+ * horizon. In particular the checkpoint record recovery restarts from is
+ * such a re-replay.
+ */
+ if (!reachedConsistency)
+ return;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ local_version = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (replayed_version == local_version)
+ {
+ /*
+ * Report convergence if this process warned before. Nothing else
+ * tells the operator that running pg_checksums on the other nodes, or
+ * a rebuild, took effect. A restart in between loses the context,
+ * but it re-arms the warning too.
+ */
+ if (last_warned_version != PG_UINT32_MAX)
+ ereport(LOG,
+ errmsg("data checksum state \"%s\" of this node now agrees with the replayed WAL",
+ get_checksum_state_string(local_version)));
+
+ last_warned_version = PG_UINT32_MAX;
+ return;
+ }
+
+ /* the nodes legitimately differ while an online transition runs */
+ if (replayed_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ replayed_version == PG_DATA_CHECKSUM_INPROGRESS_OFF ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ return;
+
+ if (replayed_version == last_warned_version)
+ return;
+ last_warned_version = replayed_version;
+
+ ereport(WARNING,
+ errmsg("data checksum state \"%s\" of this node does not match the state \"%s\" in the replayed WAL",
+ get_checksum_state_string(local_version),
+ get_checksum_state_string(replayed_version)),
+ errdetail("The data checksum state was most likely changed with pg_checksums on another node."),
+ errhint("Apply the same change with pg_checksums on the primary and all standby servers, or rebuild this server from a base backup."));
+}
+
+/*
+ * Adopt the data checksum state found at the redo point of backup label
+ * recovery. Persist it immediately so that a crash before the first
+ * restartpoint does not resurrect the state copied with the backup; a crash
+ * at this point restarts from the same redo point, so the control file does
+ * not run ahead of the replay position. If the value is unchanged the
+ * control file already carries it, so both the barrier and the persist are
+ * skipped.
+ *
+ * lsn is the location of the checkpoint-family record the state was taken
+ * from and becomes the new watermark: the adopted state covers everything
+ * below the redo point, which replay never revisits.
+ */
+static void
+AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn)
+{
+ bool changed = false;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (XLogCtl->data_checksum_version != new_version)
+ {
+ XLogCtl->data_checksum_version = new_version;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(new_version);
+ changed = true;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (!changed)
+ return;
+
+ EmitAndWaitDataChecksumsBarrier(new_version);
+
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = new_version;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+}
+
/* guc hook */
const char *
show_data_checksums(void)
@@ -5454,6 +5617,8 @@ XLOGShmemInit(void *arg)
/* Use the checksum info from control file */
XLogCtl->data_checksum_version = ControlFile->data_checksum_version;
+ XLogCtl->data_checksum_lsn = ControlFile->data_checksum_lsn;
+ XLogCtl->data_checksum_is_local = ControlFile->data_checksum_is_local;
SetLocalDataChecksumState(XLogCtl->data_checksum_version);
SpinLockInit(&XLogCtl->Insert.insertpos_lck);
@@ -6031,6 +6196,53 @@ StartupXLOG(void)
SetCommitTsLimit(checkPoint.oldestCommitTsXid,
checkPoint.newestCommitTsXid);
+ /*
+ * When recovery starts from a base backup, the control file was copied at
+ * an arbitrary moment and its data checksum state may differ from the
+ * state at the redo point, which is what the WAL from there on was
+ * written under. Adopt the state of the starting checkpoint: a shutdown
+ * checkpoint is not replayed, so take it from the record read above; the
+ * redo point of an online checkpoint is its CHECKPOINT_REDO record, so
+ * let the replay of that record adopt it. Check backupStartPoint in
+ * addition to the label: on a crash restart during backup recovery the
+ * label file is already renamed away, but the start point persists until
+ * the backup end record.
+ *
+ * Not for a base backup taken from a standby, though. Its starting
+ * checkpoint is the standby's last restartpoint, a record written by the
+ * upstream primary, whose state is not the one the copied files were
+ * written under. The copied control file is right there, as a standby
+ * persists its state only at restartpoint horizons and so never claims
+ * more than what reached disk. Such backups are recognized by
+ * backupEndPoint together with backupEndRequired; backupEndPoint is only
+ * set for "BACKUP FROM: standby" labels and persists across a crash
+ * restart. pg_rewind writes a standby label as well, but no
+ * backupEndPoint, and its recovery keeps adopting: the control file it
+ * installs carries the target's own checksum state, which can lag the
+ * redo point of the last common checkpoint the same way a restartpoint
+ * horizon can.
+ *
+ * Never adopt over a state the control file's watermark or local flag
+ * marks as newer than the starting checkpoint. A pg_checksums change is
+ * local to the node and generates no WAL, so nothing in the replayed WAL
+ * could ever restore it once overwritten; and a control file whose
+ * watermark lies above the redo point already contains the effect of
+ * every transition record up to there, including the state the starting
+ * checkpoint carries.
+ */
+ if ((haveBackupLabel || XLogRecPtrIsValid(ControlFile->backupStartPoint)) &&
+ !(XLogRecPtrIsValid(ControlFile->backupEndPoint) &&
+ ControlFile->backupEndRequired) &&
+ !ControlFile->data_checksum_is_local &&
+ checkPoint.redo > ControlFile->data_checksum_lsn)
+ {
+ if (wasShutdown)
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ checkPoint.redo);
+ else
+ adoptChecksumStateFromNextCheckpoint = true;
+ }
+
/*
* Clear out any old relcache cache files. This is *necessary* if we do
* any WAL replay, since that would probably result in the cache files
@@ -6633,11 +6845,7 @@ StartupXLOG(void)
if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
ereport(WARNING,
@@ -6654,11 +6862,7 @@ StartupXLOG(void)
else if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
}
@@ -6814,6 +7018,23 @@ static bool
PerformRecoveryXLogAction(void)
{
bool promoted = false;
+ bool flushForChecksums;
+ uint32 checksum_state;
+
+ /*
+ * The end-of-recovery record persists the data checksum state without
+ * flushing the buffer pool, but the control file may only claim "on" once
+ * every page on disk carries a checksum. If replay entered that state
+ * without a restartpoint following it, the pages rewritten by the
+ * transition are still only in the buffer pool, so take the full
+ * checkpoint below instead of the lightweight record.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ flushForChecksums = (checksum_state == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile->data_checksum_version != checksum_state);
/*
* Perform a checkpoint to update all our recovery activity to disk.
@@ -6829,7 +7050,7 @@ PerformRecoveryXLogAction(void)
* fully out of recovery mode and already accepting queries.
*/
if (ArchiveRecoveryRequested && IsUnderPostmaster &&
- PromoteIsTriggered())
+ PromoteIsTriggered() && !flushForChecksums)
{
promoted = true;
@@ -7436,6 +7657,7 @@ CreateCheckPoint(int flags)
uint32 freespace;
XLogRecPtr PriorRedoPtr;
XLogRecPtr last_important_lsn;
+ XLogRecPtr checksumLsn;
VirtualTransactionId *vxids;
int nvxids;
int oldXLogAllowed = 0;
@@ -7548,11 +7770,14 @@ CreateCheckPoint(int flags)
checkPoint.wal_level = wal_level;
/*
- * Get the current data_checksum_version value from xlogctl, valid at the
- * time of the checkpoint.
+ * Get the current data_checksum_version value from xlogctl. This is
+ * final only for a shutdown checkpoint, where no concurrent transition
+ * is possible; an online checkpoint resamples it together with the redo
+ * record below.
*/
SpinLockAcquire(&XLogCtl->info_lck);
checkPoint.dataChecksumState = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
if (shutdown)
@@ -7609,10 +7834,21 @@ CreateCheckPoint(int flags)
{
xl_checkpoint_redo redo_rec;
+ /*
+ * Sample the data checksum state and insert the redo record under
+ * DataChecksumTransitionLock, so that a concurrent transition cannot
+ * insert its XLOG2_CHECKSUMS record between the sampling and the
+ * insertion below. Without this, the redo record could follow the
+ * transition record in WAL while carrying the pre-transition state,
+ * and recovery resuming here would never learn about the transition.
+ * See XLogChecksums().
+ */
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
WALInsertLockAcquire();
redo_rec.wal_level = wal_level;
SpinLockAcquire(&XLogCtl->info_lck);
redo_rec.data_checksum_version = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
WALInsertLockRelease();
@@ -7620,6 +7856,14 @@ CreateCheckPoint(int flags)
XLogBeginInsert();
XLogRegisterData(&redo_rec, sizeof(xl_checkpoint_redo));
(void) XLogInsert(RM_XLOG_ID, XLOG_CHECKPOINT_REDO);
+ LWLockRelease(DataChecksumTransitionLock);
+
+ /*
+ * The checkpoint record must carry the same state as the redo record
+ * just inserted: the sample taken before redo determination can be
+ * stale by now, and the pair would otherwise disagree.
+ */
+ checkPoint.dataChecksumState = redo_rec.data_checksum_version;
/*
* XLogInsertRecord will have updated XLogCtl->Insert.RedoRecPtr in
@@ -7823,6 +8067,40 @@ CreateCheckPoint(int flags)
ControlFile->minRecoveryPoint = InvalidXLogRecPtr;
ControlFile->minRecoveryPointTLI = 0;
+ /*
+ * Persist the data checksum state this node runs under. Only the
+ * top-level field tracks this node; ControlFile->checkPointCopy above is
+ * a historical record used to resume replay.
+ *
+ * checkPoint.dataChecksumState was sampled under
+ * DataChecksumTransitionLock together with the redo record, so it is the
+ * state in effect at the redo point. If it was "on",
+ * the XLOG2_CHECKSUMS record announcing that precedes the redo point and
+ * every page the transition rewrote was dirtied before it, so
+ * CheckPointGuts() has just written all of them out. Recording the state
+ * here is what keeps a finished transition from being resolved as
+ * interrupted when this checkpoint is the one crash recovery resumes
+ * from: replay never sees the record announcing it.
+ *
+ * Persist it only if the state did not change while the flush was in
+ * progress. If it changed in between, the pages written out straddle two
+ * states, and the newer one could claim checksums that pages already on
+ * disk do not carry; leave the field to the next checkpoint then.
+ * SetDataChecksumsOff() persists the states that are safe to enter
+ * without a flush already. Compare the watermark rather than the state:
+ * a state comparison could not tell a full round trip back to the
+ * sampled value apart from no change at all, and the flushed pages
+ * straddle the intermediate states all the same.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksumLsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checkPoint.dataChecksumState;
+ ControlFile->data_checksum_lsn = checksumLsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
/*
* Persist unloggedLSN value. It's reset on crash recovery, so this goes
* unused on non-shutdown checkpoints, but seems useful to store it always
@@ -7967,9 +8245,11 @@ CreateEndOfRecoveryRecord(void)
ControlFile->minRecoveryPoint = recptr;
ControlFile->minRecoveryPointTLI = xlrec.ThisTimeLineID;
- /* start with the latest checksum version (as of the end of recovery) */
+ /* persist the data checksum state this node ended recovery with */
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
@@ -8174,6 +8454,9 @@ CreateRestartPoint(int flags)
XLogRecPtr endptr;
XLogSegNo _logSegNo;
TimestampTz xtime;
+ uint32 checksum_state;
+ XLogRecPtr checksum_lsn;
+ bool checksum_is_local;
/* Concurrent checkpoint/restartpoint cannot happen */
Assert(!IsUnderPostmaster || MyBackendType == B_CHECKPOINTER);
@@ -8220,8 +8503,45 @@ CreateRestartPoint(int flags)
UpdateMinRecoveryPoint(InvalidXLogRecPtr, true);
if (flags & CHECKPOINT_IS_SHUTDOWN)
{
+ bool catchUpChecksums;
+
+ /*
+ * There is no new restartpoint to persist the data checksum state
+ * with, but a cleanly stopped node should not leave the control
+ * file behind the state replay reached: pg_checksums and
+ * pg_rewind read it, and an in-progress state there makes them
+ * refuse to run. Catching it up needs the same guarantee a
+ * restartpoint gives: that every page on disk carries a checksum,
+ * so flush the buffer pool first. Only a transition to
+ * "on" that no restartpoint followed can get here; the other
+ * states are already persisted by XLOG2_CHECKSUMS replay. Replay
+ * has ended by now, so the state cannot change under us.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ checksum_is_local = XLogCtl->data_checksum_is_local;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ catchUpChecksums =
+ (checksum_lsn != ControlFile->data_checksum_lsn &&
+ XLogRecPtrIsValid(lastCheckPointRecPtr));
+
+ if (catchUpChecksums)
+ {
+ MemSet(&CheckpointStats, 0, sizeof(CheckpointStats));
+ CheckpointStats.ckpt_start_t = GetCurrentTimestamp();
+ CheckPointGuts(lastCheckPoint.redo, flags);
+ }
+
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
+ if (catchUpChecksums)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = checksum_is_local;
+ }
UpdateControlFile();
LWLockRelease(ControlFileLock);
}
@@ -8262,6 +8582,17 @@ CreateRestartPoint(int flags)
/* Update the process title */
update_checkpoint_display(flags, true, false);
+ /*
+ * Note the data checksum state the flush below starts under. Replay runs
+ * concurrently and can change the state while the flush is in progress,
+ * in which case the flush covers pages written under both states; see
+ * where the state is persisted further down.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
CheckPointGuts(lastCheckPoint.redo, flags);
/*
@@ -8321,8 +8652,26 @@ CreateRestartPoint(int flags)
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
}
- /* we shall start with the latest checksum version */
- ControlFile->data_checksum_version = lastCheckPoint.dataChecksumState;
+ /*
+ * Persist the data checksum state of this node. Not the state of the
+ * replayed checkpoint: that one belongs to the node that wrote it and
+ * may differ after an offline change on either side.
+ * ControlFile->checkPointCopy above keeps the replayed value on
+ * purpose, being a historical record used to resume replay rather
+ * than a tracker of node state.
+ *
+ * Persist only if the flush above ran under one state throughout; see
+ * CreateCheckPoint() for why, including why this compares the
+ * watermark and not the state.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksum_lsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
}
@@ -8763,9 +9112,21 @@ XLogReportParameters(void)
}
/*
- * Log the new state of checksums
+ * Log and publish the new state of checksums
+ *
+ * Inserting the record and publishing the new state must be atomic with
+ * respect to a checkpoint sampling the state for its XLOG_CHECKPOINT_REDO
+ * record: without that, a checkpoint could read the old state after the
+ * record is already in WAL and insert a redo record that both precedes the
+ * transition in WAL order and carries the pre-transition state. Recovery
+ * resuming from such a redo point would never replay the transition record
+ * and resolve the finished transition as interrupted.
+ * DataChecksumTransitionLock serializes the two; see CreateCheckPoint().
+ *
+ * Returns the end LSN of the inserted record, which the caller persists
+ * together with the new state as the data checksum watermark.
*/
-static void
+static XLogRecPtr
XLogChecksums(uint32 new_type)
{
xl_checksum_state xlrec;
@@ -8773,12 +9134,28 @@ XLogChecksums(uint32 new_type)
xlrec.new_checksum_state = new_type;
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
+
XLogBeginInsert();
XLogRegisterData((char *) &xlrec, sizeof(xl_checksum_state));
recptr = XLogInsert(RM_XLOG2_ID, XLOG2_CHECKSUMS);
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, recptr);
+
+ /* only loaded by SetDataChecksumsOn(), a no-op for the other callers */
+ INJECTION_POINT_CACHED("datachecksums-on-before-publish", NULL);
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ XLogCtl->data_checksum_version = new_type;
+ XLogCtl->data_checksum_lsn = recptr;
+ XLogCtl->data_checksum_is_local = false;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockRelease(DataChecksumTransitionLock);
+
XLogFlush(recptr);
+
+ return recptr;
}
/*
@@ -8966,11 +9343,19 @@ xlog_redo(XLogReaderState *record)
/* ControlFile->checkPointCopy always tracks the latest ckpt XID */
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->checkPointCopy.nextXid = checkPoint.nextXid;
- ControlFile->data_checksum_version = checkPoint.dataChecksumState;
UpdateControlFile();
LWLockRelease(ControlFileLock);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(checkPoint.dataChecksumState);
+
/*
* We should've already switched to the new TLI before replaying this
* record.
@@ -9206,19 +9591,17 @@ xlog_redo(XLogReaderState *record)
else if (info == XLOG_CHECKPOINT_REDO)
{
xl_checkpoint_redo redo_rec;
- bool new_state = false;
memcpy(&redo_rec, XLogRecGetData(record), sizeof(xl_checkpoint_redo));
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = redo_rec.data_checksum_version;
- SetLocalDataChecksumState(redo_rec.data_checksum_version);
- if (redo_rec.data_checksum_version != ControlFile->data_checksum_version)
- new_state = true;
- SpinLockRelease(&XLogCtl->info_lck);
-
- if (new_state)
- EmitAndWaitDataChecksumsBarrier(redo_rec.data_checksum_version);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(redo_rec.data_checksum_version,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(redo_rec.data_checksum_version);
}
else if (info == XLOG_LOGICAL_DECODING_STATUS_CHANGE)
{
@@ -9280,25 +9663,63 @@ xlog2_redo(XLogReaderState *record)
{
xl_checksum_state state;
XLogRecPtr lsn = record->EndRecPtr;
+ XLogRecPtr watermark;
memcpy(&state, XLogRecGetData(record), sizeof(xl_checksum_state));
+ SpinLockAcquire(&XLogCtl->info_lck);
+ watermark = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ /*
+ * Skip records this node has already applied. The control file
+ * carries the watermark, so this holds across restarts: recovery
+ * resuming below a record whose effect the control file already
+ * contains must not re-apply it, or it would revert a state change
+ * made with pg_checksums in between, which moves the state without
+ * writing any record of its own.
+ */
+ if (lsn <= watermark)
+ return;
+
/* advertise the location before the new state becomes visible */
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, lsn);
SpinLockAcquire(&XLogCtl->info_lck);
XLogCtl->data_checksum_version = state.new_checksum_state;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(state.new_checksum_state);
SpinLockRelease(&XLogCtl->info_lck);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = state.new_checksum_state;
+
+ /*
+ * Persist the new state, except when it is "on". Only "on" verifies
+ * checksums during reads, and between the last restartpoint and this
+ * record there may be pages on disk flushed under the old state; a
+ * crash-restart initializes verification from the control file and
+ * replay reads those pages back, so the control file may only say
+ * "on" once everything written under the transition has been flushed,
+ * as restartpoints and the end of recovery do.
+ * The opposite direction cannot wait for the restartpoint: once this
+ * record is replayed, evicted pages are written without checksums,
+ * and a control file still saying "on" would fail verification on
+ * exactly those pages after a crash.
+ */
+ if (state.new_checksum_state != PG_DATA_CHECKSUM_VERSION)
+ {
+ ControlFile->data_checksum_version = state.new_checksum_state;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ }
/*
* Update minRecoveryPoint to ensure that if recovery is aborted, we
* recover back up to this point before allowing hot standby again.
- * The new state is durable in pg_control while its location is only
- * tracked in shared memory; a standby becoming consistent below this
- * record would let base backups resume checksum verification with the
+ * The change location is only tracked in shared memory and is lost
+ * over a restart; a standby becoming consistent below this record
+ * would let base backups resume checksum verification with the
* location unknown. The local copies cannot be updated as long as
* crash recovery is happening and we expect all the WAL to be
* replayed.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 256b3a3c02e..77fa543f9d8 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -371,6 +371,7 @@ WaitLSN "Waiting to read or update shared Wait-for-LSN state."
LogicalDecodingControl "Waiting to read or update logical decoding status information."
DataChecksumsWorker "Waiting for data checksums worker."
AioWorkerControl "Waiting to update AIO worker information."
+DataChecksumTransition "Waiting for a data checksum state transition to be written to WAL."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 3b3ae23f1a6..20a99112838 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -648,6 +648,16 @@ main(int argc, char *argv[])
ControlFile->data_checksum_version =
(mode == PG_MODE_ENABLE) ? PG_DATA_CHECKSUM_VERSION : PG_DATA_CHECKSUM_OFF;
+ /*
+ * Mark the state as changed locally, without a WAL record. Recovery
+ * then knows the state is newer than anything the WAL carries and
+ * does not let a replayed checkpoint overwrite it. The watermark is
+ * left alone: any XLOG2_CHECKSUMS record this node had applied stays
+ * covered, and only records above it, written after this change, take
+ * effect again.
+ */
+ ControlFile->data_checksum_is_local = true;
+
if (do_sync)
{
pg_log_info("syncing data directory");
diff --git a/src/bin/pg_controldata/pg_controldata.c b/src/bin/pg_controldata/pg_controldata.c
index 6fc87ed114d..c363ce3dbb7 100644
--- a/src/bin/pg_controldata/pg_controldata.c
+++ b/src/bin/pg_controldata/pg_controldata.c
@@ -349,6 +349,10 @@ main(int argc, char *argv[])
(ControlFile->float8ByVal ? _("by value") : _("by reference")));
printf(_("Data page checksum version: %u\n"),
ControlFile->data_checksum_version);
+ printf(_("Data checksum watermark: %X/%08X\n"),
+ LSN_FORMAT_ARGS(ControlFile->data_checksum_lsn));
+ printf(_("Data checksum state is node-local: %s\n"),
+ (ControlFile->data_checksum_is_local ? _("yes") : _("no")));
printf(_("Default char data signedness: %s\n"),
(ControlFile->default_char_signedness ? _("signed") : _("unsigned")));
printf(_("Mock authentication nonce: %s\n"),
diff --git a/src/bin/pg_resetwal/pg_resetwal.c b/src/bin/pg_resetwal/pg_resetwal.c
index 1542a56ca4b..cdfb9898066 100644
--- a/src/bin/pg_resetwal/pg_resetwal.c
+++ b/src/bin/pg_resetwal/pg_resetwal.c
@@ -923,6 +923,13 @@ RewriteControlFile(void)
ControlFile.backupEndPoint = InvalidXLogRecPtr;
ControlFile.backupEndRequired = false;
+ /*
+ * The old WAL is gone and the new position may lie below the old
+ * watermark, which would make replay ignore future checksum transition
+ * records. The state itself is kept.
+ */
+ ControlFile.data_checksum_lsn = InvalidXLogRecPtr;
+
/*
* Force the defaults for max_* settings. The values don't really matter
* as long as wal_level='minimal'; the postmaster will reset these fields
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 2e86fd158d0..936469ea5f8 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -738,6 +738,20 @@ perform_rewind(filemap_t *filemap, rewind_source *source,
ControlFile_new.minRecoveryPoint = endrec;
ControlFile_new.minRecoveryPointTLI = endtli;
ControlFile_new.state = DB_IN_ARCHIVE_RECOVERY;
+
+ /*
+ * Keep the target's own data checksum state. Most of the data directory
+ * is still the target's: only blocks it changed since the divergence
+ * were copied from the source, so the source's state says nothing about
+ * the pages that stay. Replay from the last common checkpoint applies
+ * any WAL-logged transition the target has not seen (the watermark tells
+ * them apart), which converges the rewound server to the source's state
+ * whenever the WAL carries it.
+ */
+ ControlFile_new.data_checksum_version = ControlFile_target.data_checksum_version;
+ ControlFile_new.data_checksum_lsn = ControlFile_target.data_checksum_lsn;
+ ControlFile_new.data_checksum_is_local = ControlFile_target.data_checksum_is_local;
+
if (!dry_run)
update_controlfile(datadir_target, &ControlFile_new, do_sync);
}
diff --git a/src/include/catalog/pg_control.h b/src/include/catalog/pg_control.h
index 7b5404460ec..f898447f195 100644
--- a/src/include/catalog/pg_control.h
+++ b/src/include/catalog/pg_control.h
@@ -22,7 +22,7 @@
/* Version identifier for this pg_control format */
-#define PG_CONTROL_VERSION 1903
+#define PG_CONTROL_VERSION 1904
/* Nonce key length, see below */
#define MOCK_AUTH_NONCE_LEN 32
@@ -237,6 +237,25 @@ typedef struct ControlFileData
/* Current data checksums state */
uint32 data_checksum_version;
+ /*
+ * End of the newest XLOG2_CHECKSUMS record this node has written or
+ * applied. Replay ignores XLOG2_CHECKSUMS records at or below this
+ * point: their effect is already contained in data_checksum_version, or
+ * an offline pg_checksums change made after they were first applied
+ * supersedes them. InvalidXLogRecPtr if the node has never written or
+ * applied such a record.
+ */
+ XLogRecPtr data_checksum_lsn;
+
+ /*
+ * True when data_checksum_version was last set by pg_checksums rather
+ * than by a WAL-logged transition. Such a state is local to this node
+ * and newer than anything the WAL carries, so recovery must not replace
+ * it with a state taken from a checkpoint record. Cleared by the next
+ * WAL-logged transition.
+ */
+ bool data_checksum_is_local;
+
/*
* True if the default signedness of char is "signed" on a platform where
* the cluster is initialized.
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index d7eb648bd27..8d858be9927 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -89,6 +89,7 @@ PG_LWLOCK(54, WaitLSN)
PG_LWLOCK(55, LogicalDecodingControl)
PG_LWLOCK(56, DataChecksumsWorker)
PG_LWLOCK(57, AioWorkerControl)
+PG_LWLOCK(58, DataChecksumTransition)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/test/modules/test_checksums/Makefile b/src/test/modules/test_checksums/Makefile
index 71455cd5577..80f54bf6d8c 100644
--- a/src/test/modules/test_checksums/Makefile
+++ b/src/test/modules/test_checksums/Makefile
@@ -9,7 +9,7 @@
#
#-------------------------------------------------------------------------
-EXTRA_INSTALL = src/test/modules/injection_points
+EXTRA_INSTALL = contrib/pg_buffercache src/test/modules/injection_points
export enable_injection_points
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index fb7129d796f..fb33149b2db 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -35,6 +35,23 @@ tests += {
't/009_fpi.pl',
't/010_backup_straddle.pl',
't/011_standby_straddle.pl',
+ 't/012_offline_standby.pl',
+ 't/013_rewind.pl',
+ 't/014_lockstep.pl',
+ 't/015_backup_online_enable.pl',
+ 't/016_backup_from_standby.pl',
+ 't/017_standby_crash_after_disable.pl',
+ 't/018_promote_enable_crash.pl',
+ 't/019_restartpoint_race.pl',
+ 't/020_primary_enable_crash.pl',
+ 't/021_standby_shutdown_catchup.pl',
+ 't/022_resident_enable_crash.pl',
+ 't/023_concurrent_checkpoint_enable.pl',
+ 't/024_enable_crash_after_checkpoint.pl',
+ 't/025_cascade_divergence.pl',
+ 't/029_checkpoint_transition_race.pl',
+ 't/030_offline_survives_rereplay.pl',
+ 't/031_rewind_offline_enable.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
new file mode 100644
index 00000000000..1655ccded94
--- /dev/null
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -0,0 +1,194 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Offline checksum changes with pg_checksums are local to one node. A
+# standby must neither adopt the state of the primary from replayed
+# checkpoint records, nor lose its own offline change to them.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', 'autovacuum = off');
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Scenario 1: enable offline on the primary only. The standby must
+# stay off, warn about the mismatch, and remain readable.
+$standby->stop;
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+test_checksum_state($primary, 'on');
+test_checksum_state($standby, 'off');
+
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after offline enable on the primary');
+
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+my @warnings = $log =~ /(does not match the state)/g;
+is(scalar(@warnings), 1, 'mismatch warned once per remote value');
+
+# Matching states re-arm the warning: undo the divergence on the primary,
+# then diverge again to the same value, all without restarting the standby.
+$primary->stop;
+$primary->checksum_disable_offline;
+$logstart = -s $standby->logfile;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no warning while the states match again');
+like($log, qr/now agrees with the replayed WAL/,
+ 'convergence is reported once the states match again');
+
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+test_checksum_state($standby, 'off');
+
+# The local state survives both clean and immediate restarts.
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+# Converge the cluster: enable offline on the standby too.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+$primary->wait_for_catchup($standby);
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after converging');
+
+# Scenario 2: disable offline on the standby only. The replayed
+# checkpoint records of the still-enabled primary must not override it.
+$standby->stop;
+$standby->checksum_disable_offline;
+$standby->start;
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'off');
+
+# Restartpoints must persist the local state, not the replayed copy.
+$standby->safe_psql('postgres', "CHECKPOINT;");
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable with checksums disabled locally');
+
+# Scenario 3: crash-restart right after an online transition, before the
+# next restartpoint. Replay then resumes from an older restartpoint whose
+# checkpoint records still carry the pre-transition state. Those must
+# still match the state seeded from the control file, and the transition
+# itself must be re-established by re-replaying the XLOG2_CHECKSUMS record,
+# without a spurious mismatch warning along the way.
+
+# Converge first: bring the standby back to "on" offline.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+disable_data_checksums($primary, wait => 'off');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+# Crash-restart the standby immediately, before any restartpoint has had a
+# chance to persist the new state to its control file.
+$logstart = -s $standby->logfile;
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001',
+ 'standby readable after crash-restart across an online transition');
+
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no spurious mismatch warning after crash-restart across an online transition'
+);
+
+# Scenario 4: a standby stopped while replaying an interrupted online
+# transition keeps the interrupted state in its own control file. A
+# primary is never caught this way, as its checksums launcher resolves
+# inprogress-on back to off from its exit cleanup. A standby has no
+# launcher; it carries forward whatever the last replayed record left
+# it in.
+
+# Block an online enable on the primary at inprogress-on with a
+# blocking temp table, same trick as in 004_offline.pl.
+my $bsession = $primary->background_psql('postgres');
+$bsession->query_safe('CREATE TEMPORARY TABLE tt (a integer);');
+enable_data_checksums($primary, wait => 'inprogress-on');
+
+# The standby picks up the in-progress state from the XLOG2_CHECKSUMS record.
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# Stop the standby cleanly; its restartpoint persists inprogress-on to
+# its own control file, since nothing on a standby resolves it away.
+$standby->stop;
+$standby->start;
+$bsession->quit;
+wait_for_checksum_state($primary, 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable once the transition completes');
+
+$standby->stop;
+$primary->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/013_rewind.pl b/src/test/modules/test_checksums/t/013_rewind.pl
new file mode 100644
index 00000000000..a791e24317d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/013_rewind.pl
@@ -0,0 +1,201 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test pg_rewind across an online data checksum enable.
+#
+# A clean switchover leaves the shutdown checkpoint of the old primary
+# as the last common checkpoint between the two nodes. When data
+# checksums are enabled online on the new primary before the old one is
+# rewound, pg_rewind installs the control file of the new primary, which
+# already claims checksums are fully enabled, while replay begins at the
+# shutdown checkpoint whose record still carries the old state. Replay
+# of the WAL stretch from before the enable must run with checksums off,
+# as recorded in the checkpoint, else it would verify pages which never
+# had checksums written and fail recovery.
+#
+# The new primary requests a checkpoint right after promotion, and
+# replaying its CHECKPOINT_REDO record would repair the state before
+# any interesting WAL is reached. Hold the checkpointer on the standby
+# in a restartpoint over the promotion, like in the recovery test
+# 041_checkpoint_at_promote, so that the post-promotion writes end up
+# in WAL before the first checkpoint of the new timeline.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Old primary. full_page_writes is off so that the updates done on the
+# promoted node do not carry full page images, forcing replay on the
+# rewound node to read the pages from disk. wal_log_hints is required
+# by pg_rewind on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+full_page_writes = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+log_checkpoints = on
+]);
+$node_a->start;
+
+if (!$node_a->check_extension('injection_points'))
+{
+ plan skip_all => 'Extension injection_points not installed';
+}
+
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+$node_a->safe_psql('postgres', "CREATE EXTENSION injection_points;");
+
+# Set the hint bits on t before taking the backup, so that reads on the
+# promoted node do not emit full page images for its pages later.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->safe_psql('postgres', "SELECT count(*) FROM t;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_a, 'off');
+test_checksum_state($node_b, 'off');
+
+# Hold the next restartpoint on the standby.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point', 'wait');");
+
+# Give the restartpoint a checkpoint record to work on, then start it
+# in a background session; it will block on the injection point with
+# the checkpointer busy until released.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+my $bg_psql = $node_b->background_psql('postgres', on_error_stop => 0);
+$bg_psql->query_until(
+ qr/starting_restartpoint/, q(
+ \echo starting_restartpoint
+ CHECKPOINT;
+));
+$node_b->wait_for_event('checkpointer', 'create-restart-point');
+
+# Clean switchover: the shutdown checkpoint of A streams to B and
+# becomes the last common checkpoint.
+$node_a->stop('fast');
+
+my ($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+$stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+my $shutdown_ckpt = $1;
+$node_b->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_ckpt'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+
+my $logstart = -s $node_b->logfile;
+$node_b->promote;
+
+# Accidental restart of the old primary, diverging its timeline.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->stop('fast');
+
+# Updates on the new primary before its first checkpoint; replay of
+# these on the rewound node has to read the pages from disk.
+$node_b->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 25 = 0;");
+ok( !$node_b->log_contains("checkpoint complete", $logstart),
+ "no checkpoint on the new timeline before the updates");
+
+# Release the checkpointer; the queued post-promotion checkpoint runs
+# after the updates.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+$bg_psql->quit;
+
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+# pg_rewind refuses to run with full_page_writes disabled on the
+# source; the updates it was disabled for are already in WAL.
+$node_b->safe_psql('postgres', "ALTER SYSTEM SET full_page_writes = on;");
+$node_b->safe_psql('postgres', "SELECT pg_reload_conf();");
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the new primary');
+
+# Replay on the rewound node must start at the shutdown checkpoint of
+# the switchover, with the control file of the new primary.
+my $backup_label = slurp_file($node_a->data_dir . '/backup_label');
+$backup_label =~ /^CHECKPOINT LOCATION: ([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from backup_label";
+is($1, $shutdown_ckpt, 'replay starts at the switchover checkpoint');
+
+($stdout, $stderr) = run_command(
+ [
+ 'pg_waldump',
+ '-p' => $node_a->data_dir . '/pg_wal',
+ '-t' => 1,
+ '-s' => $shutdown_ckpt,
+ '-n' => 1,
+ ]);
+like($stdout, qr/CHECKPOINT_SHUTDOWN/,
+ 'last common checkpoint is a shutdown checkpoint');
+
+# pg_rewind keeps the target's own checksum state in the control file it
+# writes; the online enable reaches the rewound node through WAL replay
+# below, not through the copied control file.
+($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+like(
+ $stdout,
+ qr/^Data page checksum version:\s*0$/m,
+ 'rewound node keeps its own checksum state in the control file');
+
+# Start the rewound node as a standby of the new primary. Replay runs
+# through the pre-enable WAL stretch and the online enable.
+#
+# The rewind replaced the configuration files with those of the new
+# primary, so put the port back.
+my $connstr = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_b->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/014_lockstep.pl b/src/test/modules/test_checksums/t/014_lockstep.pl
new file mode 100644
index 00000000000..89c3b9da2cc
--- /dev/null
+++ b/src/test/modules/test_checksums/t/014_lockstep.pl
@@ -0,0 +1,89 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# The lockstep procedure for offline checksum changes in a replication
+# setup: stop all nodes, run pg_checksums on all of them, restart.
+# Replay of checkpoint records written before the change must not
+# revert the state of the standby.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Stop the standby first: the WAL written after this point is replayed
+# only after the offline switch, and every checkpoint record in it
+# still carries the old state.
+$standby->stop;
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 0;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 1;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+# The lockstep procedure.
+$primary->checksum_enable_offline;
+$standby->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to off.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "off" in the replayed WAL/,
+ 0);
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep enable');
+
+# Crash the standby and replay the same stretch again.
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'on');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after crash restart');
+
+# Once a post-switch checkpoint has been replayed the states match and
+# no warning may be logged.
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no mismatch warning once the states match');
+
+# Every page the standby wrote in this window must carry a checksum.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$standby->stop;
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'checksums valid on the standby');
+
+$primary->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/015_backup_online_enable.pl b/src/test/modules/test_checksums/t/015_backup_online_enable.pl
new file mode 100644
index 00000000000..8ea6679dc63
--- /dev/null
+++ b/src/test/modules/test_checksums/t/015_backup_online_enable.pl
@@ -0,0 +1,71 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A base backup taken while an online checksum enable is running. The
+# control file copied with the backup, and the redo point of the backup
+# checkpoint, both carry the in-progress state; replay of the
+# XLOG2_CHECKSUMS records completes the transition on the new standby.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', 'autovacuum = off');
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+# Hold the enable at inprogress-on with a pre-existing temp table the
+# checksum worker has to wait out, same trick as in 004_offline.pl and
+# scenario 4 of 012_offline_standby.pl.
+my $bsession = $primary->background_psql('postgres');
+$bsession->query_safe('CREATE TEMPORARY TABLE tt (a integer);');
+
+enable_data_checksums($primary, wait => 'inprogress-on');
+
+# The backup's checkpoint is taken while the primary sits at
+# inprogress-on, so both the control file and the redo point's
+# XLOG_CHECKPOINT_REDO record carry that state.
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+
+# Backup label recovery adopts the redo point's state immediately, before
+# any further WAL is replayed.
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# Release the barrier; the transition completes on the primary and
+# replicates. Nothing before this point can legitimately disagree: the
+# standby starts out at the same inprogress-on state as the primary, so
+# there is nothing yet to warn about.
+$bsession->quit;
+wait_for_checksum_state($primary, 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after the transition completed');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no mismatch warning at any point during a backup taken mid-transition');
+
+# No extra CHECKPOINT needed: the transition's completion checkpoint wrote
+# every page with a checksum, and the shutdown restartpoint flushes the rest.
+$standby->stop;
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'checksums valid on the standby');
+
+$primary->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/016_backup_from_standby.pl b/src/test/modules/test_checksums/t/016_backup_from_standby.pl
new file mode 100644
index 00000000000..3d01259e424
--- /dev/null
+++ b/src/test/modules/test_checksums/t/016_backup_from_standby.pl
@@ -0,0 +1,100 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that a base backup taken from a standby keeps the standby's data
+# checksum state.
+#
+# A backup taken on a standby uses the last restartpoint as its starting
+# checkpoint (do_pg_backup_start()), so the record at the redo point was
+# written by the upstream primary and carries the primary's state. Recovery
+# from such a backup must not adopt that state: the files were copied from
+# the standby, and with the standby diverged to "off" under an "on" primary
+# the new node would come up verifying checksums its files do not have.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', 'autovacuum = off');
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Offline enable on the primary only. Per 012_offline_standby.pl this is a
+# divergence the standby is expected to survive: it keeps its own "off"
+# state, warns once, and stays readable.
+$standby->stop;
+$primary->stop;
+system_or_bail('pg_checksums', '--enable', '--pgdata', $primary->data_dir);
+$primary->start;
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'on');
+test_checksum_state($standby, 'off');
+
+# Make sure the standby has written pages under its own "off" state, so its
+# files really do lack checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,50000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', "SELECT count(*) FROM t2;");
+
+# Now take a base backup *from the standby* and bring the copy up.
+$standby->backup('from_standby');
+my $newnode = PostgreSQL::Test::Cluster->new('newnode');
+$newnode->init_from_backup($standby, 'from_standby', has_streaming => 1);
+$newnode->append_conf('postgresql.conf',
+ "primary_conninfo = '" . $primary->connstr . "'");
+$newnode->start;
+
+# The new node's files all came from a cluster running with checksums off.
+# Anything but "off" here means it adopted the primary's state through the
+# checkpoint record at the redo point.
+my ($rc, $stdout, $stderr) = $newnode->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'the node copied from the standby accepts connections')
+ or diag("stderr: $stderr");
+is($stdout, 'off', 'backup of an "off" standby comes up with checksums off');
+
+# And it must be able to read the pages the standby wrote without checksums.
+($rc, $stdout, $stderr) =
+ $newnode->psql('postgres', "SELECT count(*) FROM t2;");
+is($rc, 0, 'pages copied from the standby are readable')
+ or diag("stderr: $stderr");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($newnode->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the node copied from the standby');
+
+$newnode->stop('immediate');
+
+($stdout, $stderr) = run_command([ 'pg_controldata', $newnode->data_dir ]);
+my ($ctl_state) = $stdout =~ /Data page checksum version:\s+(\d+)/;
+note("newnode control file data checksum version: $ctl_state");
+
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/017_standby_crash_after_disable.pl b/src/test/modules/test_checksums/t/017_standby_crash_after_disable.pl
new file mode 100644
index 00000000000..315b27b93d4
--- /dev/null
+++ b/src/test/modules/test_checksums/t/017_standby_crash_after_disable.pl
@@ -0,0 +1,125 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that a standby crashing after a replayed online disable, but before
+# its next restartpoint, restarts cleanly.
+#
+# Pages dirtied before the XLOG2_CHECKSUMS record and evicted after it are
+# written out under the new "off" state, without checksums. The control
+# file must follow the record immediately in this direction: a crash-restart
+# initializes checksum verification from the control file, and one still
+# saying "on" would fail verification on exactly those pages while replaying
+# records older than the state change.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$primary->start;
+
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,1000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+
+# Small buffer pool so that replaying a bulk load evicts the pages we care
+# about, and no restartpoints so the control file stays where it is.
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 1MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_delay = 10000
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'on');
+
+# No full page images, so that replay has to read the pages of "t" back from
+# disk instead of overwriting them from the WAL.
+$primary->append_conf('postgresql.conf', 'full_page_writes = off');
+$primary->reload;
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" on the standby, *before* the state change. They are
+# not flushed: the standby has no restartpoint from here on.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+# Online disable. The standby replays it and switches to "off", but its
+# control file is not updated.
+$primary->safe_psql('postgres', 'SELECT pg_disable_data_checksums();');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+my ($ctl_before) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($ctl_state) = $ctl_before =~ /Data page checksum version:\s+(\d+)/;
+note(
+ "standby control file data_checksum_version after the disable: "
+ . "$ctl_state (0 = off, 1 = on)");
+
+# Force the standby to evict the dirty pages of "t" now that it is "off":
+# they are written back without checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE filler AS SELECT generate_series(1,300000) AS a;");
+$primary->wait_for_catchup($standby);
+
+# Crash the standby before it gets a chance to run a restartpoint.
+$standby->stop('immediate');
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'standby restarts after crashing between a checksum state '
+ . 'change and the next restartpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'the pages written while "off" are readable')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("standby log tail:\n" . join("\n", @lines[ -12 .. -1 ]))
+ if @lines >= 12;
+ fail('the pages written while "off" are readable');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/018_promote_enable_crash.pl b/src/test/modules/test_checksums/t/018_promote_enable_crash.pl
new file mode 100644
index 00000000000..97f84da5685
--- /dev/null
+++ b/src/test/modules/test_checksums/t/018_promote_enable_crash.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a standby promoted right after replaying an online enable, before any
+# restartpoint has flushed the rewritten pages.
+#
+# The end-of-recovery record persists the checksum state without flushing the
+# buffer pool, while the control file still points at a restartpoint older than
+# the transition. Persisting "on" there would make a crash before the
+# post-promotion checkpoint verify checksums over the pages that were flushed
+# while the state was still "off"; promotion must take the full checkpoint
+# instead.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Dirty the pages of "t" without full page images, then push them out to disk
+# on the standby while checksums are still off.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Online enable on the primary; the standby replays it into its own buffers,
+# where the rewritten pages stay dirty (no restartpoint, no bgwriter).
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+note("standby control file before the promotion: $before");
+
+# Promote. The end-of-recovery record persists the live state without any
+# buffer flush; the checkpoint requested afterwards is not immediate.
+$standby->promote;
+$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery();')
+ or die 'timed out waiting for the promotion';
+$standby->stop('immediate');
+
+($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($ckpt) = $ctl =~ /Latest checkpoint location:\s+(\S+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note(
+ "standby control file after the promotion: $after, checkpoint $ckpt, redo $redo"
+);
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'promoted node restarts after crashing before its first checkpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/019_restartpoint_race.pl b/src/test/modules/test_checksums/t/019_restartpoint_race.pl
new file mode 100644
index 00000000000..65aa834f366
--- /dev/null
+++ b/src/test/modules/test_checksums/t/019_restartpoint_race.pl
@@ -0,0 +1,138 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a restartpoint whose flush races the replay of an online enable.
+#
+# Replay keeps running while CheckPointGuts() writes out the buffer pool, so
+# the state at the end of the flush can be newer than the one the flush ran
+# under. Persisting it would claim checksums for pages the flush wrote out
+# before the transition, against a redo pointer that predates it.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# C1: the checkpoint the pending restartpoint will target.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" without full page images and push them to disk on
+# the standby while it is still "off".
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Start a restartpoint for C1 and hold it right after CheckPointGuts().
+$standby->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point','wait');");
+my $bg = $standby->background_psql('postgres');
+$bg->query_until(qr//, "\\echo restartpoint\nCHECKPOINT;\n");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint injection point';
+
+# The startup process keeps replaying while the checkpointer is held: the
+# whole online enable lands, and the rewritten pages stay dirty.
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+# Let the restartpoint finish; it persists the state it samples now.
+$standby->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint to finish';
+$standby->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+
+$standby->stop('immediate');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note("standby control file after the restartpoint: $after, redo $redo");
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started, 'standby restarts after a restartpoint that raced the enable');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/020_primary_enable_crash.pl b/src/test/modules/test_checksums/t/020_primary_enable_crash.pl
new file mode 100644
index 00000000000..196914bc40c
--- /dev/null
+++ b/src/test/modules/test_checksums/t/020_primary_enable_crash.pl
@@ -0,0 +1,118 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a primary crashing inside SetDataChecksumsOn(), just before the forced
+# checkpoint that flushes the rewritten pages, with crash recovery resuming
+# from a checkpoint older than the transition.
+#
+# The control file must still say "inprogress-on" there. Replay does not
+# adopt the state of the checkpoint record it resumes from, so an "on" written
+# before the flush would stay in effect while replay reads pages whose rewrite
+# never reached disk. Here the pages went out through the rewriting worker's
+# ring buffer, so only the control file state discriminates; see
+# 022_resident_enable_crash.pl for the case where they did not.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('primary_enable_crash');
+$node->init(no_data_checksums => 1);
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+shared_buffers = 512MB
+bgwriter_lru_maxpages = 0
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$node->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+# Establish the checkpoint that crash recovery will resume from.
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+# Dirty the pages of "t" without full page images, then push them out to disk
+# while checksums are still off.
+$node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+my $relpath =
+ $node->safe_psql('postgres', "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $node->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on t: $evicted, relpath $relpath");
+
+# Stop the enable right after the control file has been updated to "on" but
+# before the checkpoint that flushes the rewritten pages.
+$node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+$node->safe_psql('postgres', 'SELECT pg_enable_data_checksums();');
+
+$node->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'datachecksums-on-before-checkpoint';"
+) or die 'timed out waiting for the injection point';
+
+my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+note("control file data_checksum_version before the crash: $ctl_state");
+is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+$node->stop('immediate');
+
+# Show the on-disk checksum field of the first page of "t".
+my $page;
+open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $node->start(fail_ok => 1);
+ok($started, 'primary restarts after crashing inside the online enable');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $node->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/021_standby_shutdown_catchup.pl b/src/test/modules/test_checksums/t/021_standby_shutdown_catchup.pl
new file mode 100644
index 00000000000..699dae4542a
--- /dev/null
+++ b/src/test/modules/test_checksums/t/021_standby_shutdown_catchup.pl
@@ -0,0 +1,93 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a standby stopped after replaying an online enable, but before the
+# checkpoint record that follows it on the primary.
+#
+# The shutdown restartpoint has no new checkpoint record to work from and is
+# skipped, so the control file would keep the in-progress state the transition
+# passed through, even though replay left the node at "on" and rewrote every
+# page. pg_checksums reads that field and refuses to run on an in-progress
+# state, so the shutdown has to flush and catch it up instead.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Anchor the standby's last restartpoint here, so that nothing replayed from
+# now on gives the shutdown restartpoint a newer checkpoint record to use.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Hold the enable right after the state change record, before the checkpoint
+# it requests once the transition is complete.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+enable_data_checksums($primary);
+$primary->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'datachecksums-on-before-checkpoint';"
+) or die 'timed out waiting for the injection point';
+wait_for_checksum_state($primary, 'on');
+
+# Flush the state change record out to the standby without writing a
+# checkpoint record of any kind.
+$primary->safe_psql('postgres', 'CREATE TABLE flush_marker (a int);');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+$standby->stop;
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($state, '1',
+ 'shutdown catches the control file up with the replayed state');
+
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'pg_checksums verifies the stopped standby');
+
+$primary->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-checkpoint');");
+$primary->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-checkpoint');");
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/022_resident_enable_crash.pl b/src/test/modules/test_checksums/t/022_resident_enable_crash.pl
new file mode 100644
index 00000000000..95f230ba39e
--- /dev/null
+++ b/src/test/modules/test_checksums/t/022_resident_enable_crash.pl
@@ -0,0 +1,138 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a primary crashing inside SetDataChecksumsOn(), just before the forced
+# checkpoint that flushes the rewritten pages, when those pages are resident
+# in shared buffers.
+#
+# The rewriting worker reads through a BAS_VACUUM ring, which writes the pages
+# back as the ring recycles, but a page already resident in shared buffers is
+# not read through the ring: ReadBufferExtended() hands back the existing
+# buffer, and it stays dirty until a checkpoint. The control file may
+# therefore not say "on" before that checkpoint has run, or crash recovery
+# would resume from a checkpoint older than the transition with verification
+# already enabled, and replay records that read those still-unchecksummed
+# pages. full_page_writes is off so that the records do not simply overwrite
+# the pages with a full page image.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('resident_enable_crash');
+$node->init(no_data_checksums => 1);
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+wal_level = replica
+full_page_writes = off
+wal_log_hints = off
+shared_buffers = 512MB
+bgwriter_lru_maxpages = 0
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$node->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,100000) AS a;');
+my $relpath =
+ $node->safe_psql('postgres', "SELECT pg_relation_filepath('t'::regclass);");
+
+# Establish the checkpoint that crash recovery will resume from, with the
+# pages of "t" written out while checksums are still off.
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+# Dirty those pages again without emitting full page images, and leave them in
+# shared buffers. Replay of these records has to read the pages from disk.
+$node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+
+my $dirty_before = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+cmp_ok($dirty_before, '>', 0,
+ 'pages of t are resident and dirty before enabling checksums');
+
+# Hold the enabling right before the checkpoint that flushes the rewritten
+# pages.
+$node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+
+enable_data_checksums($node);
+$node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+note("control file data_checksum_version before the crash: $ctl_state");
+is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+# The rewritten pages must still be sitting dirty in shared buffers, or the
+# window this test is about does not exist.
+my $dirty_after = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+cmp_ok($dirty_after, '>', 0,
+ 'rewritten pages of t are still dirty before the checkpoint');
+
+$node->stop('immediate');
+
+# The on-disk copy is the one written before the transition, without a
+# checksum.
+my $page;
+open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+is($pd_checksum, 0, 'block 0 of t on disk carries no checksum');
+
+my $started = $node->start(fail_ok => 1);
+ok($started, 'primary restarts after crashing inside the online enable');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $node->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/023_concurrent_checkpoint_enable.pl b/src/test/modules/test_checksums/t/023_concurrent_checkpoint_enable.pl
new file mode 100644
index 00000000000..2d1b78fc5cd
--- /dev/null
+++ b/src/test/modules/test_checksums/t/023_concurrent_checkpoint_enable.pl
@@ -0,0 +1,114 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A checkpoint that runs between the XLOG2_CHECKSUMS("on") record and the
+# control file write at the end of SetDataChecksumsOn() must not make a crash
+# throw the completed transition away.
+#
+# SetDataChecksumsOn() writes the record, flips shared memory to "on", emits
+# the barrier and only then requests the checkpoint that flushes the rewritten
+# pages; the control file is written after that checkpoint returns. A crash in
+# that window is harmless only as long as recovery still starts before the
+# record. Any checkpoint completing in the window moves the redo point past
+# the record, so recovery would never see it, would come up with the control
+# file's "inprogress-on" and StartupXLOG() would demote that to "off", even
+# though every page on disk carries a checksum by then.
+#
+# CreateCheckPoint() therefore persists the state the checkpoint ran under, the
+# same way CreateRestartPoint() does. The window is naturally reachable:
+# checkpoint_timeout, max_wal_size, an explicit CHECKPOINT, pg_basebackup or
+# pg_backup_start can all fire there. Here it is made deterministic by holding
+# the launcher at the datachecksums-on-before-checkpoint injection point and
+# checkpointing from another session.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('concurrent_checkpoint');
+$node->init(no_data_checksums => 1);
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+my $relpath =
+ $node->safe_psql('postgres', "SELECT pg_relation_filepath('t'::regclass);");
+
+# Hold the launcher after the record, the shared memory flip and the barrier,
+# but before the checkpoint SetDataChecksumsOn() requests itself.
+$node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+
+enable_data_checksums($node);
+$node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+# Every backend already sees "on" and writes checksums.
+test_checksum_state($node, 'on');
+
+# ... while the control file still says "inprogress-on".
+my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($before, '3', 'control file says "inprogress-on" inside the window');
+
+# A concurrent checkpoint. It flushes the rewritten pages and moves the redo
+# point past the XLOG2_CHECKSUMS("on") record, so it has to record the "on"
+# state in the control file as well.
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($after, '1',
+ 'the concurrent checkpoint records "on" in the control file');
+
+$node->stop('immediate');
+
+# The checkpoint flushed the rewritten pages, so they carry a checksum on
+# disk: the transition really did complete.
+my $page;
+open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+isnt($pd_checksum, 0, 'block 0 of t on disk carries a checksum');
+
+my $log_offset = -s $node->logfile;
+$node->start;
+
+# The transition is complete on disk, so the cluster has to come back "on".
+test_checksum_state($node, 'on');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+$node->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/024_enable_crash_after_checkpoint.pl b/src/test/modules/test_checksums/t/024_enable_crash_after_checkpoint.pl
new file mode 100644
index 00000000000..e75f17394cb
--- /dev/null
+++ b/src/test/modules/test_checksums/t/024_enable_crash_after_checkpoint.pl
@@ -0,0 +1,101 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A crash between the checkpoint an online enable requests and the control
+# file write that follows it must not lose the transition.
+#
+# The last steps of SetDataChecksumsOn() are
+#
+# WAL record -> shmem -> barrier -> checkpoint -> persist
+#
+# The checkpoint is what makes "on" safe to persist, but it also moves the
+# redo point above the XLOG2_CHECKSUMS record that carries the new state.
+# Crash recovery started from that checkpoint therefore never replays the
+# record, so without the state the checkpoint itself persists, a crash in the
+# remaining window would bring the cluster back at "inprogress-on", which
+# StartupXLOG() resolves to "off", discarding a transition whose pages are all
+# on disk with a checksum.
+#
+# t/023 exercises the same window through a checkpoint requested by another
+# session. This test closes it from the other side: the transition's own
+# checkpoint is the one that moves the redo point, so persisting the state
+# from CreateCheckPoint() is what has to save it, not the write below the
+# injection point.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('crash_after_checkpoint');
+$node->init(no_data_checksums => 1);
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+# Hold the launcher after the checkpoint that licenses "on" has completed and
+# before the state reaches the control file.
+$node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-after-checkpoint','wait');"
+);
+
+enable_data_checksums($node);
+$node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-after-checkpoint');
+
+# The transition is complete as far as the running cluster is concerned.
+test_checksum_state($node, 'on');
+
+# The checkpoint has already recorded it, so the pending write below the
+# injection point has nothing left to do.
+my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+my ($version) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($version, '1',
+ 'the requested checkpoint recorded "on" in the control file');
+
+# Crash before the control file write that follows the checkpoint.
+$node->stop('immediate');
+
+$node->start;
+
+# Every page on disk carries a checksum and the checkpoint that flushed them
+# completed, so the cluster has to come back verifying them.
+test_checksum_state($node, 'on');
+
+is($node->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '10000', 'relation readable after the crash');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+# The data directory must be one the offline tools accept.
+$node->stop;
+$node->command_ok([ 'pg_checksums', '--check', '-D', $node->data_dir ],
+ 'pg_checksums accepts the data directory');
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/025_cascade_divergence.pl b/src/test/modules/test_checksums/t/025_cascade_divergence.pl
new file mode 100644
index 00000000000..b4b474fa583
--- /dev/null
+++ b/src/test/modules/test_checksums/t/025_cascade_divergence.pl
@@ -0,0 +1,115 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# An offline data checksum state change made on one node of a cascading setup
+# is reported all the way down the chain.
+#
+# CheckReplayedDataChecksumState() is reached from xlog_redo() when a
+# checkpoint record (XLOG_CHECKPOINT_SHUTDOWN or XLOG_CHECKPOINT_REDO) is
+# replayed after consistency has been reached. Those records are written by
+# the root primary only and are relayed verbatim by every intermediate
+# standby, so a cascaded standby that was changed offline learns about the
+# divergence too. An intermediate standby's own restartpoints are not
+# WAL-logged and cannot surface it.
+#
+# Note that this makes detection checkpoint-driven: an idle or read-only
+# primary surfaces the divergence no sooner than its next checkpoint, so up to
+# checkpoint_timeout may pass before the operator sees anything. Closing that
+# window would require comparing the state when a walreceiver connects.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# checkpoint_timeout is deliberately long so that the only checkpoint record in
+# this test is the one requested explicitly below.
+my $conf = qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+);
+
+my $primary = PostgreSQL::Test::Cluster->new('cascade_primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', $conf);
+$primary->start;
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,1000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('cascade_standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf('postgresql.conf', $conf);
+$standby->start;
+
+$standby->backup('backup');
+my $cascade = PostgreSQL::Test::Cluster->new('cascade_cascade');
+$cascade->init_from_backup($standby, 'backup', has_streaming => 1);
+$cascade->append_conf('postgresql.conf', $conf);
+$cascade->start;
+
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+test_checksum_state($cascade, 'off');
+
+# Change the cascaded standby offline, and only it.
+$cascade->stop;
+$cascade->command_ok([ 'pg_checksums', '--enable', '-D', $cascade->data_dir ],
+ 'pg_checksums enables checksums on the cascaded standby only');
+
+my $logstart = -s $cascade->logfile;
+$cascade->start;
+
+test_checksum_state($cascade, 'on');
+
+# Ordinary WAL traffic carries no state, so the cascaded standby keeps
+# streaming from a chain whose state is "off" without noticing.
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (1);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+is($cascade->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '1001', 'cascaded standby keeps streaming from a divergent chain');
+
+# Restartpoints on the intermediate standby are not WAL-logged, so this does
+# not surface anything on the cascaded standby either.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (2);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+my $log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+unlike(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'an intermediate restartpoint does not surface the divergence');
+
+# A checkpoint on the root primary does: the record is relayed down the whole
+# chain, so both the direct and the cascaded standby compare it against their
+# own state. wait_for_catchup() below waits for replay, not just receipt, so
+# the record has been through xlog_redo() on both nodes once it returns.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+$log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+like(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'the cascaded standby reports the divergence on the primary checkpoint');
+
+$cascade->stop;
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/029_checkpoint_transition_race.pl b/src/test/modules/test_checksums/t/029_checkpoint_transition_race.pl
new file mode 100644
index 00000000000..211aab714f7
--- /dev/null
+++ b/src/test/modules/test_checksums/t/029_checkpoint_transition_race.pl
@@ -0,0 +1,126 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A checkpoint racing SetDataChecksumsOn() between the insertion of the
+# XLOG2_CHECKSUMS("on") record and the shared memory update must not insert
+# an XLOG_CHECKPOINT_REDO record that follows the transition in WAL order
+# while still carrying "inprogress-on". Recovery resuming from such a redo
+# point never replays the preceding transition record, comes up in
+# "inprogress-on" and resolves the finished transition as interrupted.
+#
+# XLogChecksums() closes the window by inserting the record and publishing
+# the new state under DataChecksumTransitionLock, which CreateCheckPoint()
+# takes around sampling the state and inserting the redo record. Here the
+# launcher is held between the two steps at the
+# datachecksums-on-before-publish injection point, and a concurrent
+# CHECKPOINT has to block on the lock instead of completing inside the
+# window.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+use IPC::Run;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('checkpoint_transition');
+$node->init(no_data_checksums => 1);
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+# The datachecksums-on-before-publish point fires inside a critical section,
+# where the wait machinery must not allocate. Waiting once at the
+# datachecksums-enable-checksums-delay point, which the launcher runs outside
+# the critical section, initializes it; see 050_redo_segment_missing.pl for
+# the same recipe around create-checkpoint-run.
+$node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-enable-checksums-delay','wait');"
+);
+
+# Hold the launcher after the XLOG2_CHECKSUMS("on") record is in WAL but
+# before the new state is published in shared memory.
+$node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-publish','wait');"
+);
+
+enable_data_checksums($node);
+$node->wait_for_event('datachecksums launcher',
+ 'datachecksums-enable-checksums-delay');
+$node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-enable-checksums-delay');");
+$node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-enable-checksums-delay');");
+
+$node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-publish');
+
+# The record is in WAL, the published state is still the old one.
+test_checksum_state($node, 'inprogress-on');
+
+# A concurrent checkpoint. It must block on DataChecksumTransitionLock
+# before inserting its redo record rather than complete inside the window.
+my $checkpointer = IPC::Run::start(
+ [ 'psql', '-XAtq', '-d', $node->connstr('postgres'), '-c', 'CHECKPOINT;' ],
+ '<' => '/dev/null',
+ '>' => '/dev/null',
+ '2>' => '/dev/null',
+ IPC::Run::timer($PostgreSQL::Test::Utils::timeout_default));
+
+ok( $node->poll_query_until(
+ 'postgres',
+ "SELECT wait_event = 'DataChecksumTransition' "
+ . "FROM pg_stat_activity WHERE backend_type = 'checkpointer';"),
+ 'concurrent checkpoint blocks on DataChecksumTransitionLock');
+
+# Release the launcher; the checkpoint then samples the published "on" and
+# its redo record follows the transition record.
+$node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-publish');");
+$node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-publish');");
+
+$checkpointer->finish;
+
+# Crash while the launcher's own checkpoint may still be in flight. Recovery
+# resumes from the concurrent checkpoint's redo point, which now lies above
+# the transition record and carries "on".
+$node->stop('immediate');
+
+my $log_offset = -s $node->logfile;
+$node->start;
+
+# The transition completed, so the cluster has to come back "on".
+wait_for_checksum_state($node, 'on');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+$node->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/030_offline_survives_rereplay.pl b/src/test/modules/test_checksums/t/030_offline_survives_rereplay.pl
new file mode 100644
index 00000000000..ea041117572
--- /dev/null
+++ b/src/test/modules/test_checksums/t/030_offline_survives_rereplay.pl
@@ -0,0 +1,100 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# An offline pg_checksums change on a standby must survive the re-replay of
+# an older XLOG2_CHECKSUMS record. A standby that replayed the final "on"
+# record and stopped cleanly before any restartpoint moved past it resumes
+# replay below the record on the next startup; without the watermark in the
+# control file, re-applying it would silently revert an offline disable made
+# while the standby was down.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(no_data_checksums => 1, allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+wal_keep_size = 1GB
+));
+$primary->start;
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf('postgresql.conf', 'checkpoint_timeout = 1h');
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# A checkpoint record for the standby's shutdown restartpoint to build on,
+# with a redo point below the transition records written next.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+
+# Hold the launcher after the XLOG2_CHECKSUMS("on") record and its barrier,
+# but before the checkpoint that would move the redo horizon past it.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+
+enable_data_checksums($primary);
+$primary->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+# The standby replays the "on" record; its restartpoint horizon stays below.
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+$standby->stop('fast');
+
+# The clean shutdown caught the control file up to "on" while the resume
+# point stays below the record.
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($version) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($version, '1', 'standby control file says "on" after the clean stop');
+
+# Disable checksums offline while the standby is down.
+$standby->checksum_disable_offline;
+
+# Let the enable finish on the primary.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-checkpoint');");
+$primary->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-checkpoint');");
+wait_for_checksum_state($primary, 'on');
+
+# The restart re-reads the WAL below the stop position, including the "on"
+# record, which the watermark now marks as already applied.
+my $log_offset = -s $standby->logfile;
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+
+# The nodes legitimately diverged, which replayed checkpoints report.
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $log_offset);
+like(
+ $log,
+ qr/data checksum state "off" of this node does not match the state "on" in the replayed WAL/,
+ 'the standby reports the divergence from the primary');
+
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/031_rewind_offline_enable.pl b/src/test/modules/test_checksums/t/031_rewind_offline_enable.pl
new file mode 100644
index 00000000000..3fccfc8f4a4
--- /dev/null
+++ b/src/test/modules/test_checksums/t/031_rewind_offline_enable.pl
@@ -0,0 +1,86 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind across offline checksum enables on both nodes. The last common
+# checkpoint still carries "off"; the rewound server must not adopt that
+# over the "on" both sides were moved to with pg_checksums, since no record
+# in the replayed WAL could ever restore it.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# wal_log_hints keeps the pair eligible for pg_rewind without checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b);
+
+# Failover to B, and divergence on A.
+$node_b->promote;
+$node_b->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_a->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_a->stop('fast');
+
+# The lockstep procedure across the divergence: both nodes get their
+# checksums enabled offline.
+$node_a->checksum_enable_offline;
+$node_b->stop;
+$node_b->checksum_enable_offline;
+$node_b->start;
+test_checksum_state($node_b, 'on');
+
+# The states match, so the rewind proceeds without complaint.
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled offline on both nodes');
+
+$node_a->append_conf('postgresql.conf', 'port = ' . $node_a->port);
+$node_a->enable_streaming($node_b);
+$node_a->set_standby_mode;
+
+my $log_offset = -s $node_a->logfile;
+$node_a->start;
+$node_b->wait_for_catchup($node_a);
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+
+# The offline enable survives: replay from the common checkpoint must not
+# resurrect the pre-divergence "off".
+test_checksum_state($node_a, 'on');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($node_a->logfile, $log_offset);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no divergence reported between the rewound server and its source');
+
+$node_a->stop;
+$node_b->stop;
+
+done_testing();
--
2.55.0
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-31 05:06 Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
parent: Zsolt Parragi <zsolt.parragi@percona.com>
0 siblings, 1 reply; 43+ messages in thread
From: Bertrand Drouvot @ 2026-08-31 05:06 UTC (permalink / raw)
To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: pgsql-hackers@lists.postgresql.org
Hi,
On Sat, Aug 29, 2026 at 10:36:14PM +0100, Zsolt Parragi wrote:
> > Thanks! I don't see the patch attached. Would you mind sharing it?
>
> Sorry, I forgot to attach it to the previous email.
Thanks!
=== 1
The v5-0001 commit message says:
"
The documented procedure for offline changes in a replication setup
becomes the lockstep one: stop all nodes, run pg_checksums on each of
them, then restart.
"
I did some more testing and realized that stopping both nodes is not sufficient
to prevent a mismatch in all cases.
For example, start a primary and standby with checksums off, with the standby's
latest replayed checksum transition at L0:
1. Stop the standby.
2. Enable and then disable checksums online on the primary. This writes:
L1: inprogress-on
L2: on
L3: inprogress-off
L4: off
3. Stop the primary.
4. Run pg_checksums --enable on both stopped nodes.
At this point:
primary: on, watermark L4
standby: on, watermark L0
The standby has not seen L1-L4. When it restarts, each record has an LSN greater
than L0 and is therefore applied. The final XLOG2_CHECKSUMS(off) changes the
standby back to off, while the primary remains on. We get a mismatch despite
both nodes being stopped when pg_checksums ran.
The mismatch remains silent until a later primary checkpoint carrying on is
replayed. FWIW, v1 has the same issue.
Fixing this would probably require recording additional ordering information for
offline changes, adding even more complexity to v5. Another option would be to
document that the standby must be fully caught up before both nodes are stopped
for the offline operation.
> > Do you see the control version change as a concern?
>
> Yes, it is another non-trivial change in an already complex patch,
> really close to RC1. It's also not an area where we could easily
> implement bug fixes in a minor version, if we discover something
> later.
Yeah, and I think the case above reinforces that concern.
=== 2
+ printf(_("Data checksum watermark: %X/%08X\n"),
+ LSN_FORMAT_ARGS(ControlFile->data_checksum_lsn));
+ printf(_("Data checksum state is node-local: %s\n"),
+ (ControlFile->data_checksum_is_local ? _("yes") : _("no")));
That produces pg_upgrade --check against a running source cluster with checksums
enabled to fail with:
"
old cluster does not use data checksums but the new one does
"
Matching "Data page checksum version:" specifically should fix it.
That makes me realize that we don't have tests for pg_upgrade --check against a
running cluster: I'll open a dedicated thread and submit a patch to add those
new tests.
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-31 07:16 Daniel Gustafsson <daniel@yesql.se>
parent: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 2 replies; 43+ messages in thread
From: Daniel Gustafsson @ 2026-08-31 07:16 UTC (permalink / raw)
To: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-hackers@lists.postgresql.org
> On 31 Aug 2026, at 07:06, Bertrand Drouvot <bertranddrouvot.pg@gmail.com> wrote:
>
> Hi,
>
> On Sat, Aug 29, 2026 at 10:36:14PM +0100, Zsolt Parragi wrote:
>>> Thanks! I don't see the patch attached. Would you mind sharing it?
>>
>> Sorry, I forgot to attach it to the previous email.
>
> Thanks!
>
> === 1
>
> The v5-0001 commit message says:
>
> "
> The documented procedure for offline changes in a replication setup
> becomes the lockstep one: stop all nodes, run pg_checksums on each of
> them, then restart.
> "
>
> I did some more testing and realized that stopping both nodes is not sufficient
> to prevent a mismatch in all cases.
>
> For example, start a primary and standby with checksums off, with the standby's
> latest replayed checksum transition at L0:
>
> 1. Stop the standby.
> 2. Enable and then disable checksums online on the primary. This writes:
>
> L1: inprogress-on
> L2: on
> L3: inprogress-off
> L4: off
>
> 3. Stop the primary.
> 4. Run pg_checksums --enable on both stopped nodes.
>
> At this point:
>
> primary: on, watermark L4
> standby: on, watermark L0
>
> The standby has not seen L1-L4. When it restarts, each record has an LSN greater
> than L0 and is therefore applied. The final XLOG2_CHECKSUMS(off) changes the
> standby back to off, while the primary remains on. We get a mismatch despite
> both nodes being stopped when pg_checksums ran.
>
> The mismatch remains silent until a later primary checkpoint carrying on is
> replayed. FWIW, v1 has the same issue.
>
> Fixing this would probably require recording additional ordering information for
> offline changes, adding even more complexity to v5. Another option would be to
> document that the standby must be fully caught up before both nodes are stopped
> for the offline operation.
I think we really need to think about documenting a lot of this, potentially
even to the point of saying that offline and online changes should not be mixed
as they work with completely different durability models.
The more I think about this the less excited I am about contorting the logic of
a feature which does proper WAL logging to cope with a tool that doesn't,
including misuses like creating mismatched clusters. We should probably start
to look at improving pg_checksums such that transitions are WAL logged rather
than shoehorning in such changes with a WAL logged flow. pg_checksums rewrites
the datadirectory without the postmaster given any information that any change
was made, which in itself should be a red flag. Making sure that StartupXLOG
can detect the offline change (or something along those lines) and properly log
it seems like a better starting point.
>>> Do you see the control version change as a concern?
>>
>> Yes, it is another non-trivial change in an already complex patch,
>> really close to RC1. It's also not an area where we could easily
>> implement bug fixes in a minor version, if we discover something
>> later.
>
> Yeah, and I think the case above reinforces that concern.
I don't think the above reinforces not wanting to do a pg_control change at
this point. I think it reinforces that changing datafiles without WAL logging
is a fairly slippery slope.
--
Daniel Gustafsson
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-31 08:01 Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
parent: Daniel Gustafsson <daniel@yesql.se>
1 sibling, 1 reply; 43+ messages in thread
From: Bertrand Drouvot @ 2026-08-31 08:01 UTC (permalink / raw)
To: Daniel Gustafsson <daniel@yesql.se>; +Cc: Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-hackers@lists.postgresql.org
Hi,
On Mon, Aug 31, 2026 at 09:16:52AM +0200, Daniel Gustafsson wrote:
> > On 31 Aug 2026, at 07:06, Bertrand Drouvot <bertranddrouvot.pg@gmail.com> wrote:
> >
> > Fixing this would probably require recording additional ordering information for
> > offline changes, adding even more complexity to v5. Another option would be to
> > document that the standby must be fully caught up before both nodes are stopped
> > for the offline operation.
>
> I think we really need to think about documenting a lot of this, potentially
> even to the point of saying that offline and online changes should not be mixed
Yeah, that could make sense, at least to make the limitation explicit and warn
users about these cases.
> as they work with completely different durability models.
Yeap.
> The more I think about this the less excited I am about contorting the logic of
> a feature which does proper WAL logging to cope with a tool that doesn't,
> including misuses like creating mismatched clusters. We should probably start
> to look at improving pg_checksums such that transitions are WAL logged rather
> than shoehorning in such changes with a WAL logged flow. pg_checksums rewrites
> the datadirectory without the postmaster given any information that any change
> was made, which in itself should be a red flag.
I agree.
> Making sure that StartupXLOG
> can detect the offline change (or something along those lines) and properly log
> it seems like a better starting point.
That sounds worth exploring. Would the idea be for pg_checksums to leave a
marker in pg_control which StartupXLOG would turn into a WAL logged transition?
> >>> Do you see the control version change as a concern?
> >>
> >> Yes, it is another non-trivial change in an already complex patch,
> >> really close to RC1. It's also not an area where we could easily
> >> implement bug fixes in a minor version, if we discover something
> >> later.
> >
> > Yeah, and I think the case above reinforces that concern.
>
> I don't think the above reinforces not wanting to do a pg_control change at
> this point. I think it reinforces that changing datafiles without WAL logging
> is a fairly slippery slope.
Yeah, I see your point that the underlying issue is the non WAL logged change.
My concern was not about a pg_control change in itself, but that addressing this
case might require further changes to the current pg_control design, which adds
risk this close to RC1.
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-31 08:01 Zsolt Parragi <zsolt.parragi@percona.com>
parent: Daniel Gustafsson <daniel@yesql.se>
1 sibling, 0 replies; 43+ messages in thread
From: Zsolt Parragi @ 2026-08-31 08:01 UTC (permalink / raw)
To: Daniel Gustafsson <daniel@yesql.se>; +Cc: pgsql-hackers@lists.postgresql.org, Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
On Mon, 31 Aug 2026, Bertrand Drouvot <bertranddrouvot.pg@gmail.com> wrote:
> I did some more testing and realized that stopping both nodes is not sufficient
> to prevent a mismatch in all cases.
...
> Fixing this would probably require recording additional ordering information for
> offline changes, adding even more complexity to v5. Another option would be to
> document that the standby must be fully caught up before both nodes are stopped
> for the offline operation.
This is one of the variations I mentioned earlier, that we can emit
spurious warnings in some cases, and later state that the state
restored, and print a log about that. And with that, this is one f the
main reasons why I thought we shouldn't try to make it an error in 19.
This is reachable in even simpler, not so engineered scenarios. I
added the logging of restored state especially for cases like this, so
we can later see that the warning resolved itself.
On Mon, 31 Aug 2026, Daniel Gustafsson <daniel@yesql.se> wrote:
> I don't think the above reinforces not wanting to do a pg_control change at
> this point. I think it reinforces that changing datafiles without WAL logging
> is a fairly slippery slope.
We can simply prevent this (in this case) by not printing out any
warnings until we reached the primary's LSN, the standby already
receives this information at connection time. I have a patch for it,
but it was again a bit more complex than I initially hoped for. (and
even then, we will still have a few corner-cases left with spurious
warnings, if I remember correctly even with that there would be still
an issue with chaining standbys, that needs another fix)
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-31 10:55 Zsolt Parragi <zsolt.parragi@percona.com>
parent: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 1 reply; 43+ messages in thread
From: Zsolt Parragi @ 2026-08-31 10:55 UTC (permalink / raw)
To: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; pgsql-hackers@lists.postgresql.org
> Yeah, I see your point that the underlying issue is the non WAL logged change.
> My concern was not about a pg_control change in itself, but that addressing this
> case might require further changes to the current pg_control design, which adds
> risk this close to RC1.
(When I read the emails previously I completely missed the part that
in this case it remains a mismatch. Sorry, that was before coffee)
I don't think we can get this completely right without both wal and
pg_control changes (and possibly some extra connection-time
communication between the standby and primary to provide early
reporting instead of delayed in the most common cases)
If we want to move toward the error direction in pg20 (which we
should) we will have to log that an offline change happened.
I attached v6 which for now documents the current behavior as-is, and
other than that I reorganized the test files.
Attachments:
[application/octet-stream] v6-0002-pg_checksums-Refuse-interrupted-transitions-note-.patch (7.7K, ../../CAN4CZFNqKg9Ts76r922cKWcOgtH32ocyghQa8NLB-qeCeC1RKg@mail.gmail.com/2-v6-0002-pg_checksums-Refuse-interrupted-transitions-note-.patch)
download | inline diff:
From 56b8cc10d353dc55990c169f54d2044abb6d61cd Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:19 +0000
Subject: [PATCH v6 2/5] pg_checksums: Refuse interrupted transitions, note
change is local
An inprogress-on or inprogress-off control file means an online
transition was cut short; running the offline tool on top of it mixes
two procedures. Refuse it in every mode, with a hint pointing at the
start-stop cycle (or, on a standby, at letting replication finish)
that resets the state. Also print that the change applies to one
data directory only, with standby-aware wording, since in a
replication setup the same change must be made on every node.
On a primary, inprogress-on cannot survive a graceful stop: the
checksums launcher resolves it from its exit cleanup, and a crashed
primary is already rejected by the existing "cluster must be shut
down" check. inprogress-off can, however: it is set by the backend
running pg_disable_data_checksums(), and a fast shutdown arriving
between its two barriers leaves it behind in a cleanly shut down
control file, where the next start-stop cycle resolves it at end of
recovery. A standby can be stopped with either state, since it has
no launcher and only carries forward whatever state the last replayed
WAL record left it in, with a restartpoint persisting that as-is.
The standby is also the deterministic way to reach the new guard,
which is why the test coverage uses one; the primary window would
need an injection point between the two barriers.
The pg_checksums page said the tool still processes all relation files
regardless of an interrupted online transition; describe the refusal
instead.
---
doc/src/sgml/ref/pg_checksums.sgml | 11 ++++---
src/bin/pg_checksums/pg_checksums.c | 21 +++++++++++++
.../test_checksums/t/012_offline_standby.pl | 31 ++++++++++++++-----
3 files changed, 51 insertions(+), 12 deletions(-)
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index bf07da09754..000940e7cb7 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -49,11 +49,12 @@ PostgreSQL documentation
</para>
<para>
- When enabling checksums with <application>pg_checksums</application>, if
- checksums were in the process of being enabled using
- <xref linkend="checksums-online-enable-disable"/> when the cluster was shut
- down, <application>pg_checksums</application> will still process all
- relation files regardless of the progress of online checksum processing.
+ If checksums were in the process of being enabled or disabled using
+ <xref linkend="checksums-online-enable-disable"/> when the cluster was
+ shut down, the control file still records that interrupted state, and
+ <application>pg_checksums</application> refuses to run in any mode.
+ Start the cluster and shut it down cleanly to reset the state, then
+ retry; on a standby, let replication complete the transition first.
</para>
<para>
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 20a99112838..a49b6f368b7 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -586,6 +586,21 @@ main(int argc, char *argv[])
ControlFile->state != DB_SHUTDOWNED_IN_RECOVERY)
pg_fatal("cluster must be shut down");
+ /*
+ * An inprogress state means an online transition was cut short. A
+ * standby stopped mid-transition carries either state; a cleanly shut
+ * down primary can still carry inprogress-off, which a fast shutdown
+ * during pg_disable_data_checksums() leaves behind, while inprogress-on
+ * is always resolved by the launcher's exit cleanup.
+ */
+ if (ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted");
+ pg_log_error_hint("Start and cleanly shut down the cluster once to reset the data checksum state, then retry. On a standby, let replication complete the transition first.");
+ exit(1);
+ }
+
if (ControlFile->data_checksum_version != PG_DATA_CHECKSUM_VERSION &&
mode == PG_MODE_CHECK)
pg_fatal("data checksums are not enabled in cluster");
@@ -673,6 +688,12 @@ main(int argc, char *argv[])
printf(_("Checksums enabled in cluster\n"));
else
printf(_("Checksums disabled in cluster\n"));
+
+ printf(_("This change applies to this data directory only.\n"));
+ if (ControlFile->state == DB_SHUTDOWNED_IN_RECOVERY)
+ printf(_("This node appears to be a standby; apply the same change to the primary and all other standbys.\n"));
+ else
+ printf(_("In a replication setup, apply the same change to every node while all are stopped, before restarting any of them.\n"));
}
return 0;
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
index 241a1282ae9..5efe11f993e 100644
--- a/src/test/modules/test_checksums/t/012_offline_standby.pl
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -149,7 +149,10 @@ $newnode->stop('immediate');
# Converge the cluster: enable offline on the standby too.
$standby->stop;
-$standby->checksum_enable_offline;
+command_checks_all(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ 0, [qr/appears to be a standby/],
+ [], 'standby-role notice on offline enable');
$standby->start;
test_checksum_state($standby, 'on');
$primary->wait_for_catchup($standby);
@@ -217,11 +220,11 @@ unlike(
);
# Scenario 4: a standby stopped while replaying an interrupted online
-# transition keeps the interrupted state in its own control file. A
-# primary is never caught this way, as its checksums launcher resolves
-# inprogress-on back to off from its exit cleanup. A standby has no
-# launcher; it carries forward whatever the last replayed record left
-# it in.
+# transition keeps the interrupted state in its own control file, and
+# pg_checksums must refuse to touch it. A primary is never caught this
+# way, as its checksums launcher resolves inprogress-on back to off from
+# its exit cleanup. A standby has no launcher; it carries forward
+# whatever the last replayed record left it in.
# Block an online enable on the primary at inprogress-on with a
# blocking temp table, same trick as in 004_offline.pl.
@@ -235,6 +238,20 @@ wait_for_checksum_state($standby, 'inprogress-on');
# Stop the standby cleanly; its restartpoint persists inprogress-on to
# its own control file, since nothing on a standby resolves it away.
$standby->stop;
+
+command_fails_like(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --enable refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --check refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--disable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --disable refuses a standby stopped mid-transition');
+
$standby->start;
wait_for_checksum_state($standby, 'inprogress-on');
@@ -257,7 +274,7 @@ wait_for_checksum_state($primary, 'on');
$primary->wait_for_catchup($standby);
wait_for_checksum_state($standby, 'on');
-is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
'10001', 'standby readable once the transition completes');
# Scenario 4 continued: the backup taken mid-transition must also
--
2.55.0
[application/octet-stream] v6-0004-pg_combinebackup-Refuse-mixed-data-checksum-state.patch (9.5K, ../../CAN4CZFNqKg9Ts76r922cKWcOgtH32ocyghQa8NLB-qeCeC1RKg@mail.gmail.com/3-v6-0004-pg_combinebackup-Refuse-mixed-data-checksum-state.patch)
download | inline diff:
From 0bf51fcd340af5da50ddde40fa9e9dc30c7b90fb Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Tue, 25 Aug 2026 15:06:58 +0000
Subject: [PATCH v6 4/5] pg_combinebackup: Refuse mixed data checksum states in
a backup chain
check_control_files() detected a chain whose backups were taken under
different data checksum states, warned, and proceeded. The output
directory keeps the last backup's control file, so a full backup taken
with checksums off combined with an incremental taken after an offline
enable produces a cluster whose control file says "on" while most of
its blocks carry no checksums; it starts, and then every connection
dies on the first unchecksummed catalog page. An offline enable
between two backups of a chain is all it takes, since it rewrites
every page without logging anything, so the incremental backup does
not re-ship the pages.
Turn the warning into an error, matching what pg_rewind does for the
equivalent combinations. The check stays asymmetric on purpose: when
the last backup was taken with checksums off, stale checksums from an
earlier backup are never verified, and that chain remains usable.
---
doc/src/sgml/ref/pg_combinebackup.sgml | 19 +--
src/bin/pg_combinebackup/pg_combinebackup.c | 14 +-
src/test/modules/test_checksums/meson.build | 1 +
.../t/023_combinebackup_mixed.pl | 134 ++++++++++++++++++
4 files changed, 155 insertions(+), 13 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/023_combinebackup_mixed.pl
diff --git a/doc/src/sgml/ref/pg_combinebackup.sgml b/doc/src/sgml/ref/pg_combinebackup.sgml
index 9a6d201e0b8..f5d7e177b36 100644
--- a/doc/src/sgml/ref/pg_combinebackup.sgml
+++ b/doc/src/sgml/ref/pg_combinebackup.sgml
@@ -306,18 +306,19 @@ PostgreSQL documentation
<para>
<literal>pg_combinebackup</literal> does not recompute page checksums when
- writing the output directory. Therefore, if any of the backups used for
- reconstruction were taken with checksums disabled, but the final backup was
- taken with checksums enabled, the resulting directory may contain pages
- with invalid checksums.
+ writing the output directory. It therefore refuses a chain in which some
+ of the backups used for reconstruction were taken with checksums disabled
+ but the final backup was taken with checksums enabled: the resulting
+ directory would contain pages with invalid checksums, and the cluster
+ would fail checksum verification as soon as it reads them. Take a new
+ full backup after enabling data checksums with
+ <xref linkend="app-pgchecksums"/>.
</para>
<para>
- To avoid this problem, taking a new full backup after changing the checksum
- state of the cluster using <xref linkend="app-pgchecksums"/> is
- recommended. Otherwise, you can disable and then optionally reenable
- checksums on the directory produced by <literal>pg_combinebackup</literal>
- in order to correct the problem.
+ The reverse case is accepted: when the final backup was taken with
+ checksums disabled, stale checksums copied from an earlier backup are
+ never verified.
</para>
</refsect1>
diff --git a/src/bin/pg_combinebackup/pg_combinebackup.c b/src/bin/pg_combinebackup/pg_combinebackup.c
index 86e2ee37c40..254a27b125b 100644
--- a/src/bin/pg_combinebackup/pg_combinebackup.c
+++ b/src/bin/pg_combinebackup/pg_combinebackup.c
@@ -670,13 +670,19 @@ check_control_files(int n_backups, char **backup_dirs)
pg_log_debug("system identifier is %" PRIu64, system_identifier);
/*
- * Warn the user if not all backups are in the same state with regards to
- * checksums.
+ * Reject a chain whose backups were not all in the same state with
+ * regards to checksums. The loop above only flags that when the last
+ * backup has checksums enabled: the blocks taken from the older backups
+ * have no checksums, and pg_combinebackup does not recompute them, so the
+ * combined cluster would fail verification as soon as it reads them. The
+ * other direction is harmless, since stale checksums are never verified
+ * while checksums are disabled.
*/
if (data_checksum_mismatch)
{
- pg_log_warning("only some backups have checksums enabled");
- pg_log_warning_hint("Disable, and optionally reenable, checksums on the output directory to avoid failures.");
+ pg_log_error("only some backups have checksums enabled");
+ pg_log_error_hint("Take a new full backup after changing the data checksum state with pg_checksums.");
+ exit(1);
}
return system_identifier;
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index a4f25bac7a5..8b85a7848bb 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -46,6 +46,7 @@ tests += {
't/020_cascade_divergence.pl',
't/021_rewind_state.pl',
't/022_rewind_standby_target.pl',
+ 't/023_combinebackup_mixed.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/023_combinebackup_mixed.pl b/src/test/modules/test_checksums/t/023_combinebackup_mixed.pl
new file mode 100644
index 00000000000..ce9dee9413b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/023_combinebackup_mixed.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_combinebackup refuses a chain whose backups were taken under
+# different data checksum states when the final state has them enabled.
+#
+# The output directory keeps the last backup's control file, so a full
+# backup taken with checksums off plus an incremental taken after an
+# offline enable would produce a directory whose control file says "on"
+# while most of its blocks have no checksums. An offline enable between
+# the two backups is enough to get there: it rewrites every page but logs
+# nothing, so the incremental backup does not re-ship the pages.
+#
+# The reverse order stays allowed: with the final backup taken with
+# checksums off, stale checksums from an earlier backup are never verified.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $node = PostgreSQL::Test::Cluster->new('node');
+$node->init(
+ no_data_checksums => 1,
+ has_archiving => 1,
+ allows_streaming => 1);
+$node->append_conf('postgresql.conf', 'summarize_wal = on');
+$node->append_conf('postgresql.conf', 'autovacuum = off');
+$node->start;
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+test_checksum_state($node, 'off');
+
+$node->backup('full');
+
+# Offline enable: rewrites every page, logs nothing.
+$node->stop;
+system_or_bail('pg_checksums', '--enable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'on');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full/backup_manifest',
+ ],
+ 'incremental backup with checksums on');
+
+# Combining the chain must be refused: the result would say "on" while the
+# blocks inherited from the full backup have no checksums.
+command_fails_like(
+ [
+ 'pg_combinebackup',
+ $node->backup_dir . '/full',
+ $node->backup_dir . '/incr',
+ '--output' => $node->backup_dir . '/combined',
+ ],
+ qr/only some backups have checksums enabled/,
+ 'pg_combinebackup refuses a chain crossing an offline enable');
+
+# The other direction: full backup with checksums on, offline disable, then
+# an incremental. The last backup wins, so the combined cluster comes up off.
+$node->backup('full2');
+
+$node->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr2',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full2/backup_manifest',
+ ],
+ 'incremental backup with checksums off');
+
+my $restored = PostgreSQL::Test::Cluster->new('restored');
+$restored->init_from_backup(
+ $node, 'incr2',
+ combine_with_prior => ['full2'],
+ has_restoring => 1,
+ standby => 0);
+
+$restored->start;
+
+my ($rc, $stdout, $stderr) = $restored->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'combined backup accepts connections') or diag("stderr: $stderr");
+is($stdout, 'off', 'combined backup follows the last backup state');
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t3;');
+is($rc, 0, 'blocks from the incremental backup are readable')
+ or diag("stderr: $stderr");
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t1;');
+is($rc, 0, 'blocks inherited from the full backup are readable')
+ or diag("stderr: $stderr");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($restored->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures in the combined backup');
+
+$restored->stop('immediate');
+$node->stop;
+
+done_testing();
--
2.55.0
[application/octet-stream] v6-0001-Do-not-adopt-data-checksum-state-from-another-nod.patch (118.6K, ../../CAN4CZFNqKg9Ts76r922cKWcOgtH32ocyghQa8NLB-qeCeC1RKg@mail.gmail.com/4-v6-0001-Do-not-adopt-data-checksum-state-from-another-nod.patch)
download | inline diff:
From 712fcf1ee3838fda8d81ab6aaba5cf506bb6959c Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:02 +0000
Subject: [PATCH v6 1/5] Do not adopt data checksum state from another node
during replay
Offline pg_checksums changes are local to one node, but replay adopted
the data checksum state carried by checkpoint records unconditionally.
After an offline enable on the primary, a standby whose pages were
never rewritten started verifying checksums it does not have; after an
offline change on a standby, the next replayed checkpoint silently
reverted it.
Instead, cross-check the replayed state against the local one and warn
once per divergent value. When the states match again, say so in the
log and re-arm the warning.
The control file now tracks this node's state alone, which makes the
moment it is written out part of the design: it may only claim "on"
once every page on disk carries a checksum, or a crash-restart would
verify pages that were flushed before the transition rewrote them.
XLOG2_CHECKSUMS replay therefore persists every state but "on" as soon
as it is replayed, none of them verifying anything, and leaves "on" to
the flush of the next restartpoint. A restartpoint only persists a
state its flush ran under from beginning to end, a promotion writes a
full checkpoint instead of an end-of-recovery record while the state is
still unpersisted, and a shutdown restartpoint with no new checkpoint
record to work from flushes before catching the control file up rather
than leaving it behind. Enabling checksums on the primary defers the
same write to after the checkpoint that flushes the rewritten pages:
the ones the worker found in shared buffers do not go out through its
ring buffer. A checkpoint persists the state its own redo point ran
under, under the same rule a restartpoint follows, since the checkpoint
that licenses "on" also moves the redo point past the record announcing
it: without that, a crash before the deferred write would resume above
the record and resolve a finished transition as interrupted. Checkpoint
records replayed below the consistency point are not cross-checked, the
persisted state being legitimately newer than what they carry.
For the redo point to be a reliable anchor, the states carried by WAL
records must match their WAL order. Transitions insert their record
and publish the new state under the new DataChecksumTransitionLock,
which a checkpoint holds across sampling the state and inserting its
XLOG_CHECKPOINT_REDO record; without that, a redo record could follow
a transition record in WAL while still carrying the pre-transition
state, and recovery resuming there would resolve the finished
transition as interrupted all the same.
pg_control also gains a watermark, the end LSN of the newest
XLOG2_CHECKSUMS record the node has written or applied, persisted
together with the state it produced, and replay skips records at or
below it. Without it, a standby that replayed a transition and stopped
cleanly before any restartpoint moved past the record would re-apply it
on the next startup, overriding a pg_checksums change made while it was
down; the offline change writes no WAL, so nothing would restore it. A
flag next to the watermark marks a state last written by pg_checksums
as local to this node, and recovery never adopts a checkpoint-borne
state over a local one, nor over a control file whose watermark already
covers the starting checkpoint. The watermark also stands in for the
state comparisons around the flushes above: record positions are
unique, so a full round trip back to the sampled state cannot alias.
A base backup copies the control file at an arbitrary moment, so its
state can be newer than the redo point replay starts from. Adopt the
state carried by the starting checkpoint record under backup-label
recovery; for a shutdown checkpoint, which replay does not see, take
it in StartupXLOG. Backups taken from a standby are the exception:
their starting checkpoint was written by the upstream primary, so they
keep the state of the control file that was copied with them.
pg_rewind keeps the target's own state, watermark and flag in the
control file it installs, since most of the data directory remains the
target's; replay from the last common checkpoint still adopts a state
its watermark does not cover, and applies any online transition the
target has not seen.
Bump PG_CONTROL_VERSION.
The documented procedure for offline changes in a replication setup
becomes the lockstep one: stop all nodes, run pg_checksums on each of
them, then restart. Tests cover the lockstep procedure, divergent
offline changes on either node and down a cascading chain,
crash-restarts and promotions around online transitions, checkpoints
racing an online enable and the transition record itself, a base backup
taken during an online enable, an offline disable on a standby
surviving the re-replay of the enable that preceded it, and pg_rewind
across an online enable as well as across offline enables on both nodes
after a divergence.
---
doc/src/sgml/ref/pg_checksums.sgml | 32 +-
doc/src/sgml/wal.sgml | 8 +
src/backend/access/transam/xlog.c | 549 ++++++++++++++++--
.../utils/activity/wait_event_names.txt | 1 +
src/bin/pg_checksums/pg_checksums.c | 10 +
src/bin/pg_controldata/pg_controldata.c | 4 +
src/bin/pg_resetwal/pg_resetwal.c | 7 +
src/bin/pg_rewind/pg_rewind.c | 14 +
src/include/catalog/pg_control.h | 21 +-
src/include/storage/lwlocklist.h | 1 +
src/test/modules/test_checksums/Makefile | 2 +-
src/test/modules/test_checksums/meson.build | 9 +
.../test_checksums/t/012_offline_standby.pl | 289 +++++++++
.../modules/test_checksums/t/013_rewind.pl | 201 +++++++
.../modules/test_checksums/t/014_lockstep.pl | 182 ++++++
.../t/015_standby_crash_after_disable.pl | 125 ++++
.../t/016_promote_enable_crash.pl | 134 +++++
.../test_checksums/t/017_restartpoint_race.pl | 138 +++++
.../t/018_enable_crash_windows.pl | 496 ++++++++++++++++
.../t/019_standby_shutdown_catchup.pl | 126 ++++
.../t/020_cascade_divergence.pl | 115 ++++
21 files changed, 2389 insertions(+), 75 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/012_offline_standby.pl
create mode 100644 src/test/modules/test_checksums/t/013_rewind.pl
create mode 100644 src/test/modules/test_checksums/t/014_lockstep.pl
create mode 100644 src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
create mode 100644 src/test/modules/test_checksums/t/016_promote_enable_crash.pl
create mode 100644 src/test/modules/test_checksums/t/017_restartpoint_race.pl
create mode 100644 src/test/modules/test_checksums/t/018_enable_crash_windows.pl
create mode 100644 src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
create mode 100644 src/test/modules/test_checksums/t/020_cascade_divergence.pl
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index ae66fad3f0f..bf07da09754 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -242,15 +242,29 @@ PostgreSQL documentation
data directory must not be started or else data loss may occur.
</para>
<para>
- When using a replication setup with tools which perform direct copies
- of relation file blocks (for example <xref linkend="app-pgrewind"/>),
- enabling or disabling checksums can lead to page corruptions in the
- shape of incorrect checksums if the operation is not done consistently
- across all nodes. When enabling or disabling checksums in a replication
- setup, it is thus recommended to stop all the clusters before switching
- them all consistently. Destroying all standbys, performing the operation
- on the primary and finally recreating the standbys from scratch is also
- safe.
+ Enabling or disabling checksums with
+ <application>pg_checksums</application> changes only the local data
+ directory; the new state does not propagate over replication. In a
+ replication setup the same change must be applied to every node: stop all
+ nodes, run <application>pg_checksums</application> on each of them, and
+ only then restart them. Tools that copy relation file blocks directly
+ between nodes, such as <xref linkend="app-pgrewind"/>, likewise require
+ both nodes to be in the same data checksum state.
+ </para>
+ <para>
+ If the change is applied inconsistently, each node keeps its own state,
+ and a standby logs a warning when the state recorded in the replayed WAL
+ differs from its own. The same warning can appear transiently while a
+ standby catches up over WAL written before a consistent change; it stops
+ once a checkpoint record carrying the new state has been replayed.
+ </para>
+ <para>
+ A standby whose data directory was never checksummed must not have
+ checksums enabled by catching up this way. Converge the cluster by
+ running <application>pg_checksums</application> on it while stopped, by
+ enabling checksums online, or by recreating it from a base backup. Note
+ that an online enable only starts from a primary whose checksums are off,
+ so if they are already enabled there, disable them online first.
</para>
<para>
If <application>pg_checksums</application> is aborted or killed while
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index ec62d17fbcc..db18a8c516e 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -317,6 +317,14 @@
verify checksums, on an offline cluster.
</para>
+ <para>
+ An offline change only affects the data directory it is run on; the
+ new state does not propagate over replication. In a replication setup
+ the same change must be applied to all nodes while all of them are
+ stopped, as described in <xref linkend="app-pgchecksums"/>. A standby
+ whose state diverges logs a warning but keeps its local setting.
+ </para>
+
</sect2>
<sect2 id="checksums-online-enable-disable" xreflabel="Online Enabling of Checksums">
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index de4c96e135f..20428ad68cf 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -556,9 +556,17 @@ typedef struct XLogCtlData
*/
XLogRecPtr lastFpwDisableRecPtr;
- /* last data_checksum_version we've seen */
+ /* current data checksum state of this node */
uint32 data_checksum_version;
+ /*
+ * In-memory copies of ControlFile->data_checksum_lsn and
+ * ControlFile->data_checksum_is_local, see there. Updated together with
+ * data_checksum_version under info_lck.
+ */
+ XLogRecPtr data_checksum_lsn;
+ bool data_checksum_is_local;
+
slock_t info_lck; /* locks shared variables shown above */
/*
@@ -690,6 +698,14 @@ static ChecksumStateType LocalDataChecksumState = 0;
*/
int data_checksums = 0;
+/*
+ * Whether replay of the next checkpoint-family record must adopt the data
+ * checksum state it carries. Set when recovery starts from a base backup,
+ * where the state at the redo point takes precedence over the control file
+ * copied with the backup later.
+ */
+static bool adoptChecksumStateFromNextCheckpoint = false;
+
/* For WALInsertLockAcquire/Release functions */
static int MyLockNo = 0;
static bool holdingAllLocks = false;
@@ -730,6 +746,8 @@ static void ValidateXLOGDirectoryStructure(void);
static void CleanupBackupHistory(void);
static void UpdateMinRecoveryPoint(XLogRecPtr lsn, bool force);
static bool PerformRecoveryXLogAction(void);
+static void CheckReplayedDataChecksumState(uint32 replayed_version);
+static void AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn);
static void InitControlFile(uint64 sysidentifier, uint32 data_checksum_version);
static void WriteControlFile(void);
static void ReadControlFile(void);
@@ -757,7 +775,7 @@ static void WALInsertLockAcquireExclusive(void);
static void WALInsertLockRelease(void);
static void WALInsertLockUpdateInsertingAt(XLogRecPtr insertingAt);
-static void XLogChecksums(uint32 new_type);
+static XLogRecPtr XLogChecksums(uint32 new_type);
/*
* Insert an XLOG record represented by an already-constructed chain of data
@@ -4774,6 +4792,7 @@ void
SetDataChecksumsOnInProgress(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
/*
* The state transition is performed in a critical section with
@@ -4782,14 +4801,12 @@ SetDataChecksumsOnInProgress(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4827,6 +4844,8 @@ void
SetDataChecksumsOn(void)
{
uint64 barrier;
+ bool persist;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4847,23 +4866,11 @@ SetDataChecksumsOn(void)
SpinLockRelease(&XLogCtl->info_lck);
INJECTION_POINT("datachecksums-enable-checksums-delay", NULL);
+ INJECTION_POINT_LOAD("datachecksums-on-before-publish");
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_VERSION);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- SpinLockRelease(&XLogCtl->info_lck);
-
- /*
- * Update the controlfile before waiting since if we have an immediate
- * shutdown while waiting we want to come back up with checksums enabled.
- */
- LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- UpdateControlFile();
- LWLockRelease(ControlFileLock);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_VERSION);
barrier = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_CHECKSUM_ON);
@@ -4873,6 +4880,39 @@ SetDataChecksumsOn(void)
INJECTION_POINT("datachecksums-on-before-checkpoint", NULL);
RequestCheckpoint(CHECKPOINT_FORCE | CHECKPOINT_WAIT | CHECKPOINT_FAST);
+
+ INJECTION_POINT("datachecksums-on-after-checkpoint", NULL);
+
+ /*
+ * Persist "on" only now that the checkpoint has flushed the pages the
+ * transition rewrote. Crash recovery initializes verification from the
+ * control file but resumes from a checkpoint that can predate the
+ * transition, so an "on" persisted earlier would have replay verify pages
+ * whose rewrite never reached disk: the pages the worker found in shared
+ * buffers are not written back by its ring buffer.
+ *
+ * The checkpoint above normally persists the state itself; this covers the
+ * case where it started before the record written above and left the field
+ * alone. Crashing before this point is safe, as replay then re-establishes
+ * "on" from the full page images of the rewrite. Skip the write if the
+ * state moved on meanwhile, since whatever moved it persists its own.
+ * Compare the watermark rather than the state: a state comparison could
+ * not tell our transition from a later round trip back to "on".
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ persist = (XLogCtl->data_checksum_lsn == recptr);
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (persist)
+ {
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+ }
+
WaitForProcSignalBarrier(barrier);
}
@@ -4893,6 +4933,7 @@ void
SetDataChecksumsOff(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4918,14 +4959,12 @@ SetDataChecksumsOff(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4956,14 +4995,12 @@ SetDataChecksumsOff(void)
/* Ensure that we don't incur a checkpoint during disabling checksums */
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -5001,6 +5038,132 @@ SetLocalDataChecksumState(uint32 data_checksum_version)
data_checksums = data_checksum_version;
}
+/*
+ * Cross-check the data checksum state carried by a replayed checkpoint record
+ * against the state of this node.
+ *
+ * The state in the record belongs to whichever node wrote the WAL, and must
+ * not be adopted: an offline state change made with pg_checksums on one node
+ * of a replication set generates no WAL, and adopting would leak it into the
+ * other nodes through replay. Backup label recovery is the one exception,
+ * see AdoptReplayedDataChecksumState(). XLOG_CHECKPOINT_ONLINE needs no call
+ * here, as an online checkpoint's state already traveled in the preceding
+ * XLOG_CHECKPOINT_REDO record.
+ *
+ * Only archive recovery can see a lasting mismatch, as only there can the WAL
+ * and the control file come from different nodes or different times. In
+ * crash recovery a mismatch means replay resumed from a restartpoint
+ * predating an already-applied XLOG2_CHECKSUMS record, and replaying forward
+ * re-establishes the same state.
+ */
+static void
+CheckReplayedDataChecksumState(uint32 replayed_version)
+{
+ /*
+ * Warn once per remote value, so a lasting mismatch does not flood the
+ * log. Matching states re-arm the warning. Backend-local state is
+ * enough: replay only runs in the startup process, and restarting it
+ * re-arms as well.
+ */
+ static uint32 last_warned_version = PG_UINT32_MAX;
+ uint32 local_version;
+
+ if (!ArchiveRecoveryRequested)
+ return;
+
+ /*
+ * Re-replayed WAL below the consistency point was already cross-checked
+ * before minRecoveryPoint was last persisted, and the persisted state can
+ * legitimately be newer than what checkpoint records there carry:
+ * XLOG2_CHECKSUMS replay persists most states ahead of the restartpoint
+ * horizon. In particular the checkpoint record recovery restarts from is
+ * such a re-replay.
+ */
+ if (!reachedConsistency)
+ return;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ local_version = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (replayed_version == local_version)
+ {
+ /*
+ * Report convergence if this process warned before. Nothing else
+ * tells the operator that running pg_checksums on the other nodes, or
+ * a rebuild, took effect. A restart in between loses the context,
+ * but it re-arms the warning too.
+ */
+ if (last_warned_version != PG_UINT32_MAX)
+ ereport(LOG,
+ errmsg("data checksum state \"%s\" of this node now agrees with the replayed WAL",
+ get_checksum_state_string(local_version)));
+
+ last_warned_version = PG_UINT32_MAX;
+ return;
+ }
+
+ /* the nodes legitimately differ while an online transition runs */
+ if (replayed_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ replayed_version == PG_DATA_CHECKSUM_INPROGRESS_OFF ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ return;
+
+ if (replayed_version == last_warned_version)
+ return;
+ last_warned_version = replayed_version;
+
+ ereport(WARNING,
+ errmsg("data checksum state \"%s\" of this node does not match the state \"%s\" in the replayed WAL",
+ get_checksum_state_string(local_version),
+ get_checksum_state_string(replayed_version)),
+ errdetail("The data checksum state was most likely changed with pg_checksums on another node."),
+ errhint("Apply the same change with pg_checksums on the primary and all standby servers, or rebuild this server from a base backup."));
+}
+
+/*
+ * Adopt the data checksum state found at the redo point of backup label
+ * recovery. Persist it immediately so that a crash before the first
+ * restartpoint does not resurrect the state copied with the backup; a crash
+ * at this point restarts from the same redo point, so the control file does
+ * not run ahead of the replay position. If the value is unchanged the
+ * control file already carries it, so both the barrier and the persist are
+ * skipped.
+ *
+ * lsn is the location of the checkpoint-family record the state was taken
+ * from and becomes the new watermark: the adopted state covers everything
+ * below the redo point, which replay never revisits.
+ */
+static void
+AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn)
+{
+ bool changed = false;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (XLogCtl->data_checksum_version != new_version)
+ {
+ XLogCtl->data_checksum_version = new_version;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(new_version);
+ changed = true;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (!changed)
+ return;
+
+ EmitAndWaitDataChecksumsBarrier(new_version);
+
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = new_version;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+}
+
/* guc hook */
const char *
show_data_checksums(void)
@@ -5454,6 +5617,8 @@ XLOGShmemInit(void *arg)
/* Use the checksum info from control file */
XLogCtl->data_checksum_version = ControlFile->data_checksum_version;
+ XLogCtl->data_checksum_lsn = ControlFile->data_checksum_lsn;
+ XLogCtl->data_checksum_is_local = ControlFile->data_checksum_is_local;
SetLocalDataChecksumState(XLogCtl->data_checksum_version);
SpinLockInit(&XLogCtl->Insert.insertpos_lck);
@@ -6031,6 +6196,53 @@ StartupXLOG(void)
SetCommitTsLimit(checkPoint.oldestCommitTsXid,
checkPoint.newestCommitTsXid);
+ /*
+ * When recovery starts from a base backup, the control file was copied at
+ * an arbitrary moment and its data checksum state may differ from the
+ * state at the redo point, which is what the WAL from there on was
+ * written under. Adopt the state of the starting checkpoint: a shutdown
+ * checkpoint is not replayed, so take it from the record read above; the
+ * redo point of an online checkpoint is its CHECKPOINT_REDO record, so
+ * let the replay of that record adopt it. Check backupStartPoint in
+ * addition to the label: on a crash restart during backup recovery the
+ * label file is already renamed away, but the start point persists until
+ * the backup end record.
+ *
+ * Not for a base backup taken from a standby, though. Its starting
+ * checkpoint is the standby's last restartpoint, a record written by the
+ * upstream primary, whose state is not the one the copied files were
+ * written under. The copied control file is right there, as a standby
+ * persists its state only at restartpoint horizons and so never claims
+ * more than what reached disk. Such backups are recognized by
+ * backupEndPoint together with backupEndRequired; backupEndPoint is only
+ * set for "BACKUP FROM: standby" labels and persists across a crash
+ * restart. pg_rewind writes a standby label as well, but no
+ * backupEndPoint, and its recovery keeps adopting: the control file it
+ * installs carries the target's own checksum state, which can lag the
+ * redo point of the last common checkpoint the same way a restartpoint
+ * horizon can.
+ *
+ * Never adopt over a state the control file's watermark or local flag
+ * marks as newer than the starting checkpoint. A pg_checksums change is
+ * local to the node and generates no WAL, so nothing in the replayed WAL
+ * could ever restore it once overwritten; and a control file whose
+ * watermark lies above the redo point already contains the effect of
+ * every transition record up to there, including the state the starting
+ * checkpoint carries.
+ */
+ if ((haveBackupLabel || XLogRecPtrIsValid(ControlFile->backupStartPoint)) &&
+ !(XLogRecPtrIsValid(ControlFile->backupEndPoint) &&
+ ControlFile->backupEndRequired) &&
+ !ControlFile->data_checksum_is_local &&
+ checkPoint.redo > ControlFile->data_checksum_lsn)
+ {
+ if (wasShutdown)
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ checkPoint.redo);
+ else
+ adoptChecksumStateFromNextCheckpoint = true;
+ }
+
/*
* Clear out any old relcache cache files. This is *necessary* if we do
* any WAL replay, since that would probably result in the cache files
@@ -6633,11 +6845,7 @@ StartupXLOG(void)
if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
ereport(WARNING,
@@ -6654,11 +6862,7 @@ StartupXLOG(void)
else if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
}
@@ -6814,6 +7018,23 @@ static bool
PerformRecoveryXLogAction(void)
{
bool promoted = false;
+ bool flushForChecksums;
+ uint32 checksum_state;
+
+ /*
+ * The end-of-recovery record persists the data checksum state without
+ * flushing the buffer pool, but the control file may only claim "on" once
+ * every page on disk carries a checksum. If replay entered that state
+ * without a restartpoint following it, the pages rewritten by the
+ * transition are still only in the buffer pool, so take the full
+ * checkpoint below instead of the lightweight record.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ flushForChecksums = (checksum_state == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile->data_checksum_version != checksum_state);
/*
* Perform a checkpoint to update all our recovery activity to disk.
@@ -6829,7 +7050,7 @@ PerformRecoveryXLogAction(void)
* fully out of recovery mode and already accepting queries.
*/
if (ArchiveRecoveryRequested && IsUnderPostmaster &&
- PromoteIsTriggered())
+ PromoteIsTriggered() && !flushForChecksums)
{
promoted = true;
@@ -7436,6 +7657,7 @@ CreateCheckPoint(int flags)
uint32 freespace;
XLogRecPtr PriorRedoPtr;
XLogRecPtr last_important_lsn;
+ XLogRecPtr checksumLsn;
VirtualTransactionId *vxids;
int nvxids;
int oldXLogAllowed = 0;
@@ -7548,11 +7770,14 @@ CreateCheckPoint(int flags)
checkPoint.wal_level = wal_level;
/*
- * Get the current data_checksum_version value from xlogctl, valid at the
- * time of the checkpoint.
+ * Get the current data_checksum_version value from xlogctl. This is
+ * final only for a shutdown checkpoint, where no concurrent transition
+ * is possible; an online checkpoint resamples it together with the redo
+ * record below.
*/
SpinLockAcquire(&XLogCtl->info_lck);
checkPoint.dataChecksumState = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
if (shutdown)
@@ -7609,10 +7834,21 @@ CreateCheckPoint(int flags)
{
xl_checkpoint_redo redo_rec;
+ /*
+ * Sample the data checksum state and insert the redo record under
+ * DataChecksumTransitionLock, so that a concurrent transition cannot
+ * insert its XLOG2_CHECKSUMS record between the sampling and the
+ * insertion below. Without this, the redo record could follow the
+ * transition record in WAL while carrying the pre-transition state,
+ * and recovery resuming here would never learn about the transition.
+ * See XLogChecksums().
+ */
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
WALInsertLockAcquire();
redo_rec.wal_level = wal_level;
SpinLockAcquire(&XLogCtl->info_lck);
redo_rec.data_checksum_version = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
WALInsertLockRelease();
@@ -7620,6 +7856,14 @@ CreateCheckPoint(int flags)
XLogBeginInsert();
XLogRegisterData(&redo_rec, sizeof(xl_checkpoint_redo));
(void) XLogInsert(RM_XLOG_ID, XLOG_CHECKPOINT_REDO);
+ LWLockRelease(DataChecksumTransitionLock);
+
+ /*
+ * The checkpoint record must carry the same state as the redo record
+ * just inserted: the sample taken before redo determination can be
+ * stale by now, and the pair would otherwise disagree.
+ */
+ checkPoint.dataChecksumState = redo_rec.data_checksum_version;
/*
* XLogInsertRecord will have updated XLogCtl->Insert.RedoRecPtr in
@@ -7823,6 +8067,40 @@ CreateCheckPoint(int flags)
ControlFile->minRecoveryPoint = InvalidXLogRecPtr;
ControlFile->minRecoveryPointTLI = 0;
+ /*
+ * Persist the data checksum state this node runs under. Only the
+ * top-level field tracks this node; ControlFile->checkPointCopy above is
+ * a historical record used to resume replay.
+ *
+ * checkPoint.dataChecksumState was sampled under
+ * DataChecksumTransitionLock together with the redo record, so it is the
+ * state in effect at the redo point. If it was "on",
+ * the XLOG2_CHECKSUMS record announcing that precedes the redo point and
+ * every page the transition rewrote was dirtied before it, so
+ * CheckPointGuts() has just written all of them out. Recording the state
+ * here is what keeps a finished transition from being resolved as
+ * interrupted when this checkpoint is the one crash recovery resumes
+ * from: replay never sees the record announcing it.
+ *
+ * Persist it only if the state did not change while the flush was in
+ * progress. If it changed in between, the pages written out straddle two
+ * states, and the newer one could claim checksums that pages already on
+ * disk do not carry; leave the field to the next checkpoint then.
+ * SetDataChecksumsOff() persists the states that are safe to enter
+ * without a flush already. Compare the watermark rather than the state:
+ * a state comparison could not tell a full round trip back to the
+ * sampled value apart from no change at all, and the flushed pages
+ * straddle the intermediate states all the same.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksumLsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checkPoint.dataChecksumState;
+ ControlFile->data_checksum_lsn = checksumLsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
/*
* Persist unloggedLSN value. It's reset on crash recovery, so this goes
* unused on non-shutdown checkpoints, but seems useful to store it always
@@ -7967,9 +8245,11 @@ CreateEndOfRecoveryRecord(void)
ControlFile->minRecoveryPoint = recptr;
ControlFile->minRecoveryPointTLI = xlrec.ThisTimeLineID;
- /* start with the latest checksum version (as of the end of recovery) */
+ /* persist the data checksum state this node ended recovery with */
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
@@ -8174,6 +8454,9 @@ CreateRestartPoint(int flags)
XLogRecPtr endptr;
XLogSegNo _logSegNo;
TimestampTz xtime;
+ uint32 checksum_state;
+ XLogRecPtr checksum_lsn;
+ bool checksum_is_local;
/* Concurrent checkpoint/restartpoint cannot happen */
Assert(!IsUnderPostmaster || MyBackendType == B_CHECKPOINTER);
@@ -8220,8 +8503,45 @@ CreateRestartPoint(int flags)
UpdateMinRecoveryPoint(InvalidXLogRecPtr, true);
if (flags & CHECKPOINT_IS_SHUTDOWN)
{
+ bool catchUpChecksums;
+
+ /*
+ * There is no new restartpoint to persist the data checksum state
+ * with, but a cleanly stopped node should not leave the control
+ * file behind the state replay reached: pg_checksums and
+ * pg_rewind read it, and an in-progress state there makes them
+ * refuse to run. Catching it up needs the same guarantee a
+ * restartpoint gives: that every page on disk carries a checksum,
+ * so flush the buffer pool first. Only a transition to
+ * "on" that no restartpoint followed can get here; the other
+ * states are already persisted by XLOG2_CHECKSUMS replay. Replay
+ * has ended by now, so the state cannot change under us.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ checksum_is_local = XLogCtl->data_checksum_is_local;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ catchUpChecksums =
+ (checksum_lsn != ControlFile->data_checksum_lsn &&
+ XLogRecPtrIsValid(lastCheckPointRecPtr));
+
+ if (catchUpChecksums)
+ {
+ MemSet(&CheckpointStats, 0, sizeof(CheckpointStats));
+ CheckpointStats.ckpt_start_t = GetCurrentTimestamp();
+ CheckPointGuts(lastCheckPoint.redo, flags);
+ }
+
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
+ if (catchUpChecksums)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = checksum_is_local;
+ }
UpdateControlFile();
LWLockRelease(ControlFileLock);
}
@@ -8262,6 +8582,17 @@ CreateRestartPoint(int flags)
/* Update the process title */
update_checkpoint_display(flags, true, false);
+ /*
+ * Note the data checksum state the flush below starts under. Replay runs
+ * concurrently and can change the state while the flush is in progress,
+ * in which case the flush covers pages written under both states; see
+ * where the state is persisted further down.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
CheckPointGuts(lastCheckPoint.redo, flags);
/*
@@ -8321,8 +8652,26 @@ CreateRestartPoint(int flags)
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
}
- /* we shall start with the latest checksum version */
- ControlFile->data_checksum_version = lastCheckPoint.dataChecksumState;
+ /*
+ * Persist the data checksum state of this node. Not the state of the
+ * replayed checkpoint: that one belongs to the node that wrote it and
+ * may differ after an offline change on either side.
+ * ControlFile->checkPointCopy above keeps the replayed value on
+ * purpose, being a historical record used to resume replay rather
+ * than a tracker of node state.
+ *
+ * Persist only if the flush above ran under one state throughout; see
+ * CreateCheckPoint() for why, including why this compares the
+ * watermark and not the state.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksum_lsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
}
@@ -8763,9 +9112,21 @@ XLogReportParameters(void)
}
/*
- * Log the new state of checksums
+ * Log and publish the new state of checksums
+ *
+ * Inserting the record and publishing the new state must be atomic with
+ * respect to a checkpoint sampling the state for its XLOG_CHECKPOINT_REDO
+ * record: without that, a checkpoint could read the old state after the
+ * record is already in WAL and insert a redo record that both precedes the
+ * transition in WAL order and carries the pre-transition state. Recovery
+ * resuming from such a redo point would never replay the transition record
+ * and resolve the finished transition as interrupted.
+ * DataChecksumTransitionLock serializes the two; see CreateCheckPoint().
+ *
+ * Returns the end LSN of the inserted record, which the caller persists
+ * together with the new state as the data checksum watermark.
*/
-static void
+static XLogRecPtr
XLogChecksums(uint32 new_type)
{
xl_checksum_state xlrec;
@@ -8773,12 +9134,28 @@ XLogChecksums(uint32 new_type)
xlrec.new_checksum_state = new_type;
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
+
XLogBeginInsert();
XLogRegisterData((char *) &xlrec, sizeof(xl_checksum_state));
recptr = XLogInsert(RM_XLOG2_ID, XLOG2_CHECKSUMS);
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, recptr);
+
+ /* only loaded by SetDataChecksumsOn(), a no-op for the other callers */
+ INJECTION_POINT_CACHED("datachecksums-on-before-publish", NULL);
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ XLogCtl->data_checksum_version = new_type;
+ XLogCtl->data_checksum_lsn = recptr;
+ XLogCtl->data_checksum_is_local = false;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockRelease(DataChecksumTransitionLock);
+
XLogFlush(recptr);
+
+ return recptr;
}
/*
@@ -8966,11 +9343,19 @@ xlog_redo(XLogReaderState *record)
/* ControlFile->checkPointCopy always tracks the latest ckpt XID */
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->checkPointCopy.nextXid = checkPoint.nextXid;
- ControlFile->data_checksum_version = checkPoint.dataChecksumState;
UpdateControlFile();
LWLockRelease(ControlFileLock);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(checkPoint.dataChecksumState);
+
/*
* We should've already switched to the new TLI before replaying this
* record.
@@ -9206,19 +9591,17 @@ xlog_redo(XLogReaderState *record)
else if (info == XLOG_CHECKPOINT_REDO)
{
xl_checkpoint_redo redo_rec;
- bool new_state = false;
memcpy(&redo_rec, XLogRecGetData(record), sizeof(xl_checkpoint_redo));
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = redo_rec.data_checksum_version;
- SetLocalDataChecksumState(redo_rec.data_checksum_version);
- if (redo_rec.data_checksum_version != ControlFile->data_checksum_version)
- new_state = true;
- SpinLockRelease(&XLogCtl->info_lck);
-
- if (new_state)
- EmitAndWaitDataChecksumsBarrier(redo_rec.data_checksum_version);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(redo_rec.data_checksum_version,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(redo_rec.data_checksum_version);
}
else if (info == XLOG_LOGICAL_DECODING_STATUS_CHANGE)
{
@@ -9280,25 +9663,63 @@ xlog2_redo(XLogReaderState *record)
{
xl_checksum_state state;
XLogRecPtr lsn = record->EndRecPtr;
+ XLogRecPtr watermark;
memcpy(&state, XLogRecGetData(record), sizeof(xl_checksum_state));
+ SpinLockAcquire(&XLogCtl->info_lck);
+ watermark = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ /*
+ * Skip records this node has already applied. The control file
+ * carries the watermark, so this holds across restarts: recovery
+ * resuming below a record whose effect the control file already
+ * contains must not re-apply it, or it would revert a state change
+ * made with pg_checksums in between, which moves the state without
+ * writing any record of its own.
+ */
+ if (lsn <= watermark)
+ return;
+
/* advertise the location before the new state becomes visible */
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, lsn);
SpinLockAcquire(&XLogCtl->info_lck);
XLogCtl->data_checksum_version = state.new_checksum_state;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(state.new_checksum_state);
SpinLockRelease(&XLogCtl->info_lck);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = state.new_checksum_state;
+
+ /*
+ * Persist the new state, except when it is "on". Only "on" verifies
+ * checksums during reads, and between the last restartpoint and this
+ * record there may be pages on disk flushed under the old state; a
+ * crash-restart initializes verification from the control file and
+ * replay reads those pages back, so the control file may only say
+ * "on" once everything written under the transition has been flushed,
+ * as restartpoints and the end of recovery do.
+ * The opposite direction cannot wait for the restartpoint: once this
+ * record is replayed, evicted pages are written without checksums,
+ * and a control file still saying "on" would fail verification on
+ * exactly those pages after a crash.
+ */
+ if (state.new_checksum_state != PG_DATA_CHECKSUM_VERSION)
+ {
+ ControlFile->data_checksum_version = state.new_checksum_state;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ }
/*
* Update minRecoveryPoint to ensure that if recovery is aborted, we
* recover back up to this point before allowing hot standby again.
- * The new state is durable in pg_control while its location is only
- * tracked in shared memory; a standby becoming consistent below this
- * record would let base backups resume checksum verification with the
+ * The change location is only tracked in shared memory and is lost
+ * over a restart; a standby becoming consistent below this record
+ * would let base backups resume checksum verification with the
* location unknown. The local copies cannot be updated as long as
* crash recovery is happening and we expect all the WAL to be
* replayed.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 256b3a3c02e..77fa543f9d8 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -371,6 +371,7 @@ WaitLSN "Waiting to read or update shared Wait-for-LSN state."
LogicalDecodingControl "Waiting to read or update logical decoding status information."
DataChecksumsWorker "Waiting for data checksums worker."
AioWorkerControl "Waiting to update AIO worker information."
+DataChecksumTransition "Waiting for a data checksum state transition to be written to WAL."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 3b3ae23f1a6..20a99112838 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -648,6 +648,16 @@ main(int argc, char *argv[])
ControlFile->data_checksum_version =
(mode == PG_MODE_ENABLE) ? PG_DATA_CHECKSUM_VERSION : PG_DATA_CHECKSUM_OFF;
+ /*
+ * Mark the state as changed locally, without a WAL record. Recovery
+ * then knows the state is newer than anything the WAL carries and
+ * does not let a replayed checkpoint overwrite it. The watermark is
+ * left alone: any XLOG2_CHECKSUMS record this node had applied stays
+ * covered, and only records above it, written after this change, take
+ * effect again.
+ */
+ ControlFile->data_checksum_is_local = true;
+
if (do_sync)
{
pg_log_info("syncing data directory");
diff --git a/src/bin/pg_controldata/pg_controldata.c b/src/bin/pg_controldata/pg_controldata.c
index 6fc87ed114d..c363ce3dbb7 100644
--- a/src/bin/pg_controldata/pg_controldata.c
+++ b/src/bin/pg_controldata/pg_controldata.c
@@ -349,6 +349,10 @@ main(int argc, char *argv[])
(ControlFile->float8ByVal ? _("by value") : _("by reference")));
printf(_("Data page checksum version: %u\n"),
ControlFile->data_checksum_version);
+ printf(_("Data checksum watermark: %X/%08X\n"),
+ LSN_FORMAT_ARGS(ControlFile->data_checksum_lsn));
+ printf(_("Data checksum state is node-local: %s\n"),
+ (ControlFile->data_checksum_is_local ? _("yes") : _("no")));
printf(_("Default char data signedness: %s\n"),
(ControlFile->default_char_signedness ? _("signed") : _("unsigned")));
printf(_("Mock authentication nonce: %s\n"),
diff --git a/src/bin/pg_resetwal/pg_resetwal.c b/src/bin/pg_resetwal/pg_resetwal.c
index 1542a56ca4b..cdfb9898066 100644
--- a/src/bin/pg_resetwal/pg_resetwal.c
+++ b/src/bin/pg_resetwal/pg_resetwal.c
@@ -923,6 +923,13 @@ RewriteControlFile(void)
ControlFile.backupEndPoint = InvalidXLogRecPtr;
ControlFile.backupEndRequired = false;
+ /*
+ * The old WAL is gone and the new position may lie below the old
+ * watermark, which would make replay ignore future checksum transition
+ * records. The state itself is kept.
+ */
+ ControlFile.data_checksum_lsn = InvalidXLogRecPtr;
+
/*
* Force the defaults for max_* settings. The values don't really matter
* as long as wal_level='minimal'; the postmaster will reset these fields
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 2e86fd158d0..936469ea5f8 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -738,6 +738,20 @@ perform_rewind(filemap_t *filemap, rewind_source *source,
ControlFile_new.minRecoveryPoint = endrec;
ControlFile_new.minRecoveryPointTLI = endtli;
ControlFile_new.state = DB_IN_ARCHIVE_RECOVERY;
+
+ /*
+ * Keep the target's own data checksum state. Most of the data directory
+ * is still the target's: only blocks it changed since the divergence
+ * were copied from the source, so the source's state says nothing about
+ * the pages that stay. Replay from the last common checkpoint applies
+ * any WAL-logged transition the target has not seen (the watermark tells
+ * them apart), which converges the rewound server to the source's state
+ * whenever the WAL carries it.
+ */
+ ControlFile_new.data_checksum_version = ControlFile_target.data_checksum_version;
+ ControlFile_new.data_checksum_lsn = ControlFile_target.data_checksum_lsn;
+ ControlFile_new.data_checksum_is_local = ControlFile_target.data_checksum_is_local;
+
if (!dry_run)
update_controlfile(datadir_target, &ControlFile_new, do_sync);
}
diff --git a/src/include/catalog/pg_control.h b/src/include/catalog/pg_control.h
index 7b5404460ec..f898447f195 100644
--- a/src/include/catalog/pg_control.h
+++ b/src/include/catalog/pg_control.h
@@ -22,7 +22,7 @@
/* Version identifier for this pg_control format */
-#define PG_CONTROL_VERSION 1903
+#define PG_CONTROL_VERSION 1904
/* Nonce key length, see below */
#define MOCK_AUTH_NONCE_LEN 32
@@ -237,6 +237,25 @@ typedef struct ControlFileData
/* Current data checksums state */
uint32 data_checksum_version;
+ /*
+ * End of the newest XLOG2_CHECKSUMS record this node has written or
+ * applied. Replay ignores XLOG2_CHECKSUMS records at or below this
+ * point: their effect is already contained in data_checksum_version, or
+ * an offline pg_checksums change made after they were first applied
+ * supersedes them. InvalidXLogRecPtr if the node has never written or
+ * applied such a record.
+ */
+ XLogRecPtr data_checksum_lsn;
+
+ /*
+ * True when data_checksum_version was last set by pg_checksums rather
+ * than by a WAL-logged transition. Such a state is local to this node
+ * and newer than anything the WAL carries, so recovery must not replace
+ * it with a state taken from a checkpoint record. Cleared by the next
+ * WAL-logged transition.
+ */
+ bool data_checksum_is_local;
+
/*
* True if the default signedness of char is "signed" on a platform where
* the cluster is initialized.
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index d7eb648bd27..8d858be9927 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -89,6 +89,7 @@ PG_LWLOCK(54, WaitLSN)
PG_LWLOCK(55, LogicalDecodingControl)
PG_LWLOCK(56, DataChecksumsWorker)
PG_LWLOCK(57, AioWorkerControl)
+PG_LWLOCK(58, DataChecksumTransition)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/test/modules/test_checksums/Makefile b/src/test/modules/test_checksums/Makefile
index 71455cd5577..80f54bf6d8c 100644
--- a/src/test/modules/test_checksums/Makefile
+++ b/src/test/modules/test_checksums/Makefile
@@ -9,7 +9,7 @@
#
#-------------------------------------------------------------------------
-EXTRA_INSTALL = src/test/modules/injection_points
+EXTRA_INSTALL = contrib/pg_buffercache src/test/modules/injection_points
export enable_injection_points
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index fb7129d796f..70870715976 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -35,6 +35,15 @@ tests += {
't/009_fpi.pl',
't/010_backup_straddle.pl',
't/011_standby_straddle.pl',
+ 't/012_offline_standby.pl',
+ 't/013_rewind.pl',
+ 't/014_lockstep.pl',
+ 't/015_standby_crash_after_disable.pl',
+ 't/016_promote_enable_crash.pl',
+ 't/017_restartpoint_race.pl',
+ 't/018_enable_crash_windows.pl',
+ 't/019_standby_shutdown_catchup.pl',
+ 't/020_cascade_divergence.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
new file mode 100644
index 00000000000..241a1282ae9
--- /dev/null
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -0,0 +1,289 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Offline checksum changes with pg_checksums are local to one node. A
+# standby must neither adopt the state of the primary from replayed
+# checkpoint records, nor lose its own offline change to them.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', 'autovacuum = off');
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Scenario 1: enable offline on the primary only. The standby must
+# stay off, warn about the mismatch, and remain readable.
+$standby->stop;
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+test_checksum_state($primary, 'on');
+test_checksum_state($standby, 'off');
+
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after offline enable on the primary');
+
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+my @warnings = $log =~ /(does not match the state)/g;
+is(scalar(@warnings), 1, 'mismatch warned once per remote value');
+
+# Matching states re-arm the warning: undo the divergence on the primary,
+# then diverge again to the same value, all without restarting the standby.
+$primary->stop;
+$primary->checksum_disable_offline;
+$logstart = -s $standby->logfile;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no warning while the states match again');
+like($log, qr/now agrees with the replayed WAL/,
+ 'convergence is reported once the states match again');
+
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+test_checksum_state($standby, 'off');
+
+# The local state survives both clean and immediate restarts.
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+# Still in scenario 1's divergence: take a base backup *from the
+# standby*. A backup taken on a standby uses the last restartpoint as
+# its starting checkpoint (do_pg_backup_start()), so the record at the
+# redo point was written by the upstream primary and carries the
+# primary's state. Recovery from such a backup must not adopt that
+# state: the files were copied from the standby, and the new node
+# would come up verifying checksums its files do not have.
+
+# Make sure the standby has written pages under its own "off" state, so
+# its files really do lack checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,50000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', "SELECT count(*) FROM t2;");
+
+$standby->backup('from_standby');
+my $newnode = PostgreSQL::Test::Cluster->new('newnode');
+$newnode->init_from_backup($standby, 'from_standby');
+
+# Stream from the primary, not from the standby the files were copied
+# from, so the new node replays the "on" primary's records.
+$newnode->enable_streaming($primary);
+$newnode->start;
+
+# The new node's files all came from a cluster running with checksums
+# off. Anything but "off" here means it adopted the primary's state
+# through the checkpoint record at the redo point.
+my ($rc, $stdout, $stderr) = $newnode->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'the node copied from the standby accepts connections')
+ or diag("stderr: $stderr");
+is($stdout, 'off', 'backup of an "off" standby comes up with checksums off');
+
+# A checkpoint record streamed from the "on" primary must not flip the
+# state either.
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($newnode);
+test_checksum_state($newnode, 'off');
+
+# And it must be able to read the pages the standby wrote without
+# checksums.
+(undef, $stdout, $stderr) =
+ $newnode->psql('postgres', "SELECT count(*) FROM t2;");
+is($stdout, '50000', 'pages copied from the standby are readable')
+ or diag("stderr: $stderr");
+
+my $newnode_log = PostgreSQL::Test::Utils::slurp_file($newnode->logfile);
+unlike(
+ $newnode_log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the node copied from the standby');
+
+$newnode->stop('immediate');
+
+# Converge the cluster: enable offline on the standby too.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+$primary->wait_for_catchup($standby);
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after converging');
+
+# Scenario 2: disable offline on the standby only. The replayed
+# checkpoint records of the still-enabled primary must not override it.
+$standby->stop;
+$standby->checksum_disable_offline;
+$standby->start;
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'off');
+
+# Restartpoints must persist the local state, not the replayed copy.
+$standby->safe_psql('postgres', "CHECKPOINT;");
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable with checksums disabled locally');
+
+# Scenario 3: crash-restart right after an online transition, before the
+# next restartpoint. Replay then resumes from an older restartpoint whose
+# checkpoint records still carry the pre-transition state. Those must
+# still match the state seeded from the control file, and the transition
+# itself must be re-established by re-replaying the XLOG2_CHECKSUMS record,
+# without a spurious mismatch warning along the way.
+
+# Converge first: bring the standby back to "on" offline.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+disable_data_checksums($primary, wait => 'off');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+# Crash-restart the standby immediately, before any restartpoint has had a
+# chance to persist the new state to its control file.
+$logstart = -s $standby->logfile;
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001',
+ 'standby readable after crash-restart across an online transition');
+
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no spurious mismatch warning after crash-restart across an online transition'
+);
+
+# Scenario 4: a standby stopped while replaying an interrupted online
+# transition keeps the interrupted state in its own control file. A
+# primary is never caught this way, as its checksums launcher resolves
+# inprogress-on back to off from its exit cleanup. A standby has no
+# launcher; it carries forward whatever the last replayed record left
+# it in.
+
+# Block an online enable on the primary at inprogress-on with a
+# blocking temp table, same trick as in 004_offline.pl.
+my $bsession = $primary->background_psql('postgres');
+$bsession->query_safe('CREATE TEMPORARY TABLE tt (a integer);');
+enable_data_checksums($primary, wait => 'inprogress-on');
+
+# The standby picks up the in-progress state from the XLOG2_CHECKSUMS record.
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# Stop the standby cleanly; its restartpoint persists inprogress-on to
+# its own control file, since nothing on a standby resolves it away.
+$standby->stop;
+$standby->start;
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# The primary still sits at inprogress-on, so a base backup taken now
+# copies a control file and a redo point that both carry the
+# in-progress state; replay of the XLOG2_CHECKSUMS records completes
+# the transition on the new standby.
+$primary->backup('inprogress_backup');
+my $standby2 = PostgreSQL::Test::Cluster->new('standby2');
+$standby2->init_from_backup($primary, 'inprogress_backup',
+ has_streaming => 1);
+$standby2->start;
+
+# Backup label recovery adopts the redo point's state immediately,
+# before any further WAL is replayed.
+wait_for_checksum_state($standby2, 'inprogress-on');
+
+$bsession->quit;
+wait_for_checksum_state($primary, 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable once the transition completes');
+
+# Scenario 4 continued: the backup taken mid-transition must also
+# complete the transition and stay readable.
+$primary->wait_for_catchup($standby2);
+wait_for_checksum_state($standby2, 'on');
+
+is($standby2->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'mid-transition backup standby readable after the transition');
+
+$standby2->stop;
+
+# Nothing along this standby's replay can legitimately disagree: it
+# started out at the same inprogress-on state as the primary.
+my $standby2_log = PostgreSQL::Test::Utils::slurp_file($standby2->logfile);
+unlike(
+ $standby2_log,
+ qr/does not match the state/,
+ 'no mismatch warning on a backup taken mid-transition');
+
+# No extra CHECKPOINT needed: the transition's completion checkpoint
+# wrote every page with a checksum, and the shutdown restartpoint
+# flushes the rest.
+command_ok([ 'pg_checksums', '--check', '-D', $standby2->data_dir ],
+ 'checksums valid on the mid-transition backup standby');
+
+$standby->stop;
+$primary->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/013_rewind.pl b/src/test/modules/test_checksums/t/013_rewind.pl
new file mode 100644
index 00000000000..a791e24317d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/013_rewind.pl
@@ -0,0 +1,201 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test pg_rewind across an online data checksum enable.
+#
+# A clean switchover leaves the shutdown checkpoint of the old primary
+# as the last common checkpoint between the two nodes. When data
+# checksums are enabled online on the new primary before the old one is
+# rewound, pg_rewind installs the control file of the new primary, which
+# already claims checksums are fully enabled, while replay begins at the
+# shutdown checkpoint whose record still carries the old state. Replay
+# of the WAL stretch from before the enable must run with checksums off,
+# as recorded in the checkpoint, else it would verify pages which never
+# had checksums written and fail recovery.
+#
+# The new primary requests a checkpoint right after promotion, and
+# replaying its CHECKPOINT_REDO record would repair the state before
+# any interesting WAL is reached. Hold the checkpointer on the standby
+# in a restartpoint over the promotion, like in the recovery test
+# 041_checkpoint_at_promote, so that the post-promotion writes end up
+# in WAL before the first checkpoint of the new timeline.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Old primary. full_page_writes is off so that the updates done on the
+# promoted node do not carry full page images, forcing replay on the
+# rewound node to read the pages from disk. wal_log_hints is required
+# by pg_rewind on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+full_page_writes = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+log_checkpoints = on
+]);
+$node_a->start;
+
+if (!$node_a->check_extension('injection_points'))
+{
+ plan skip_all => 'Extension injection_points not installed';
+}
+
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+$node_a->safe_psql('postgres', "CREATE EXTENSION injection_points;");
+
+# Set the hint bits on t before taking the backup, so that reads on the
+# promoted node do not emit full page images for its pages later.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->safe_psql('postgres', "SELECT count(*) FROM t;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_a, 'off');
+test_checksum_state($node_b, 'off');
+
+# Hold the next restartpoint on the standby.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point', 'wait');");
+
+# Give the restartpoint a checkpoint record to work on, then start it
+# in a background session; it will block on the injection point with
+# the checkpointer busy until released.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+my $bg_psql = $node_b->background_psql('postgres', on_error_stop => 0);
+$bg_psql->query_until(
+ qr/starting_restartpoint/, q(
+ \echo starting_restartpoint
+ CHECKPOINT;
+));
+$node_b->wait_for_event('checkpointer', 'create-restart-point');
+
+# Clean switchover: the shutdown checkpoint of A streams to B and
+# becomes the last common checkpoint.
+$node_a->stop('fast');
+
+my ($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+$stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+my $shutdown_ckpt = $1;
+$node_b->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_ckpt'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+
+my $logstart = -s $node_b->logfile;
+$node_b->promote;
+
+# Accidental restart of the old primary, diverging its timeline.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->stop('fast');
+
+# Updates on the new primary before its first checkpoint; replay of
+# these on the rewound node has to read the pages from disk.
+$node_b->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 25 = 0;");
+ok( !$node_b->log_contains("checkpoint complete", $logstart),
+ "no checkpoint on the new timeline before the updates");
+
+# Release the checkpointer; the queued post-promotion checkpoint runs
+# after the updates.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+$bg_psql->quit;
+
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+# pg_rewind refuses to run with full_page_writes disabled on the
+# source; the updates it was disabled for are already in WAL.
+$node_b->safe_psql('postgres', "ALTER SYSTEM SET full_page_writes = on;");
+$node_b->safe_psql('postgres', "SELECT pg_reload_conf();");
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the new primary');
+
+# Replay on the rewound node must start at the shutdown checkpoint of
+# the switchover, with the control file of the new primary.
+my $backup_label = slurp_file($node_a->data_dir . '/backup_label');
+$backup_label =~ /^CHECKPOINT LOCATION: ([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from backup_label";
+is($1, $shutdown_ckpt, 'replay starts at the switchover checkpoint');
+
+($stdout, $stderr) = run_command(
+ [
+ 'pg_waldump',
+ '-p' => $node_a->data_dir . '/pg_wal',
+ '-t' => 1,
+ '-s' => $shutdown_ckpt,
+ '-n' => 1,
+ ]);
+like($stdout, qr/CHECKPOINT_SHUTDOWN/,
+ 'last common checkpoint is a shutdown checkpoint');
+
+# pg_rewind keeps the target's own checksum state in the control file it
+# writes; the online enable reaches the rewound node through WAL replay
+# below, not through the copied control file.
+($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+like(
+ $stdout,
+ qr/^Data page checksum version:\s*0$/m,
+ 'rewound node keeps its own checksum state in the control file');
+
+# Start the rewound node as a standby of the new primary. Replay runs
+# through the pre-enable WAL stretch and the online enable.
+#
+# The rewind replaced the configuration files with those of the new
+# primary, so put the port back.
+my $connstr = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_b->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/014_lockstep.pl b/src/test/modules/test_checksums/t/014_lockstep.pl
new file mode 100644
index 00000000000..713f37fb22c
--- /dev/null
+++ b/src/test/modules/test_checksums/t/014_lockstep.pl
@@ -0,0 +1,182 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# The lockstep procedure for offline checksum changes in a replication
+# setup: stop all nodes, run pg_checksums on all of them, restart.
+# Replay of checkpoint records written before the change must not
+# revert the state of the standby.
+#
+# The same pair then runs the procedure in the disable direction, and
+# finally across a divergence: after a failover both nodes get their
+# checksums enabled offline, while the last common checkpoint still
+# carries "off". pg_rewind must keep the offline "on" on the rewound
+# node, since no record in the replayed WAL could ever restore it.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# wal_log_hints keeps the pair eligible for pg_rewind without checksums.
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Part 1: the lockstep procedure in the enable direction. Stop the
+# standby first: the WAL written after this point is replayed only
+# after the offline switch, and every checkpoint record in it still
+# carries the old state.
+$standby->stop;
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 0;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 1;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+# The lockstep procedure.
+$primary->checksum_enable_offline;
+$standby->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to off.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "off" in the replayed WAL/,
+ 0);
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep enable');
+
+# Crash the standby and replay the same stretch again.
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'on');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after crash restart');
+
+# Once a post-switch checkpoint has been replayed the states match and
+# no warning may be logged.
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no mismatch warning once the states match');
+
+# Every page the standby wrote in this window must carry a checksum.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$standby->stop;
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'checksums valid on the standby');
+
+# Part 2: the lockstep procedure in the other direction. The standby
+# is already stopped; give it a pre-switch WAL stretch to replay, with
+# every checkpoint record in it still carrying "on" -- an explicit
+# checkpoint plus the shutdown checkpoint written by stop().
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 2;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+$primary->checksum_disable_offline;
+$standby->checksum_disable_offline;
+
+my $disable_logstart = -s $standby->logfile;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to on.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $disable_logstart);
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep disable');
+
+# Part 3: failover, divergence and pg_rewind across offline enables.
+# The roles swap from here on: the standby becomes the new primary and
+# the old primary is rewound to follow it, but the variables keep the
+# names they had above.
+#
+# Failover: promote the standby, then diverge the old primary.
+$standby->promote;
+$standby->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$primary->stop;
+
+# The lockstep procedure across the divergence: both nodes get their
+# checksums enabled offline.
+$primary->checksum_enable_offline;
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+
+# The states match, so the rewind proceeds without complaint.
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $primary->data_dir,
+ '--source-server' => $standby->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled offline on both nodes');
+
+# The rewind replaced the configuration files with those of the source,
+# so put the port back. The copied file also carries the source's own
+# stale primary_conninfo; enable_streaming() below appends ours, which
+# wins as the later entry.
+$primary->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$primary->port]}
+]);
+$primary->enable_streaming($standby);
+
+my $rewind_logstart = -s $primary->logfile;
+$primary->start;
+$standby->wait_for_catchup($primary);
+
+is($primary->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+
+# The offline enable survives: replay from the common checkpoint must
+# not resurrect the pre-divergence "off".
+test_checksum_state($primary, 'on');
+
+$log =
+ PostgreSQL::Test::Utils::slurp_file($primary->logfile, $rewind_logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no divergence reported between the rewound server and its source');
+
+$primary->stop;
+$standby->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
new file mode 100644
index 00000000000..315b27b93d4
--- /dev/null
+++ b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
@@ -0,0 +1,125 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that a standby crashing after a replayed online disable, but before
+# its next restartpoint, restarts cleanly.
+#
+# Pages dirtied before the XLOG2_CHECKSUMS record and evicted after it are
+# written out under the new "off" state, without checksums. The control
+# file must follow the record immediately in this direction: a crash-restart
+# initializes checksum verification from the control file, and one still
+# saying "on" would fail verification on exactly those pages while replaying
+# records older than the state change.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$primary->start;
+
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,1000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+
+# Small buffer pool so that replaying a bulk load evicts the pages we care
+# about, and no restartpoints so the control file stays where it is.
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 1MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_delay = 10000
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'on');
+
+# No full page images, so that replay has to read the pages of "t" back from
+# disk instead of overwriting them from the WAL.
+$primary->append_conf('postgresql.conf', 'full_page_writes = off');
+$primary->reload;
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" on the standby, *before* the state change. They are
+# not flushed: the standby has no restartpoint from here on.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+# Online disable. The standby replays it and switches to "off", but its
+# control file is not updated.
+$primary->safe_psql('postgres', 'SELECT pg_disable_data_checksums();');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+my ($ctl_before) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($ctl_state) = $ctl_before =~ /Data page checksum version:\s+(\d+)/;
+note(
+ "standby control file data_checksum_version after the disable: "
+ . "$ctl_state (0 = off, 1 = on)");
+
+# Force the standby to evict the dirty pages of "t" now that it is "off":
+# they are written back without checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE filler AS SELECT generate_series(1,300000) AS a;");
+$primary->wait_for_catchup($standby);
+
+# Crash the standby before it gets a chance to run a restartpoint.
+$standby->stop('immediate');
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'standby restarts after crashing between a checksum state '
+ . 'change and the next restartpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'the pages written while "off" are readable')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("standby log tail:\n" . join("\n", @lines[ -12 .. -1 ]))
+ if @lines >= 12;
+ fail('the pages written while "off" are readable');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/016_promote_enable_crash.pl b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
new file mode 100644
index 00000000000..97f84da5685
--- /dev/null
+++ b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a standby promoted right after replaying an online enable, before any
+# restartpoint has flushed the rewritten pages.
+#
+# The end-of-recovery record persists the checksum state without flushing the
+# buffer pool, while the control file still points at a restartpoint older than
+# the transition. Persisting "on" there would make a crash before the
+# post-promotion checkpoint verify checksums over the pages that were flushed
+# while the state was still "off"; promotion must take the full checkpoint
+# instead.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Dirty the pages of "t" without full page images, then push them out to disk
+# on the standby while checksums are still off.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Online enable on the primary; the standby replays it into its own buffers,
+# where the rewritten pages stay dirty (no restartpoint, no bgwriter).
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+note("standby control file before the promotion: $before");
+
+# Promote. The end-of-recovery record persists the live state without any
+# buffer flush; the checkpoint requested afterwards is not immediate.
+$standby->promote;
+$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery();')
+ or die 'timed out waiting for the promotion';
+$standby->stop('immediate');
+
+($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($ckpt) = $ctl =~ /Latest checkpoint location:\s+(\S+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note(
+ "standby control file after the promotion: $after, checkpoint $ckpt, redo $redo"
+);
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'promoted node restarts after crashing before its first checkpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/017_restartpoint_race.pl b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
new file mode 100644
index 00000000000..65aa834f366
--- /dev/null
+++ b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
@@ -0,0 +1,138 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a restartpoint whose flush races the replay of an online enable.
+#
+# Replay keeps running while CheckPointGuts() writes out the buffer pool, so
+# the state at the end of the flush can be newer than the one the flush ran
+# under. Persisting it would claim checksums for pages the flush wrote out
+# before the transition, against a redo pointer that predates it.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# C1: the checkpoint the pending restartpoint will target.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" without full page images and push them to disk on
+# the standby while it is still "off".
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Start a restartpoint for C1 and hold it right after CheckPointGuts().
+$standby->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point','wait');");
+my $bg = $standby->background_psql('postgres');
+$bg->query_until(qr//, "\\echo restartpoint\nCHECKPOINT;\n");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint injection point';
+
+# The startup process keeps replaying while the checkpointer is held: the
+# whole online enable lands, and the rewritten pages stay dirty.
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+# Let the restartpoint finish; it persists the state it samples now.
+$standby->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint to finish';
+$standby->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+
+$standby->stop('immediate');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note("standby control file after the restartpoint: $after, redo $redo");
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started, 'standby restarts after a restartpoint that raced the enable');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/018_enable_crash_windows.pl b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
new file mode 100644
index 00000000000..17f96ff6272
--- /dev/null
+++ b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
@@ -0,0 +1,496 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Crash and checkpoint windows inside an online data checksum enable on
+# a single node. Each scenario starts from "off" on the same cluster,
+# runs the enable into a held injection point, and checks what a crash
+# or a concurrent checkpoint in that window may and may not do to the
+# control file and the pages on disk.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+use IPC::Run;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Bring the node back to "off" with no leftover table between
+# scenarios. The crash restarts already resolve an interrupted enable
+# to "off", so only disable when something is left to disable. Those
+# restarts also clear any attached injection points. The checkpoint
+# bounds the WAL the next scenario's crash recovery has to replay.
+sub reset_scenario
+{
+ my ($node) = @_;
+
+ BAIL_OUT('node is not running after the previous scenario')
+ unless $node->is_alive;
+
+ my $state = $node->safe_psql('postgres', 'SHOW data_checksums;');
+ disable_data_checksums($node, wait => 'off') if $state ne 'off';
+ $node->safe_psql('postgres', 'DROP TABLE IF EXISTS t;');
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+ test_checksum_state($node, 'off');
+}
+
+my $node = PostgreSQL::Test::Cluster->new('crash_windows');
+$node->init(no_data_checksums => 1);
+
+# Scenario 1 needs the pages of "t" to reach disk without a full page
+# image, hence full_page_writes and wal_log_hints off. Scenario 2
+# needs them to stay dirty in shared buffers until a checkpoint, hence
+# no bgwriter and a large shared_buffers. wal_level is pinned because
+# init() writes "minimal" for a non-streaming node. The rest merely
+# tolerate the settings.
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+wal_level = replica
+max_wal_size = 10GB
+full_page_writes = off
+wal_log_hints = off
+shared_buffers = 512MB
+bgwriter_lru_maxpages = 0
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$node->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+
+test_checksum_state($node, 'off');
+
+# Scenario 1: a primary crashes inside SetDataChecksumsOn(), just before the
+# forced checkpoint that flushes the rewritten pages, with crash recovery
+# resuming from a checkpoint older than the transition. The control file
+# must still say "inprogress-on" there: replay does not adopt the state of
+# the checkpoint record it resumes from, so an "on" written before the flush
+# would stay in effect while replay reads pages whose rewrite never reached
+# disk. Here the pages went out through the rewriting worker's ring buffer,
+# so only the control file state discriminates; see the next scenario for
+# the case where they did not.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Establish the checkpoint that crash recovery will resume from.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty the pages of "t" without full page images, then push them out
+ # to disk while checksums are still off.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+ my $evicted = $node->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+ note("evict_relation on t: $evicted, relpath $relpath");
+
+ # Stop the enable right after the control file has been updated to
+ # "on" but before the checkpoint that flushes the rewritten pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+ $node->safe_psql('postgres', 'SELECT pg_enable_data_checksums();');
+
+ $node->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'datachecksums-on-before-checkpoint';"
+ ) or die 'timed out waiting for the injection point';
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ $node->stop('immediate');
+
+ # Show the on-disk checksum field of the first page of "t".
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 2: the same crash as above, but with the rewritten pages resident
+# in shared buffers instead of gone out through the ring. The rewriting
+# worker reads through a BAS_VACUUM ring, which writes pages back as the
+# ring recycles, but a page already resident in shared buffers is not read
+# through the ring: ReadBufferExtended() hands back the existing buffer, and
+# it stays dirty until a checkpoint. The control file may therefore not
+# say "on" before that checkpoint has run, or crash recovery would resume
+# from a checkpoint older than the transition with verification already
+# enabled, and replay records that read those still-unchecksummed pages.
+# full_page_writes is off so that the records do not simply overwrite the
+# pages with a full page image.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,100000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Establish the checkpoint that crash recovery will resume from, with the
+ # pages of "t" written out while checksums are still off.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty those pages again without emitting full page images, and leave
+ # them in shared buffers. Replay of these records has to read the
+ # pages from disk.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+
+ my $dirty_before = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_before, '>', 0,
+ 'pages of t are resident and dirty before enabling checksums');
+
+ # Hold the enabling right before the checkpoint that flushes the rewritten
+ # pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ # The rewritten pages must still be sitting dirty in shared buffers, or
+ # the window this test is about does not exist.
+ my $dirty_after = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_after, '>', 0,
+ 'rewritten pages of t are still dirty before the checkpoint');
+
+ $node->stop('immediate');
+
+ # The on-disk copy is the one written before the transition, without a
+ # checksum.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ is($pd_checksum, 0, 'block 0 of t on disk carries no checksum');
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 3: a checkpoint that runs between the XLOG2_CHECKSUMS("on")
+# record and the control file write at the end of SetDataChecksumsOn() must
+# not make a crash throw the completed transition away. SetDataChecksumsOn()
+# writes the record, flips shared memory to "on", emits the barrier and only
+# then requests the checkpoint that flushes the rewritten pages; the control
+# file is written after that checkpoint returns. A crash in that window is
+# harmless only as long as recovery still starts before the record. Any
+# checkpoint completing in the window moves the redo point past the record,
+# so recovery would never see it, would come up with the control file's
+# "inprogress-on" and StartupXLOG() would demote that to "off", even though
+# every page on disk carries a checksum by then. CreateCheckPoint()
+# therefore persists the state the checkpoint ran under, the same way
+# CreateRestartPoint() does. Here the window is made deterministic by
+# holding the launcher at the datachecksums-on-before-checkpoint injection
+# point and checkpointing from another session.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Hold the launcher after the record, the shared memory flip and the
+ # barrier, but before the checkpoint SetDataChecksumsOn() requests
+ # itself.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ # Every backend already sees "on" and writes checksums.
+ test_checksum_state($node, 'on');
+
+ # ... while the control file still says "inprogress-on".
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($before, '3', 'control file says "inprogress-on" inside the window');
+
+ # A concurrent checkpoint. It flushes the rewritten pages and moves
+ # the redo point past the XLOG2_CHECKSUMS("on") record, so it has to
+ # record the "on" state in the control file as well.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($after, '1',
+ 'the concurrent checkpoint records "on" in the control file');
+
+ $node->stop('immediate');
+
+ # The checkpoint flushed the rewritten pages, so they carry a checksum on
+ # disk: the transition really did complete.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ isnt($pd_checksum, 0, 'block 0 of t on disk carries a checksum');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition is complete on disk, so the cluster has to come back
+ # "on".
+ test_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+}
+reset_scenario($node);
+
+# Scenario 4: a crash between the checkpoint an online enable requests and
+# the control file write that follows it must not lose the transition. The
+# last steps of SetDataChecksumsOn() are
+#
+# WAL record -> shmem -> barrier -> checkpoint -> persist
+#
+# The checkpoint is what makes "on" safe to persist, but it also moves the
+# redo point above the XLOG2_CHECKSUMS record that carries the new state.
+# Crash recovery started from that checkpoint therefore never replays the
+# record, so without the state the checkpoint itself persists, a crash in
+# the remaining window would bring the cluster back at "inprogress-on",
+# which StartupXLOG() resolves to "off", discarding a transition whose
+# pages are all on disk with a checksum. The previous scenario exercises
+# the same window through a checkpoint requested by another session; this
+# one closes it from the other side: the transition's own checkpoint is the
+# one that moves the redo point, so persisting the state from
+# CreateCheckPoint() is what has to save it, not the write below the
+# injection point.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Hold the launcher after the checkpoint that licenses "on" has
+ # completed and before the state reaches the control file.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-after-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-after-checkpoint');
+
+ # The transition is complete as far as the running cluster is concerned.
+ test_checksum_state($node, 'on');
+
+ # The checkpoint has already recorded it, so the pending write below the
+ # injection point has nothing left to do.
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($version) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($version, '1',
+ 'the requested checkpoint recorded "on" in the control file');
+
+ # Crash before the control file write that follows the checkpoint.
+ $node->stop('immediate');
+
+ $node->start;
+
+ # Every page on disk carries a checksum and the checkpoint that
+ # flushed them completed, so the cluster has to come back verifying
+ # them.
+ test_checksum_state($node, 'on');
+
+ is($node->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '10000', 'relation readable after the crash');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ # The data directory must be one the offline tools accept.
+ $node->stop;
+ $node->command_ok([ 'pg_checksums', '--check', '-D', $node->data_dir ],
+ 'pg_checksums accepts the data directory');
+
+ # Leave the node running for the reset.
+ $node->start;
+}
+reset_scenario($node);
+
+# Scenario 5: a checkpoint racing SetDataChecksumsOn() between the insertion
+# of the XLOG2_CHECKSUMS("on") record and the shared memory update must not
+# insert an XLOG_CHECKPOINT_REDO record that follows the transition in WAL
+# order while still carrying "inprogress-on". Recovery resuming from such a
+# redo point never replays the preceding transition record, comes up in
+# "inprogress-on" and resolves the finished transition as interrupted.
+# XLogChecksums() closes the window by inserting the record and publishing
+# the new state under DataChecksumTransitionLock, which CreateCheckPoint()
+# takes around sampling the state and inserting the redo record. Here the
+# launcher is held between the two steps at the
+# datachecksums-on-before-publish injection point, and a concurrent
+# CHECKPOINT has to block on the lock instead of completing inside the
+# window.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # The datachecksums-on-before-publish point fires inside a critical
+ # section, where the wait machinery must not allocate. Waiting once
+ # at the datachecksums-enable-checksums-delay point, which the
+ # launcher runs outside the critical section, initializes it; see
+ # 050_redo_segment_missing.pl for the same recipe around
+ # create-checkpoint-run.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-enable-checksums-delay','wait');"
+ );
+
+ # Hold the launcher after the XLOG2_CHECKSUMS("on") record is in WAL but
+ # before the new state is published in shared memory.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-publish','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-enable-checksums-delay');
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-enable-checksums-delay');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-enable-checksums-delay');");
+
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-publish');
+
+ # The record is in WAL, the published state is still the old one.
+ test_checksum_state($node, 'inprogress-on');
+
+ # A concurrent checkpoint. It must block on DataChecksumTransitionLock
+ # before inserting its redo record rather than complete inside the window.
+ my $checkpointer = IPC::Run::start(
+ [ 'psql', '-XAtq', '-d', $node->connstr('postgres'), '-c', 'CHECKPOINT;' ],
+ '<' => '/dev/null',
+ '>' => '/dev/null',
+ '2>' => '/dev/null',
+ IPC::Run::timer($PostgreSQL::Test::Utils::timeout_default));
+
+ ok( $node->poll_query_until(
+ 'postgres',
+ "SELECT wait_event = 'DataChecksumTransition' "
+ . "FROM pg_stat_activity WHERE backend_type = 'checkpointer';"),
+ 'concurrent checkpoint blocks on DataChecksumTransitionLock');
+
+ # Release the launcher; the checkpoint then samples the published "on" and
+ # its redo record follows the transition record.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-publish');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-publish');");
+
+ $checkpointer->finish;
+
+ # Crash while the launcher's own checkpoint may still be in flight.
+ # Recovery resumes from the concurrent checkpoint's redo point, which
+ # now lies above the transition record and carries "on".
+ $node->stop('immediate');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition completed, so the cluster has to come back "on".
+ wait_for_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ $node->stop;
+}
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
new file mode 100644
index 00000000000..aaa35f8f7dd
--- /dev/null
+++ b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
@@ -0,0 +1,126 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A standby stopped after replaying an online enable, but before the
+# checkpoint record that follows it on the primary.
+#
+# The shutdown restartpoint has no new checkpoint record to work from
+# and is skipped, so the control file would keep the in-progress state
+# the transition passed through, even though replay left the node at
+# "on" and rewrote every page. pg_checksums reads that field and
+# refuses to run on an in-progress state, so the shutdown has to flush
+# and catch it up instead.
+#
+# The caught-up control file then carries the watermark of the "on"
+# record, which the second half of this test relies on: an offline
+# disable made while the standby is down must survive the re-replay of
+# that record on the next startup, or the change would be silently
+# reverted.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+wal_keep_size = 1GB
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Anchor the standby's last restartpoint here, so that nothing replayed
+# from now on gives the shutdown restartpoint a newer checkpoint record
+# to use.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Hold the enable right after the state change record, before the
+# checkpoint it requests once the transition is complete.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+enable_data_checksums($primary);
+$primary->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+wait_for_checksum_state($primary, 'on');
+
+# The state change record is already flushed and streams on its own;
+# no checkpoint record follows while the launcher is held.
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+$standby->stop('fast');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($state, '1',
+ 'shutdown catches the control file up with the replayed state');
+
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'pg_checksums verifies the stopped standby');
+
+# Disable checksums offline while the standby is down. The restart
+# below resumes replay under the last restartpoint, re-reading the "on"
+# record; the watermark persisted with the catch-up above marks it as
+# already applied, so the offline change must survive.
+$standby->checksum_disable_offline;
+
+# Let the enable finish on the primary.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-checkpoint');");
+$primary->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-checkpoint');");
+
+# The launcher exits only after the checkpoint it requests completes,
+# so a checkpoint record now follows the transition in WAL.
+$primary->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_catalog.pg_stat_activity "
+ . "WHERE backend_type = 'datachecksums launcher';");
+
+my $log_offset = -s $standby->logfile;
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+
+# The nodes legitimately diverged, which replayed checkpoints report.
+$standby->wait_for_log(
+ qr/data checksum state "off" of this node does not match the state "on" in the replayed WAL/,
+ $log_offset);
+
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/020_cascade_divergence.pl b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
new file mode 100644
index 00000000000..b4b474fa583
--- /dev/null
+++ b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
@@ -0,0 +1,115 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# An offline data checksum state change made on one node of a cascading setup
+# is reported all the way down the chain.
+#
+# CheckReplayedDataChecksumState() is reached from xlog_redo() when a
+# checkpoint record (XLOG_CHECKPOINT_SHUTDOWN or XLOG_CHECKPOINT_REDO) is
+# replayed after consistency has been reached. Those records are written by
+# the root primary only and are relayed verbatim by every intermediate
+# standby, so a cascaded standby that was changed offline learns about the
+# divergence too. An intermediate standby's own restartpoints are not
+# WAL-logged and cannot surface it.
+#
+# Note that this makes detection checkpoint-driven: an idle or read-only
+# primary surfaces the divergence no sooner than its next checkpoint, so up to
+# checkpoint_timeout may pass before the operator sees anything. Closing that
+# window would require comparing the state when a walreceiver connects.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# checkpoint_timeout is deliberately long so that the only checkpoint record in
+# this test is the one requested explicitly below.
+my $conf = qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+);
+
+my $primary = PostgreSQL::Test::Cluster->new('cascade_primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', $conf);
+$primary->start;
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,1000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('cascade_standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf('postgresql.conf', $conf);
+$standby->start;
+
+$standby->backup('backup');
+my $cascade = PostgreSQL::Test::Cluster->new('cascade_cascade');
+$cascade->init_from_backup($standby, 'backup', has_streaming => 1);
+$cascade->append_conf('postgresql.conf', $conf);
+$cascade->start;
+
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+test_checksum_state($cascade, 'off');
+
+# Change the cascaded standby offline, and only it.
+$cascade->stop;
+$cascade->command_ok([ 'pg_checksums', '--enable', '-D', $cascade->data_dir ],
+ 'pg_checksums enables checksums on the cascaded standby only');
+
+my $logstart = -s $cascade->logfile;
+$cascade->start;
+
+test_checksum_state($cascade, 'on');
+
+# Ordinary WAL traffic carries no state, so the cascaded standby keeps
+# streaming from a chain whose state is "off" without noticing.
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (1);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+is($cascade->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '1001', 'cascaded standby keeps streaming from a divergent chain');
+
+# Restartpoints on the intermediate standby are not WAL-logged, so this does
+# not surface anything on the cascaded standby either.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (2);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+my $log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+unlike(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'an intermediate restartpoint does not surface the divergence');
+
+# A checkpoint on the root primary does: the record is relayed down the whole
+# chain, so both the direct and the cascaded standby compare it against their
+# own state. wait_for_catchup() below waits for replay, not just receipt, so
+# the record has been through xlog_redo() on both nodes once it returns.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+$log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+like(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'the cascaded standby reports the divergence on the primary checkpoint');
+
+$cascade->stop;
+$standby->stop;
+$primary->stop;
+
+done_testing();
--
2.55.0
[application/octet-stream] v6-0005-doc-Explain-how-offline-and-online-checksum-chang.patch (4.2K, ../../CAN4CZFNqKg9Ts76r922cKWcOgtH32ocyghQa8NLB-qeCeC1RKg@mail.gmail.com/5-v6-0005-doc-Explain-how-offline-and-online-checksum-chang.patch)
download | inline diff:
From 2182d31568dd652a936f664b2c800f429b0ea7db Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Mon, 31 Aug 2026 08:40:44 +0000
Subject: [PATCH v6 5/5] doc: Explain how offline and online checksum changes
interact
The lockstep procedure for offline checksum changes in a replication
setup was documented without preconditions. It is not sufficient on
its own: an offline change is recorded only in the control file and
has no ordering against WAL the node has not replayed yet, so a node
stopped before replaying an online state transition applies it on
restart and overrides the offline change. The nodes then silently
diverge even though the change was applied to all of them while
stopped.
Document the two durability models next to each other in wal.sgml,
and extend the pg_checksums notes: require standbys to have replayed
all WAL of their upstream node before stopping them, show how to
check that, and recommend not mixing online and offline changes.
---
doc/src/sgml/ref/pg_checksums.sgml | 25 ++++++++++++++++++++++---
doc/src/sgml/wal.sgml | 13 +++++++++++++
2 files changed, 35 insertions(+), 3 deletions(-)
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index 000940e7cb7..b337950be3d 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -248,9 +248,28 @@ PostgreSQL documentation
directory; the new state does not propagate over replication. In a
replication setup the same change must be applied to every node: stop all
nodes, run <application>pg_checksums</application> on each of them, and
- only then restart them. Tools that copy relation file blocks directly
- between nodes, such as <xref linkend="app-pgrewind"/>, likewise require
- both nodes to be in the same data checksum state.
+ only then restart them. Before stopping a standby, make sure it has
+ replayed all WAL of its upstream node, for example by stopping the
+ primary first and comparing
+ <function>pg_last_wal_replay_lsn()</function> with
+ <function>pg_last_wal_receive_lsn()</function> on the standby. Tools
+ that copy relation file blocks directly between nodes, such as
+ <xref linkend="app-pgrewind"/>, likewise require both nodes to be in
+ the same data checksum state.
+ </para>
+ <para>
+ The replay requirement exists because an offline change is recorded
+ only in the control file and has no defined ordering against WAL the
+ node has not replayed yet; see
+ <xref linkend="checksums-offline-enable-disable"/>. A node stopped
+ before replaying an online checksum state change applies that change
+ when it is restarted, overriding the offline change, and the states of
+ the nodes silently diverge until a later checkpoint record triggers
+ the warning described below. Because of this it is best not to mix
+ the two mechanisms: change the state of a replication setup either
+ with the offline procedure above or with an online transition, and
+ make sure the previous change has reached every node before starting
+ the next one.
</para>
<para>
If the change is applied inconsistently, each node keeps its own state,
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index db18a8c516e..d88a832c2a8 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -317,6 +317,19 @@
verify checksums, on an offline cluster.
</para>
+ <para>
+ An offline change provides durability differently from an
+ <link linkend="checksums-online-enable-disable">online change</link>.
+ An online transition is WAL-logged: it is ordered against all other
+ WAL records, it is replayed after a crash, and it propagates to
+ standbys. An offline change is recorded only in the cluster's
+ control file: it writes no WAL, it is invisible to replication, and
+ it has no defined ordering against WAL the node has not replayed
+ yet. When a node later replays WAL that contains an online checksum
+ state change, that change takes effect on the node even if it was
+ written before the offline change was made.
+ </para>
+
<para>
An offline change only affects the data directory it is run on; the
new state does not propagate over replication. In a replication setup
--
2.55.0
[application/octet-stream] v6-0003-pg_rewind-Check-the-data-checksum-states-of-sourc.patch (21.6K, ../../CAN4CZFNqKg9Ts76r922cKWcOgtH32ocyghQa8NLB-qeCeC1RKg@mail.gmail.com/6-v6-0003-pg_rewind-Check-the-data-checksum-states-of-sourc.patch)
download | inline diff:
From 1ac1c682e1320857284a2b8e8a3909e040809fdf Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 18:40:22 +0000
Subject: [PATCH v6 3/5] pg_rewind: Check the data checksum states of source
and target
pg_rewind installs the source's control file on the target, but every
block it does not copy keeps the target's content. With checksums
enabled on the target and disabled on the source, the rewound server
claims enabled checksums while the blocks copied from the source have
none, and fails checksum verification as soon as it reads them; in the
worst case every connection attempt dies on an unverifiable catalog
page.
Comparing the control files is not enough. Replay on the rewound
server resumes from the last common checkpoint and adopts the data
checksum state recorded there, so an offline disable on the source
after the divergence leaves both control files saying "off" while the
rewound server still resumes with verification enabled. Check the
state carried by the divergence checkpoint as well, which pg_rewind
already reads.
Refuse both cases, and an interrupted online transition on either
side, same as pg_checksums does. The opposite mismatch, checksums on
the source only, is allowed with a warning: recovery under the backup
label adopts the divergence state, so the rewound server keeps
checksums disabled (or converges through the replayed WAL if the
source enabled them online), and no page can fail verification.
---
doc/src/sgml/ref/pg_rewind.sgml | 14 ++
src/bin/pg_rewind/parsexlog.c | 4 +-
src/bin/pg_rewind/pg_rewind.c | 93 +++++++++++-
src/bin/pg_rewind/pg_rewind.h | 1 +
src/test/modules/test_checksums/meson.build | 2 +
.../test_checksums/t/021_rewind_state.pl | 133 +++++++++++++++++
.../t/022_rewind_standby_target.pl | 141 ++++++++++++++++++
7 files changed, 386 insertions(+), 2 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/021_rewind_state.pl
create mode 100644 src/test/modules/test_checksums/t/022_rewind_standby_target.pl
diff --git a/doc/src/sgml/ref/pg_rewind.sgml b/doc/src/sgml/ref/pg_rewind.sgml
index b95e3868e9d..ac3d0c9328f 100644
--- a/doc/src/sgml/ref/pg_rewind.sgml
+++ b/doc/src/sgml/ref/pg_rewind.sgml
@@ -124,6 +124,20 @@ PostgreSQL documentation
<literal>on</literal>, but is enabled by default.
</para>
+ <para>
+ The data checksum states of the source and target server must be
+ compatible. <application>pg_rewind</application> refuses to run if an
+ online checksum state transition was interrupted on either server, or
+ if data checksums are enabled on the target server, or were enabled at
+ the point of divergence, while the source server runs without them: the
+ rewound server would fail checksum verification on the blocks copied
+ from the source. The opposite combination is allowed with a warning;
+ the rewound server keeps data checksums disabled unless the source
+ server enabled them online after the point of divergence. See
+ <xref linkend="app-pgchecksums"/> for how to change the state
+ consistently across a replication setup.
+ </para>
+
<warning>
<title>Warning: Failures While Rewinding</title>
<para>
diff --git a/src/bin/pg_rewind/parsexlog.c b/src/bin/pg_rewind/parsexlog.c
index 023e23b063c..6e87b00f8c2 100644
--- a/src/bin/pg_rewind/parsexlog.c
+++ b/src/bin/pg_rewind/parsexlog.c
@@ -167,7 +167,8 @@ readOneRecord(const char *datadir, XLogRecPtr ptr, int tliIndex,
void
findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
- XLogRecPtr *lastchkptredo, const char *restoreCommand)
+ XLogRecPtr *lastchkptredo, uint32 *lastchkptdatachecksums,
+ const char *restoreCommand)
{
/* Walk backwards, starting from the given record */
XLogRecord *record;
@@ -255,6 +256,7 @@ findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
*lastchkptrec = searchptr;
*lastchkpttli = checkPoint.ThisTimeLineID;
*lastchkptredo = checkPoint.redo;
+ *lastchkptdatachecksums = checkPoint.dataChecksumState;
break;
}
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 936469ea5f8..8c1357c718f 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -144,6 +144,7 @@ main(int argc, char **argv)
XLogRecPtr chkptrec;
TimeLineID chkpttli;
XLogRecPtr chkptredo;
+ uint32 chkptdatachecksums;
TimeLineID source_tli;
TimeLineID target_tli;
XLogRecPtr target_wal_endrec;
@@ -471,10 +472,51 @@ main(int argc, char **argv)
keepwal_init();
findLastCheckpoint(datadir_target, divergerec, lastcommontliIndex,
- &chkptrec, &chkpttli, &chkptredo, restore_command);
+ &chkptrec, &chkpttli, &chkptredo, &chkptdatachecksums,
+ restore_command);
pg_log_info("rewinding from last common checkpoint at %X/%08X on timeline %u",
LSN_FORMAT_ARGS(chkptrec), chkpttli);
+ /*
+ * Replay on the rewound server resumes from the last common checkpoint
+ * and adopts the data checksum state recorded there, not the state in the
+ * control file installed from the source. If checksums were enabled at
+ * the divergence point but the source runs without them, the rewound
+ * server would verify checksums while the blocks copied from the source
+ * have none. The control files cannot reveal this: an offline disable on
+ * the source after the divergence leaves both of them saying "off".
+ *
+ * Only the fully enabled state needs checking. An in-progress state at
+ * the divergence point means a transition was still running there, and
+ * all of its page rewrites are logged after that point, so replay brings
+ * the rewound server to whatever state the copied WAL ends in.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the source server or recreate the target server from a base backup.");
+ exit(1);
+ }
+
+ /*
+ * The same comparison is needed against the target: every block the
+ * rewind does not copy keeps the target's content. The control files
+ * cannot reveal this case either, in the other direction: a standby's
+ * checkpoints are written by its upstream primary, so a standby whose
+ * checksums were disabled offline still has "on" checkpoints in its WAL,
+ * while both control files may agree.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_target.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the target server");
+ pg_log_error_detail("Blocks kept from the target would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the target server with pg_checksums or recreate the target server from a base backup.");
+ exit(1);
+ }
+
/* Initialize the hash table to track the status of each file */
filehash_init();
@@ -784,6 +826,55 @@ sanityChecks(void)
pg_fatal("target server needs to use either data checksums or \"wal_log_hints = on\"");
}
+ /*
+ * The rewound target keeps its control file fields from the source, but
+ * every block the rewind does not copy keeps the target's content. If
+ * checksums are enabled on the target and disabled on the source, the
+ * result would claim enabled checksums while the blocks copied from the
+ * source have none, and the target would fail checksum verification as
+ * soon as it reads them. Refuse that combination, and refuse an
+ * interrupted online transition on either side, same as pg_checksums.
+ *
+ * The opposite mismatch is allowed: recovery under the backup label
+ * written by pg_rewind adopts the checksum state as of the divergence
+ * point, so a target whose own checkpoints carry no checksums stays
+ * without them, or converges through the replayed WAL if the source
+ * enabled them online. Only warn about it, so that an offline change on
+ * the source is not overlooked.
+ *
+ * That reasoning does not hold when the target is a standby, whose
+ * checkpoints were written by its upstream primary; see the checks after
+ * findLastCheckpoint().
+ */
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted on the target server");
+ pg_log_error_hint("Start the server and shut it down cleanly to reset the state, then retry; on a standby, let replication complete the transition first.");
+ exit(1);
+ }
+ if (ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition is incomplete on the source server");
+ pg_log_error_hint("Let the transition complete, or reset the state with a clean restart, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_OFF)
+ {
+ pg_log_error("data checksums are enabled on the target server but disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, and the target would fail checksum verification after the rewind.");
+ pg_log_error_hint("Either disable data checksums on the target server or enable them on the source server, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_OFF &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_warning("data checksums are disabled on the target server but enabled on the source server");
+ pg_log_warning_detail("The rewound server will keep data checksums disabled unless the source server enabled them online after the point of divergence.");
+ }
+
/*
* Target cluster better not be running. This doesn't guard against
* someone starting the cluster concurrently. Also, this is probably more
diff --git a/src/bin/pg_rewind/pg_rewind.h b/src/bin/pg_rewind/pg_rewind.h
index 9a981f7f246..9187c3bf72f 100644
--- a/src/bin/pg_rewind/pg_rewind.h
+++ b/src/bin/pg_rewind/pg_rewind.h
@@ -39,6 +39,7 @@ extern void findLastCheckpoint(const char *datadir, XLogRecPtr forkptr,
int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
XLogRecPtr *lastchkptredo,
+ uint32 *lastchkptdatachecksums,
const char *restoreCommand);
extern XLogRecPtr readOneRecord(const char *datadir, XLogRecPtr ptr,
int tliIndex, const char *restoreCommand);
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index 70870715976..a4f25bac7a5 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -44,6 +44,8 @@ tests += {
't/018_enable_crash_windows.pl',
't/019_standby_shutdown_catchup.pl',
't/020_cascade_divergence.pl',
+ 't/021_rewind_state.pl',
+ 't/022_rewind_standby_target.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/021_rewind_state.pl b/src/test/modules/test_checksums/t/021_rewind_state.pl
new file mode 100644
index 00000000000..b7adc968c4d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/021_rewind_state.pl
@@ -0,0 +1,133 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind checks the data checksum states of source and target.
+# Checksums enabled on the target, or at the point of divergence, with
+# a source running without them are refused: the rewound server would
+# verify checksums on blocks copied from a source that has none. The
+# opposite mismatch only warns; replay keeps the target's own state.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# Scenarios 1 and 2: checksums on from initdb. wal_log_hints keeps the
+# target eligible for pg_rewind after checksums are disabled on it.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b);
+
+# Failover to B, and divergence on A.
+$node_b->promote;
+$node_b->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_a->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_a->stop('fast');
+
+# Offline disable on the source only: enabled target, disabled source.
+$node_b->stop;
+$node_b->checksum_disable_offline;
+$node_b->start;
+test_checksum_state($node_b, 'off');
+
+my @rewind_cmd = (
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'));
+
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums are enabled on the target server but disabled on the source server/,
+ 'refuses an enabled target with a disabled source');
+
+# The refusal happens before any modification: the target still starts
+# on its own timeline.
+$node_a->start;
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'target untouched by the refused rewind');
+$node_a->stop('fast');
+
+# Scenario 2: disabling the target too makes the control files match,
+# but checksums were still enabled at the point of divergence, which is
+# the state replay on the rewound server would resume with.
+$node_a->checksum_disable_offline;
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums were enabled at the point of divergence but are disabled on the source server/,
+ 'refuses when checksums were enabled at the point of divergence');
+
+# Scenario 3: fresh pair without checksums; offline enable on the
+# source only. Allowed with a warning, and the rewound server keeps
+# checksums disabled.
+my $node_c = PostgreSQL::Test::Cluster->new('node_c');
+$node_c->init(allows_streaming => 1, no_data_checksums => 1);
+$node_c->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_c->start;
+$node_c->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_c->backup('backup');
+my $node_d = PostgreSQL::Test::Cluster->new('node_d');
+$node_d->init_from_backup($node_c, 'backup', has_streaming => 1);
+$node_d->start;
+$node_c->wait_for_catchup($node_d);
+
+$node_d->promote;
+$node_d->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_c->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_c->stop('fast');
+
+$node_d->stop;
+$node_d->checksum_enable_offline;
+$node_d->start;
+test_checksum_state($node_d, 'on');
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_c->data_dir,
+ '--source-server' => $node_d->connstr('postgres'),
+ ]);
+like(
+ $stderr,
+ qr/data checksums are disabled on the target server but enabled on the source server/,
+ 'warns for a disabled target with an enabled source');
+like($stderr, qr/Done!/, 'rewind completed despite the warning');
+
+# The rewound server follows D and keeps its own state.
+$node_c->append_conf('postgresql.conf', 'port = ' . $node_c->port);
+$node_c->enable_streaming($node_d);
+$node_c->set_standby_mode;
+$node_c->start;
+$node_d->wait_for_catchup($node_c);
+is($node_c->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+test_checksum_state($node_c, 'off');
+
+$node_c->stop;
+$node_d->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/022_rewind_standby_target.pl b/src/test/modules/test_checksums/t/022_rewind_standby_target.pl
new file mode 100644
index 00000000000..3f5a3c6be8b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/022_rewind_standby_target.pl
@@ -0,0 +1,141 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_rewind compares the data checksum state at the point of
+# divergence against the target as well as the source.
+#
+# A standby's checkpoints are written by its upstream primary, so a standby
+# whose checksums were disabled offline still has "on" checkpoints in its
+# WAL. Rewinding it onto a promoted sibling passes every control file
+# check (target off + source on only warns), yet recovery after the rewind
+# adopts the divergence checkpoint's state and the rewound server would
+# verify checksums over the pages it wrote while it was locally "off".
+# pg_rewind must refuse; after an offline enable on the target, which
+# rewrites all of its pages, the rewind goes through.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t0 AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+
+my $lagging = PostgreSQL::Test::Cluster->new('lagging');
+$lagging->init_from_backup($primary, 'backup', has_streaming => 1);
+$lagging->start;
+
+my $failover = PostgreSQL::Test::Cluster->new('failover');
+$failover->init_from_backup($primary, 'backup', has_streaming => 1);
+$failover->start;
+
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+
+test_checksum_state($primary, 'on');
+test_checksum_state($lagging, 'on');
+test_checksum_state($failover, 'on');
+
+# Offline-disable checksums on the lagging standby only. This is a
+# divergence 012_offline_standby.pl declares survivable: the standby keeps
+# its own state, warns, and stays readable.
+$lagging->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $lagging->data_dir);
+$lagging->start;
+test_checksum_state($lagging, 'off');
+
+# Give the lagging standby plenty of pages to write while it is locally
+# "off", and force a restartpoint so they reach disk without checksums.
+# The other standby replays the same WAL with checksums on.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+$lagging->safe_psql('postgres', "CHECKPOINT;");
+
+# Take the "failover" standby out of the picture. It is promoted from
+# this position later, so everything replayed from here on is the part of
+# the WAL that diverges and that pg_rewind will roll back. t1 is already
+# replayed everywhere, so its blocks are *not* rolled back.
+$failover->stop;
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$primary->wait_for_catchup($lagging);
+
+# Failover: promote the other standby, which forks the timeline behind the
+# replay position of the lagging standby.
+$primary->stop('fast');
+$failover->start;
+$failover->promote;
+$failover->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$failover->safe_psql('postgres', "CHECKPOINT;");
+
+# Re-attaching the lagging standby with pg_rewind must be refused: the
+# divergence checkpoint says "on" while the target is "off", so the rewound
+# server would verify checksums over the blocks it keeps.
+$lagging->stop;
+
+command_fails_like(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ],
+ qr/data checksums were enabled at the point of divergence but are disabled on the target server/,
+ 'pg_rewind refuses a target whose divergence checkpoint has checksums');
+
+# Enabling checksums offline on the target rewrites all of its pages, after
+# which the rewind is safe.
+system_or_bail('pg_checksums', '--enable', '--pgdata', $lagging->data_dir);
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ]);
+like($stderr, qr/Done!/, 'pg_rewind completes after the offline enable');
+
+$lagging->append_conf('postgresql.conf', 'port = ' . $lagging->port);
+$lagging->enable_streaming($failover);
+$lagging->set_standby_mode;
+$lagging->start;
+
+my ($rc, $out, $err) =
+ $lagging->psql('postgres',
+ "SELECT setting FROM pg_settings " . "WHERE name = 'data_checksums';");
+is($rc, 0, 'rewound standby accepts connections') or diag("stderr: $err");
+is($out, 'on', 'rewound standby resumes with checksums on');
+
+($rc, $out, $err) = $lagging->psql('postgres', "SELECT count(*) FROM t1;");
+is($rc, 0, 'blocks kept from the target stay readable')
+ or diag("stderr: $err");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($lagging->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the rewound standby');
+
+$lagging->stop('immediate');
+$failover->stop;
+done_testing();
--
2.55.0
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-31 11:38 Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
parent: Zsolt Parragi <zsolt.parragi@percona.com>
0 siblings, 1 reply; 43+ messages in thread
From: Bertrand Drouvot @ 2026-08-31 11:38 UTC (permalink / raw)
To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; pgsql-hackers@lists.postgresql.org
Hi,
On Mon, Aug 31, 2026 at 11:55:01AM +0100, Zsolt Parragi wrote:
> > Yeah, I see your point that the underlying issue is the non WAL logged change.
> > My concern was not about a pg_control change in itself, but that addressing this
> > case might require further changes to the current pg_control design, which adds
> > risk this close to RC1.
>
> I don't think we can get this completely right without both wal and
> pg_control changes (and possibly some extra connection-time
> communication between the standby and primary to provide early
> reporting instead of delayed in the most common cases)
> If we want to move toward the error direction in pg20 (which we
> should) we will have to log that an offline change happened.
>
> I attached v6 which for now documents the current behavior as-is, and
> other than that I reorganized the test files.
Thanks! I did not look in details but it looks like that pg_upgrade --check against
a running source cluster with checksums enabled will still fail (finding === 2 in
[0], tested with the test shared in [1]).
[0]: https://postgr.es/m/apUL3N4IE934qJ08%40bdtpg
[1]: https://postgr.es/m/apU4/hmRv/4gv20W%40bdtpg
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-31 13:10 Zsolt Parragi <zsolt.parragi@percona.com>
parent: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 1 reply; 43+ messages in thread
From: Zsolt Parragi @ 2026-08-31 13:10 UTC (permalink / raw)
To: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; pgsql-hackers@lists.postgresql.org
> Thanks! I did not look in details but it looks like that pg_upgrade --check against
> a running source cluster with checksums enabled will still fail (finding === 2 in
> [0], tested with the test shared in [1]).
Yes, I missed that in the previous version, v7 fixes it.
Attachments:
[application/octet-stream] v7-0005-doc-Explain-how-offline-and-online-checksum-chang.patch (4.2K, ../../CAN4CZFMFcgfgJ99RYhax-T+YJH=CWLy6ZGANMxf77SrSirw3VQ@mail.gmail.com/2-v7-0005-doc-Explain-how-offline-and-online-checksum-chang.patch)
download | inline diff:
From ce65e9f5e674d85987fbe1743d39bd0637eeb7eb Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Mon, 31 Aug 2026 08:40:44 +0000
Subject: [PATCH v7 5/5] doc: Explain how offline and online checksum changes
interact
The lockstep procedure for offline checksum changes in a replication
setup was documented without preconditions. It is not sufficient on
its own: an offline change is recorded only in the control file and
has no ordering against WAL the node has not replayed yet, so a node
stopped before replaying an online state transition applies it on
restart and overrides the offline change. The nodes then silently
diverge even though the change was applied to all of them while
stopped.
Document the two durability models next to each other in wal.sgml,
and extend the pg_checksums notes: require standbys to have replayed
all WAL of their upstream node before stopping them, show how to
check that, and recommend not mixing online and offline changes.
---
doc/src/sgml/ref/pg_checksums.sgml | 25 ++++++++++++++++++++++---
doc/src/sgml/wal.sgml | 13 +++++++++++++
2 files changed, 35 insertions(+), 3 deletions(-)
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index 000940e7cb7..b337950be3d 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -248,9 +248,28 @@ PostgreSQL documentation
directory; the new state does not propagate over replication. In a
replication setup the same change must be applied to every node: stop all
nodes, run <application>pg_checksums</application> on each of them, and
- only then restart them. Tools that copy relation file blocks directly
- between nodes, such as <xref linkend="app-pgrewind"/>, likewise require
- both nodes to be in the same data checksum state.
+ only then restart them. Before stopping a standby, make sure it has
+ replayed all WAL of its upstream node, for example by stopping the
+ primary first and comparing
+ <function>pg_last_wal_replay_lsn()</function> with
+ <function>pg_last_wal_receive_lsn()</function> on the standby. Tools
+ that copy relation file blocks directly between nodes, such as
+ <xref linkend="app-pgrewind"/>, likewise require both nodes to be in
+ the same data checksum state.
+ </para>
+ <para>
+ The replay requirement exists because an offline change is recorded
+ only in the control file and has no defined ordering against WAL the
+ node has not replayed yet; see
+ <xref linkend="checksums-offline-enable-disable"/>. A node stopped
+ before replaying an online checksum state change applies that change
+ when it is restarted, overriding the offline change, and the states of
+ the nodes silently diverge until a later checkpoint record triggers
+ the warning described below. Because of this it is best not to mix
+ the two mechanisms: change the state of a replication setup either
+ with the offline procedure above or with an online transition, and
+ make sure the previous change has reached every node before starting
+ the next one.
</para>
<para>
If the change is applied inconsistently, each node keeps its own state,
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index db18a8c516e..d88a832c2a8 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -317,6 +317,19 @@
verify checksums, on an offline cluster.
</para>
+ <para>
+ An offline change provides durability differently from an
+ <link linkend="checksums-online-enable-disable">online change</link>.
+ An online transition is WAL-logged: it is ordered against all other
+ WAL records, it is replayed after a crash, and it propagates to
+ standbys. An offline change is recorded only in the cluster's
+ control file: it writes no WAL, it is invisible to replication, and
+ it has no defined ordering against WAL the node has not replayed
+ yet. When a node later replays WAL that contains an online checksum
+ state change, that change takes effect on the node even if it was
+ written before the offline change was made.
+ </para>
+
<para>
An offline change only affects the data directory it is run on; the
new state does not propagate over replication. In a replication setup
--
2.55.0
[application/octet-stream] v7-0002-pg_checksums-Refuse-interrupted-transitions-note-.patch (7.7K, ../../CAN4CZFMFcgfgJ99RYhax-T+YJH=CWLy6ZGANMxf77SrSirw3VQ@mail.gmail.com/3-v7-0002-pg_checksums-Refuse-interrupted-transitions-note-.patch)
download | inline diff:
From 816cbbe9c619bd6c5a6660fa13f08b40f84fc6cd Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:19 +0000
Subject: [PATCH v7 2/5] pg_checksums: Refuse interrupted transitions, note
change is local
An inprogress-on or inprogress-off control file means an online
transition was cut short; running the offline tool on top of it mixes
two procedures. Refuse it in every mode, with a hint pointing at the
start-stop cycle (or, on a standby, at letting replication finish)
that resets the state. Also print that the change applies to one
data directory only, with standby-aware wording, since in a
replication setup the same change must be made on every node.
On a primary, inprogress-on cannot survive a graceful stop: the
checksums launcher resolves it from its exit cleanup, and a crashed
primary is already rejected by the existing "cluster must be shut
down" check. inprogress-off can, however: it is set by the backend
running pg_disable_data_checksums(), and a fast shutdown arriving
between its two barriers leaves it behind in a cleanly shut down
control file, where the next start-stop cycle resolves it at end of
recovery. A standby can be stopped with either state, since it has
no launcher and only carries forward whatever state the last replayed
WAL record left it in, with a restartpoint persisting that as-is.
The standby is also the deterministic way to reach the new guard,
which is why the test coverage uses one; the primary window would
need an injection point between the two barriers.
The pg_checksums page said the tool still processes all relation files
regardless of an interrupted online transition; describe the refusal
instead.
---
doc/src/sgml/ref/pg_checksums.sgml | 11 ++++---
src/bin/pg_checksums/pg_checksums.c | 21 +++++++++++++
.../test_checksums/t/012_offline_standby.pl | 31 ++++++++++++++-----
3 files changed, 51 insertions(+), 12 deletions(-)
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index bf07da09754..000940e7cb7 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -49,11 +49,12 @@ PostgreSQL documentation
</para>
<para>
- When enabling checksums with <application>pg_checksums</application>, if
- checksums were in the process of being enabled using
- <xref linkend="checksums-online-enable-disable"/> when the cluster was shut
- down, <application>pg_checksums</application> will still process all
- relation files regardless of the progress of online checksum processing.
+ If checksums were in the process of being enabled or disabled using
+ <xref linkend="checksums-online-enable-disable"/> when the cluster was
+ shut down, the control file still records that interrupted state, and
+ <application>pg_checksums</application> refuses to run in any mode.
+ Start the cluster and shut it down cleanly to reset the state, then
+ retry; on a standby, let replication complete the transition first.
</para>
<para>
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 20a99112838..a49b6f368b7 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -586,6 +586,21 @@ main(int argc, char *argv[])
ControlFile->state != DB_SHUTDOWNED_IN_RECOVERY)
pg_fatal("cluster must be shut down");
+ /*
+ * An inprogress state means an online transition was cut short. A
+ * standby stopped mid-transition carries either state; a cleanly shut
+ * down primary can still carry inprogress-off, which a fast shutdown
+ * during pg_disable_data_checksums() leaves behind, while inprogress-on
+ * is always resolved by the launcher's exit cleanup.
+ */
+ if (ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted");
+ pg_log_error_hint("Start and cleanly shut down the cluster once to reset the data checksum state, then retry. On a standby, let replication complete the transition first.");
+ exit(1);
+ }
+
if (ControlFile->data_checksum_version != PG_DATA_CHECKSUM_VERSION &&
mode == PG_MODE_CHECK)
pg_fatal("data checksums are not enabled in cluster");
@@ -673,6 +688,12 @@ main(int argc, char *argv[])
printf(_("Checksums enabled in cluster\n"));
else
printf(_("Checksums disabled in cluster\n"));
+
+ printf(_("This change applies to this data directory only.\n"));
+ if (ControlFile->state == DB_SHUTDOWNED_IN_RECOVERY)
+ printf(_("This node appears to be a standby; apply the same change to the primary and all other standbys.\n"));
+ else
+ printf(_("In a replication setup, apply the same change to every node while all are stopped, before restarting any of them.\n"));
}
return 0;
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
index 241a1282ae9..5efe11f993e 100644
--- a/src/test/modules/test_checksums/t/012_offline_standby.pl
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -149,7 +149,10 @@ $newnode->stop('immediate');
# Converge the cluster: enable offline on the standby too.
$standby->stop;
-$standby->checksum_enable_offline;
+command_checks_all(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ 0, [qr/appears to be a standby/],
+ [], 'standby-role notice on offline enable');
$standby->start;
test_checksum_state($standby, 'on');
$primary->wait_for_catchup($standby);
@@ -217,11 +220,11 @@ unlike(
);
# Scenario 4: a standby stopped while replaying an interrupted online
-# transition keeps the interrupted state in its own control file. A
-# primary is never caught this way, as its checksums launcher resolves
-# inprogress-on back to off from its exit cleanup. A standby has no
-# launcher; it carries forward whatever the last replayed record left
-# it in.
+# transition keeps the interrupted state in its own control file, and
+# pg_checksums must refuse to touch it. A primary is never caught this
+# way, as its checksums launcher resolves inprogress-on back to off from
+# its exit cleanup. A standby has no launcher; it carries forward
+# whatever the last replayed record left it in.
# Block an online enable on the primary at inprogress-on with a
# blocking temp table, same trick as in 004_offline.pl.
@@ -235,6 +238,20 @@ wait_for_checksum_state($standby, 'inprogress-on');
# Stop the standby cleanly; its restartpoint persists inprogress-on to
# its own control file, since nothing on a standby resolves it away.
$standby->stop;
+
+command_fails_like(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --enable refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --check refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--disable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --disable refuses a standby stopped mid-transition');
+
$standby->start;
wait_for_checksum_state($standby, 'inprogress-on');
@@ -257,7 +274,7 @@ wait_for_checksum_state($primary, 'on');
$primary->wait_for_catchup($standby);
wait_for_checksum_state($standby, 'on');
-is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
'10001', 'standby readable once the transition completes');
# Scenario 4 continued: the backup taken mid-transition must also
--
2.55.0
[application/octet-stream] v7-0003-pg_rewind-Check-the-data-checksum-states-of-sourc.patch (21.6K, ../../CAN4CZFMFcgfgJ99RYhax-T+YJH=CWLy6ZGANMxf77SrSirw3VQ@mail.gmail.com/4-v7-0003-pg_rewind-Check-the-data-checksum-states-of-sourc.patch)
download | inline diff:
From a0d15a29ea6df24bb04be620a7bf2207af7330a4 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 18:40:22 +0000
Subject: [PATCH v7 3/5] pg_rewind: Check the data checksum states of source
and target
pg_rewind installs the source's control file on the target, but every
block it does not copy keeps the target's content. With checksums
enabled on the target and disabled on the source, the rewound server
claims enabled checksums while the blocks copied from the source have
none, and fails checksum verification as soon as it reads them; in the
worst case every connection attempt dies on an unverifiable catalog
page.
Comparing the control files is not enough. Replay on the rewound
server resumes from the last common checkpoint and adopts the data
checksum state recorded there, so an offline disable on the source
after the divergence leaves both control files saying "off" while the
rewound server still resumes with verification enabled. Check the
state carried by the divergence checkpoint as well, which pg_rewind
already reads.
Refuse both cases, and an interrupted online transition on either
side, same as pg_checksums does. The opposite mismatch, checksums on
the source only, is allowed with a warning: recovery under the backup
label adopts the divergence state, so the rewound server keeps
checksums disabled (or converges through the replayed WAL if the
source enabled them online), and no page can fail verification.
---
doc/src/sgml/ref/pg_rewind.sgml | 14 ++
src/bin/pg_rewind/parsexlog.c | 4 +-
src/bin/pg_rewind/pg_rewind.c | 93 +++++++++++-
src/bin/pg_rewind/pg_rewind.h | 1 +
src/test/modules/test_checksums/meson.build | 2 +
.../test_checksums/t/021_rewind_state.pl | 133 +++++++++++++++++
.../t/022_rewind_standby_target.pl | 141 ++++++++++++++++++
7 files changed, 386 insertions(+), 2 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/021_rewind_state.pl
create mode 100644 src/test/modules/test_checksums/t/022_rewind_standby_target.pl
diff --git a/doc/src/sgml/ref/pg_rewind.sgml b/doc/src/sgml/ref/pg_rewind.sgml
index b95e3868e9d..ac3d0c9328f 100644
--- a/doc/src/sgml/ref/pg_rewind.sgml
+++ b/doc/src/sgml/ref/pg_rewind.sgml
@@ -124,6 +124,20 @@ PostgreSQL documentation
<literal>on</literal>, but is enabled by default.
</para>
+ <para>
+ The data checksum states of the source and target server must be
+ compatible. <application>pg_rewind</application> refuses to run if an
+ online checksum state transition was interrupted on either server, or
+ if data checksums are enabled on the target server, or were enabled at
+ the point of divergence, while the source server runs without them: the
+ rewound server would fail checksum verification on the blocks copied
+ from the source. The opposite combination is allowed with a warning;
+ the rewound server keeps data checksums disabled unless the source
+ server enabled them online after the point of divergence. See
+ <xref linkend="app-pgchecksums"/> for how to change the state
+ consistently across a replication setup.
+ </para>
+
<warning>
<title>Warning: Failures While Rewinding</title>
<para>
diff --git a/src/bin/pg_rewind/parsexlog.c b/src/bin/pg_rewind/parsexlog.c
index 023e23b063c..6e87b00f8c2 100644
--- a/src/bin/pg_rewind/parsexlog.c
+++ b/src/bin/pg_rewind/parsexlog.c
@@ -167,7 +167,8 @@ readOneRecord(const char *datadir, XLogRecPtr ptr, int tliIndex,
void
findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
- XLogRecPtr *lastchkptredo, const char *restoreCommand)
+ XLogRecPtr *lastchkptredo, uint32 *lastchkptdatachecksums,
+ const char *restoreCommand)
{
/* Walk backwards, starting from the given record */
XLogRecord *record;
@@ -255,6 +256,7 @@ findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
*lastchkptrec = searchptr;
*lastchkpttli = checkPoint.ThisTimeLineID;
*lastchkptredo = checkPoint.redo;
+ *lastchkptdatachecksums = checkPoint.dataChecksumState;
break;
}
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 936469ea5f8..8c1357c718f 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -144,6 +144,7 @@ main(int argc, char **argv)
XLogRecPtr chkptrec;
TimeLineID chkpttli;
XLogRecPtr chkptredo;
+ uint32 chkptdatachecksums;
TimeLineID source_tli;
TimeLineID target_tli;
XLogRecPtr target_wal_endrec;
@@ -471,10 +472,51 @@ main(int argc, char **argv)
keepwal_init();
findLastCheckpoint(datadir_target, divergerec, lastcommontliIndex,
- &chkptrec, &chkpttli, &chkptredo, restore_command);
+ &chkptrec, &chkpttli, &chkptredo, &chkptdatachecksums,
+ restore_command);
pg_log_info("rewinding from last common checkpoint at %X/%08X on timeline %u",
LSN_FORMAT_ARGS(chkptrec), chkpttli);
+ /*
+ * Replay on the rewound server resumes from the last common checkpoint
+ * and adopts the data checksum state recorded there, not the state in the
+ * control file installed from the source. If checksums were enabled at
+ * the divergence point but the source runs without them, the rewound
+ * server would verify checksums while the blocks copied from the source
+ * have none. The control files cannot reveal this: an offline disable on
+ * the source after the divergence leaves both of them saying "off".
+ *
+ * Only the fully enabled state needs checking. An in-progress state at
+ * the divergence point means a transition was still running there, and
+ * all of its page rewrites are logged after that point, so replay brings
+ * the rewound server to whatever state the copied WAL ends in.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the source server or recreate the target server from a base backup.");
+ exit(1);
+ }
+
+ /*
+ * The same comparison is needed against the target: every block the
+ * rewind does not copy keeps the target's content. The control files
+ * cannot reveal this case either, in the other direction: a standby's
+ * checkpoints are written by its upstream primary, so a standby whose
+ * checksums were disabled offline still has "on" checkpoints in its WAL,
+ * while both control files may agree.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_target.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the target server");
+ pg_log_error_detail("Blocks kept from the target would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the target server with pg_checksums or recreate the target server from a base backup.");
+ exit(1);
+ }
+
/* Initialize the hash table to track the status of each file */
filehash_init();
@@ -784,6 +826,55 @@ sanityChecks(void)
pg_fatal("target server needs to use either data checksums or \"wal_log_hints = on\"");
}
+ /*
+ * The rewound target keeps its control file fields from the source, but
+ * every block the rewind does not copy keeps the target's content. If
+ * checksums are enabled on the target and disabled on the source, the
+ * result would claim enabled checksums while the blocks copied from the
+ * source have none, and the target would fail checksum verification as
+ * soon as it reads them. Refuse that combination, and refuse an
+ * interrupted online transition on either side, same as pg_checksums.
+ *
+ * The opposite mismatch is allowed: recovery under the backup label
+ * written by pg_rewind adopts the checksum state as of the divergence
+ * point, so a target whose own checkpoints carry no checksums stays
+ * without them, or converges through the replayed WAL if the source
+ * enabled them online. Only warn about it, so that an offline change on
+ * the source is not overlooked.
+ *
+ * That reasoning does not hold when the target is a standby, whose
+ * checkpoints were written by its upstream primary; see the checks after
+ * findLastCheckpoint().
+ */
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted on the target server");
+ pg_log_error_hint("Start the server and shut it down cleanly to reset the state, then retry; on a standby, let replication complete the transition first.");
+ exit(1);
+ }
+ if (ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition is incomplete on the source server");
+ pg_log_error_hint("Let the transition complete, or reset the state with a clean restart, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_OFF)
+ {
+ pg_log_error("data checksums are enabled on the target server but disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, and the target would fail checksum verification after the rewind.");
+ pg_log_error_hint("Either disable data checksums on the target server or enable them on the source server, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_OFF &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_warning("data checksums are disabled on the target server but enabled on the source server");
+ pg_log_warning_detail("The rewound server will keep data checksums disabled unless the source server enabled them online after the point of divergence.");
+ }
+
/*
* Target cluster better not be running. This doesn't guard against
* someone starting the cluster concurrently. Also, this is probably more
diff --git a/src/bin/pg_rewind/pg_rewind.h b/src/bin/pg_rewind/pg_rewind.h
index 9a981f7f246..9187c3bf72f 100644
--- a/src/bin/pg_rewind/pg_rewind.h
+++ b/src/bin/pg_rewind/pg_rewind.h
@@ -39,6 +39,7 @@ extern void findLastCheckpoint(const char *datadir, XLogRecPtr forkptr,
int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
XLogRecPtr *lastchkptredo,
+ uint32 *lastchkptdatachecksums,
const char *restoreCommand);
extern XLogRecPtr readOneRecord(const char *datadir, XLogRecPtr ptr,
int tliIndex, const char *restoreCommand);
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index 70870715976..a4f25bac7a5 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -44,6 +44,8 @@ tests += {
't/018_enable_crash_windows.pl',
't/019_standby_shutdown_catchup.pl',
't/020_cascade_divergence.pl',
+ 't/021_rewind_state.pl',
+ 't/022_rewind_standby_target.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/021_rewind_state.pl b/src/test/modules/test_checksums/t/021_rewind_state.pl
new file mode 100644
index 00000000000..b7adc968c4d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/021_rewind_state.pl
@@ -0,0 +1,133 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind checks the data checksum states of source and target.
+# Checksums enabled on the target, or at the point of divergence, with
+# a source running without them are refused: the rewound server would
+# verify checksums on blocks copied from a source that has none. The
+# opposite mismatch only warns; replay keeps the target's own state.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# Scenarios 1 and 2: checksums on from initdb. wal_log_hints keeps the
+# target eligible for pg_rewind after checksums are disabled on it.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b);
+
+# Failover to B, and divergence on A.
+$node_b->promote;
+$node_b->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_a->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_a->stop('fast');
+
+# Offline disable on the source only: enabled target, disabled source.
+$node_b->stop;
+$node_b->checksum_disable_offline;
+$node_b->start;
+test_checksum_state($node_b, 'off');
+
+my @rewind_cmd = (
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'));
+
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums are enabled on the target server but disabled on the source server/,
+ 'refuses an enabled target with a disabled source');
+
+# The refusal happens before any modification: the target still starts
+# on its own timeline.
+$node_a->start;
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'target untouched by the refused rewind');
+$node_a->stop('fast');
+
+# Scenario 2: disabling the target too makes the control files match,
+# but checksums were still enabled at the point of divergence, which is
+# the state replay on the rewound server would resume with.
+$node_a->checksum_disable_offline;
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums were enabled at the point of divergence but are disabled on the source server/,
+ 'refuses when checksums were enabled at the point of divergence');
+
+# Scenario 3: fresh pair without checksums; offline enable on the
+# source only. Allowed with a warning, and the rewound server keeps
+# checksums disabled.
+my $node_c = PostgreSQL::Test::Cluster->new('node_c');
+$node_c->init(allows_streaming => 1, no_data_checksums => 1);
+$node_c->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_c->start;
+$node_c->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_c->backup('backup');
+my $node_d = PostgreSQL::Test::Cluster->new('node_d');
+$node_d->init_from_backup($node_c, 'backup', has_streaming => 1);
+$node_d->start;
+$node_c->wait_for_catchup($node_d);
+
+$node_d->promote;
+$node_d->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_c->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_c->stop('fast');
+
+$node_d->stop;
+$node_d->checksum_enable_offline;
+$node_d->start;
+test_checksum_state($node_d, 'on');
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_c->data_dir,
+ '--source-server' => $node_d->connstr('postgres'),
+ ]);
+like(
+ $stderr,
+ qr/data checksums are disabled on the target server but enabled on the source server/,
+ 'warns for a disabled target with an enabled source');
+like($stderr, qr/Done!/, 'rewind completed despite the warning');
+
+# The rewound server follows D and keeps its own state.
+$node_c->append_conf('postgresql.conf', 'port = ' . $node_c->port);
+$node_c->enable_streaming($node_d);
+$node_c->set_standby_mode;
+$node_c->start;
+$node_d->wait_for_catchup($node_c);
+is($node_c->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+test_checksum_state($node_c, 'off');
+
+$node_c->stop;
+$node_d->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/022_rewind_standby_target.pl b/src/test/modules/test_checksums/t/022_rewind_standby_target.pl
new file mode 100644
index 00000000000..3f5a3c6be8b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/022_rewind_standby_target.pl
@@ -0,0 +1,141 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_rewind compares the data checksum state at the point of
+# divergence against the target as well as the source.
+#
+# A standby's checkpoints are written by its upstream primary, so a standby
+# whose checksums were disabled offline still has "on" checkpoints in its
+# WAL. Rewinding it onto a promoted sibling passes every control file
+# check (target off + source on only warns), yet recovery after the rewind
+# adopts the divergence checkpoint's state and the rewound server would
+# verify checksums over the pages it wrote while it was locally "off".
+# pg_rewind must refuse; after an offline enable on the target, which
+# rewrites all of its pages, the rewind goes through.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t0 AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+
+my $lagging = PostgreSQL::Test::Cluster->new('lagging');
+$lagging->init_from_backup($primary, 'backup', has_streaming => 1);
+$lagging->start;
+
+my $failover = PostgreSQL::Test::Cluster->new('failover');
+$failover->init_from_backup($primary, 'backup', has_streaming => 1);
+$failover->start;
+
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+
+test_checksum_state($primary, 'on');
+test_checksum_state($lagging, 'on');
+test_checksum_state($failover, 'on');
+
+# Offline-disable checksums on the lagging standby only. This is a
+# divergence 012_offline_standby.pl declares survivable: the standby keeps
+# its own state, warns, and stays readable.
+$lagging->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $lagging->data_dir);
+$lagging->start;
+test_checksum_state($lagging, 'off');
+
+# Give the lagging standby plenty of pages to write while it is locally
+# "off", and force a restartpoint so they reach disk without checksums.
+# The other standby replays the same WAL with checksums on.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+$lagging->safe_psql('postgres', "CHECKPOINT;");
+
+# Take the "failover" standby out of the picture. It is promoted from
+# this position later, so everything replayed from here on is the part of
+# the WAL that diverges and that pg_rewind will roll back. t1 is already
+# replayed everywhere, so its blocks are *not* rolled back.
+$failover->stop;
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$primary->wait_for_catchup($lagging);
+
+# Failover: promote the other standby, which forks the timeline behind the
+# replay position of the lagging standby.
+$primary->stop('fast');
+$failover->start;
+$failover->promote;
+$failover->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$failover->safe_psql('postgres', "CHECKPOINT;");
+
+# Re-attaching the lagging standby with pg_rewind must be refused: the
+# divergence checkpoint says "on" while the target is "off", so the rewound
+# server would verify checksums over the blocks it keeps.
+$lagging->stop;
+
+command_fails_like(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ],
+ qr/data checksums were enabled at the point of divergence but are disabled on the target server/,
+ 'pg_rewind refuses a target whose divergence checkpoint has checksums');
+
+# Enabling checksums offline on the target rewrites all of its pages, after
+# which the rewind is safe.
+system_or_bail('pg_checksums', '--enable', '--pgdata', $lagging->data_dir);
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ]);
+like($stderr, qr/Done!/, 'pg_rewind completes after the offline enable');
+
+$lagging->append_conf('postgresql.conf', 'port = ' . $lagging->port);
+$lagging->enable_streaming($failover);
+$lagging->set_standby_mode;
+$lagging->start;
+
+my ($rc, $out, $err) =
+ $lagging->psql('postgres',
+ "SELECT setting FROM pg_settings " . "WHERE name = 'data_checksums';");
+is($rc, 0, 'rewound standby accepts connections') or diag("stderr: $err");
+is($out, 'on', 'rewound standby resumes with checksums on');
+
+($rc, $out, $err) = $lagging->psql('postgres', "SELECT count(*) FROM t1;");
+is($rc, 0, 'blocks kept from the target stay readable')
+ or diag("stderr: $err");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($lagging->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the rewound standby');
+
+$lagging->stop('immediate');
+$failover->stop;
+done_testing();
--
2.55.0
[application/octet-stream] v7-0004-pg_combinebackup-Refuse-mixed-data-checksum-state.patch (9.5K, ../../CAN4CZFMFcgfgJ99RYhax-T+YJH=CWLy6ZGANMxf77SrSirw3VQ@mail.gmail.com/5-v7-0004-pg_combinebackup-Refuse-mixed-data-checksum-state.patch)
download | inline diff:
From 198df7ca02a53f4b7841e2e7d2ac1e8f36ad1a12 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Tue, 25 Aug 2026 15:06:58 +0000
Subject: [PATCH v7 4/5] pg_combinebackup: Refuse mixed data checksum states in
a backup chain
check_control_files() detected a chain whose backups were taken under
different data checksum states, warned, and proceeded. The output
directory keeps the last backup's control file, so a full backup taken
with checksums off combined with an incremental taken after an offline
enable produces a cluster whose control file says "on" while most of
its blocks carry no checksums; it starts, and then every connection
dies on the first unchecksummed catalog page. An offline enable
between two backups of a chain is all it takes, since it rewrites
every page without logging anything, so the incremental backup does
not re-ship the pages.
Turn the warning into an error, matching what pg_rewind does for the
equivalent combinations. The check stays asymmetric on purpose: when
the last backup was taken with checksums off, stale checksums from an
earlier backup are never verified, and that chain remains usable.
---
doc/src/sgml/ref/pg_combinebackup.sgml | 19 +--
src/bin/pg_combinebackup/pg_combinebackup.c | 14 +-
src/test/modules/test_checksums/meson.build | 1 +
.../t/023_combinebackup_mixed.pl | 134 ++++++++++++++++++
4 files changed, 155 insertions(+), 13 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/023_combinebackup_mixed.pl
diff --git a/doc/src/sgml/ref/pg_combinebackup.sgml b/doc/src/sgml/ref/pg_combinebackup.sgml
index 9a6d201e0b8..f5d7e177b36 100644
--- a/doc/src/sgml/ref/pg_combinebackup.sgml
+++ b/doc/src/sgml/ref/pg_combinebackup.sgml
@@ -306,18 +306,19 @@ PostgreSQL documentation
<para>
<literal>pg_combinebackup</literal> does not recompute page checksums when
- writing the output directory. Therefore, if any of the backups used for
- reconstruction were taken with checksums disabled, but the final backup was
- taken with checksums enabled, the resulting directory may contain pages
- with invalid checksums.
+ writing the output directory. It therefore refuses a chain in which some
+ of the backups used for reconstruction were taken with checksums disabled
+ but the final backup was taken with checksums enabled: the resulting
+ directory would contain pages with invalid checksums, and the cluster
+ would fail checksum verification as soon as it reads them. Take a new
+ full backup after enabling data checksums with
+ <xref linkend="app-pgchecksums"/>.
</para>
<para>
- To avoid this problem, taking a new full backup after changing the checksum
- state of the cluster using <xref linkend="app-pgchecksums"/> is
- recommended. Otherwise, you can disable and then optionally reenable
- checksums on the directory produced by <literal>pg_combinebackup</literal>
- in order to correct the problem.
+ The reverse case is accepted: when the final backup was taken with
+ checksums disabled, stale checksums copied from an earlier backup are
+ never verified.
</para>
</refsect1>
diff --git a/src/bin/pg_combinebackup/pg_combinebackup.c b/src/bin/pg_combinebackup/pg_combinebackup.c
index 86e2ee37c40..254a27b125b 100644
--- a/src/bin/pg_combinebackup/pg_combinebackup.c
+++ b/src/bin/pg_combinebackup/pg_combinebackup.c
@@ -670,13 +670,19 @@ check_control_files(int n_backups, char **backup_dirs)
pg_log_debug("system identifier is %" PRIu64, system_identifier);
/*
- * Warn the user if not all backups are in the same state with regards to
- * checksums.
+ * Reject a chain whose backups were not all in the same state with
+ * regards to checksums. The loop above only flags that when the last
+ * backup has checksums enabled: the blocks taken from the older backups
+ * have no checksums, and pg_combinebackup does not recompute them, so the
+ * combined cluster would fail verification as soon as it reads them. The
+ * other direction is harmless, since stale checksums are never verified
+ * while checksums are disabled.
*/
if (data_checksum_mismatch)
{
- pg_log_warning("only some backups have checksums enabled");
- pg_log_warning_hint("Disable, and optionally reenable, checksums on the output directory to avoid failures.");
+ pg_log_error("only some backups have checksums enabled");
+ pg_log_error_hint("Take a new full backup after changing the data checksum state with pg_checksums.");
+ exit(1);
}
return system_identifier;
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index a4f25bac7a5..8b85a7848bb 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -46,6 +46,7 @@ tests += {
't/020_cascade_divergence.pl',
't/021_rewind_state.pl',
't/022_rewind_standby_target.pl',
+ 't/023_combinebackup_mixed.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/023_combinebackup_mixed.pl b/src/test/modules/test_checksums/t/023_combinebackup_mixed.pl
new file mode 100644
index 00000000000..ce9dee9413b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/023_combinebackup_mixed.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_combinebackup refuses a chain whose backups were taken under
+# different data checksum states when the final state has them enabled.
+#
+# The output directory keeps the last backup's control file, so a full
+# backup taken with checksums off plus an incremental taken after an
+# offline enable would produce a directory whose control file says "on"
+# while most of its blocks have no checksums. An offline enable between
+# the two backups is enough to get there: it rewrites every page but logs
+# nothing, so the incremental backup does not re-ship the pages.
+#
+# The reverse order stays allowed: with the final backup taken with
+# checksums off, stale checksums from an earlier backup are never verified.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $node = PostgreSQL::Test::Cluster->new('node');
+$node->init(
+ no_data_checksums => 1,
+ has_archiving => 1,
+ allows_streaming => 1);
+$node->append_conf('postgresql.conf', 'summarize_wal = on');
+$node->append_conf('postgresql.conf', 'autovacuum = off');
+$node->start;
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+test_checksum_state($node, 'off');
+
+$node->backup('full');
+
+# Offline enable: rewrites every page, logs nothing.
+$node->stop;
+system_or_bail('pg_checksums', '--enable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'on');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full/backup_manifest',
+ ],
+ 'incremental backup with checksums on');
+
+# Combining the chain must be refused: the result would say "on" while the
+# blocks inherited from the full backup have no checksums.
+command_fails_like(
+ [
+ 'pg_combinebackup',
+ $node->backup_dir . '/full',
+ $node->backup_dir . '/incr',
+ '--output' => $node->backup_dir . '/combined',
+ ],
+ qr/only some backups have checksums enabled/,
+ 'pg_combinebackup refuses a chain crossing an offline enable');
+
+# The other direction: full backup with checksums on, offline disable, then
+# an incremental. The last backup wins, so the combined cluster comes up off.
+$node->backup('full2');
+
+$node->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr2',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full2/backup_manifest',
+ ],
+ 'incremental backup with checksums off');
+
+my $restored = PostgreSQL::Test::Cluster->new('restored');
+$restored->init_from_backup(
+ $node, 'incr2',
+ combine_with_prior => ['full2'],
+ has_restoring => 1,
+ standby => 0);
+
+$restored->start;
+
+my ($rc, $stdout, $stderr) = $restored->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'combined backup accepts connections') or diag("stderr: $stderr");
+is($stdout, 'off', 'combined backup follows the last backup state');
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t3;');
+is($rc, 0, 'blocks from the incremental backup are readable')
+ or diag("stderr: $stderr");
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t1;');
+is($rc, 0, 'blocks inherited from the full backup are readable')
+ or diag("stderr: $stderr");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($restored->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures in the combined backup');
+
+$restored->stop('immediate');
+$node->stop;
+
+done_testing();
--
2.55.0
[application/octet-stream] v7-0001-Do-not-adopt-data-checksum-state-from-another-nod.patch (119.2K, ../../CAN4CZFMFcgfgJ99RYhax-T+YJH=CWLy6ZGANMxf77SrSirw3VQ@mail.gmail.com/6-v7-0001-Do-not-adopt-data-checksum-state-from-another-nod.patch)
download | inline diff:
From 63f8f9757fcc85b236403f21cbc03f24e6f23fd2 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:02 +0000
Subject: [PATCH v7 1/5] Do not adopt data checksum state from another node
during replay
Offline pg_checksums changes are local to one node, but replay adopted
the data checksum state carried by checkpoint records unconditionally.
After an offline enable on the primary, a standby whose pages were
never rewritten started verifying checksums it does not have; after an
offline change on a standby, the next replayed checkpoint silently
reverted it.
Instead, cross-check the replayed state against the local one and warn
once per divergent value. When the states match again, say so in the
log and re-arm the warning.
The control file now tracks this node's state alone, which makes the
moment it is written out part of the design: it may only claim "on"
once every page on disk carries a checksum, or a crash-restart would
verify pages that were flushed before the transition rewrote them.
XLOG2_CHECKSUMS replay therefore persists every state but "on" as soon
as it is replayed, none of them verifying anything, and leaves "on" to
the flush of the next restartpoint. A restartpoint only persists a
state its flush ran under from beginning to end, a promotion writes a
full checkpoint instead of an end-of-recovery record while the state is
still unpersisted, and a shutdown restartpoint with no new checkpoint
record to work from flushes before catching the control file up rather
than leaving it behind. Enabling checksums on the primary defers the
same write to after the checkpoint that flushes the rewritten pages:
the ones the worker found in shared buffers do not go out through its
ring buffer. A checkpoint persists the state its own redo point ran
under, under the same rule a restartpoint follows, since the checkpoint
that licenses "on" also moves the redo point past the record announcing
it: without that, a crash before the deferred write would resume above
the record and resolve a finished transition as interrupted. Checkpoint
records replayed below the consistency point are not cross-checked, the
persisted state being legitimately newer than what they carry.
For the redo point to be a reliable anchor, the states carried by WAL
records must match their WAL order. Transitions insert their record
and publish the new state under the new DataChecksumTransitionLock,
which a checkpoint holds across sampling the state and inserting its
XLOG_CHECKPOINT_REDO record; without that, a redo record could follow
a transition record in WAL while still carrying the pre-transition
state, and recovery resuming there would resolve the finished
transition as interrupted all the same.
pg_control also gains a watermark, the end LSN of the newest
XLOG2_CHECKSUMS record the node has written or applied, persisted
together with the state it produced, and replay skips records at or
below it. Without it, a standby that replayed a transition and stopped
cleanly before any restartpoint moved past the record would re-apply it
on the next startup, overriding a pg_checksums change made while it was
down; the offline change writes no WAL, so nothing would restore it. A
flag next to the watermark marks a state last written by pg_checksums
as local to this node, and recovery never adopts a checkpoint-borne
state over a local one, nor over a control file whose watermark already
covers the starting checkpoint. The watermark also stands in for the
state comparisons around the flushes above: record positions are
unique, so a full round trip back to the sampled state cannot alias.
A base backup copies the control file at an arbitrary moment, so its
state can be newer than the redo point replay starts from. Adopt the
state carried by the starting checkpoint record under backup-label
recovery; for a shutdown checkpoint, which replay does not see, take
it in StartupXLOG. Backups taken from a standby are the exception:
their starting checkpoint was written by the upstream primary, so they
keep the state of the control file that was copied with them.
pg_rewind keeps the target's own state, watermark and flag in the
control file it installs, since most of the data directory remains the
target's; replay from the last common checkpoint still adopts a state
its watermark does not cover, and applies any online transition the
target has not seen.
Bump PG_CONTROL_VERSION.
The documented procedure for offline changes in a replication setup
becomes the lockstep one: stop all nodes, run pg_checksums on each of
them, then restart. Tests cover the lockstep procedure, divergent
offline changes on either node and down a cascading chain,
crash-restarts and promotions around online transitions, checkpoints
racing an online enable and the transition record itself, a base backup
taken during an online enable, an offline disable on a standby
surviving the re-replay of the enable that preceded it, and pg_rewind
across an online enable as well as across offline enables on both nodes
after a divergence.
---
doc/src/sgml/ref/pg_checksums.sgml | 32 +-
doc/src/sgml/wal.sgml | 8 +
src/backend/access/transam/xlog.c | 549 ++++++++++++++++--
.../utils/activity/wait_event_names.txt | 1 +
src/bin/pg_checksums/pg_checksums.c | 10 +
src/bin/pg_controldata/pg_controldata.c | 4 +
src/bin/pg_resetwal/pg_resetwal.c | 7 +
src/bin/pg_rewind/pg_rewind.c | 14 +
src/bin/pg_upgrade/controldata.c | 2 +-
src/include/catalog/pg_control.h | 21 +-
src/include/storage/lwlocklist.h | 1 +
src/test/modules/test_checksums/Makefile | 2 +-
src/test/modules/test_checksums/meson.build | 9 +
.../test_checksums/t/012_offline_standby.pl | 289 +++++++++
.../modules/test_checksums/t/013_rewind.pl | 201 +++++++
.../modules/test_checksums/t/014_lockstep.pl | 182 ++++++
.../t/015_standby_crash_after_disable.pl | 125 ++++
.../t/016_promote_enable_crash.pl | 134 +++++
.../test_checksums/t/017_restartpoint_race.pl | 138 +++++
.../t/018_enable_crash_windows.pl | 496 ++++++++++++++++
.../t/019_standby_shutdown_catchup.pl | 126 ++++
.../t/020_cascade_divergence.pl | 115 ++++
22 files changed, 2390 insertions(+), 76 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/012_offline_standby.pl
create mode 100644 src/test/modules/test_checksums/t/013_rewind.pl
create mode 100644 src/test/modules/test_checksums/t/014_lockstep.pl
create mode 100644 src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
create mode 100644 src/test/modules/test_checksums/t/016_promote_enable_crash.pl
create mode 100644 src/test/modules/test_checksums/t/017_restartpoint_race.pl
create mode 100644 src/test/modules/test_checksums/t/018_enable_crash_windows.pl
create mode 100644 src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
create mode 100644 src/test/modules/test_checksums/t/020_cascade_divergence.pl
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index ae66fad3f0f..bf07da09754 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -242,15 +242,29 @@ PostgreSQL documentation
data directory must not be started or else data loss may occur.
</para>
<para>
- When using a replication setup with tools which perform direct copies
- of relation file blocks (for example <xref linkend="app-pgrewind"/>),
- enabling or disabling checksums can lead to page corruptions in the
- shape of incorrect checksums if the operation is not done consistently
- across all nodes. When enabling or disabling checksums in a replication
- setup, it is thus recommended to stop all the clusters before switching
- them all consistently. Destroying all standbys, performing the operation
- on the primary and finally recreating the standbys from scratch is also
- safe.
+ Enabling or disabling checksums with
+ <application>pg_checksums</application> changes only the local data
+ directory; the new state does not propagate over replication. In a
+ replication setup the same change must be applied to every node: stop all
+ nodes, run <application>pg_checksums</application> on each of them, and
+ only then restart them. Tools that copy relation file blocks directly
+ between nodes, such as <xref linkend="app-pgrewind"/>, likewise require
+ both nodes to be in the same data checksum state.
+ </para>
+ <para>
+ If the change is applied inconsistently, each node keeps its own state,
+ and a standby logs a warning when the state recorded in the replayed WAL
+ differs from its own. The same warning can appear transiently while a
+ standby catches up over WAL written before a consistent change; it stops
+ once a checkpoint record carrying the new state has been replayed.
+ </para>
+ <para>
+ A standby whose data directory was never checksummed must not have
+ checksums enabled by catching up this way. Converge the cluster by
+ running <application>pg_checksums</application> on it while stopped, by
+ enabling checksums online, or by recreating it from a base backup. Note
+ that an online enable only starts from a primary whose checksums are off,
+ so if they are already enabled there, disable them online first.
</para>
<para>
If <application>pg_checksums</application> is aborted or killed while
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index ec62d17fbcc..db18a8c516e 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -317,6 +317,14 @@
verify checksums, on an offline cluster.
</para>
+ <para>
+ An offline change only affects the data directory it is run on; the
+ new state does not propagate over replication. In a replication setup
+ the same change must be applied to all nodes while all of them are
+ stopped, as described in <xref linkend="app-pgchecksums"/>. A standby
+ whose state diverges logs a warning but keeps its local setting.
+ </para>
+
</sect2>
<sect2 id="checksums-online-enable-disable" xreflabel="Online Enabling of Checksums">
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index de4c96e135f..20428ad68cf 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -556,9 +556,17 @@ typedef struct XLogCtlData
*/
XLogRecPtr lastFpwDisableRecPtr;
- /* last data_checksum_version we've seen */
+ /* current data checksum state of this node */
uint32 data_checksum_version;
+ /*
+ * In-memory copies of ControlFile->data_checksum_lsn and
+ * ControlFile->data_checksum_is_local, see there. Updated together with
+ * data_checksum_version under info_lck.
+ */
+ XLogRecPtr data_checksum_lsn;
+ bool data_checksum_is_local;
+
slock_t info_lck; /* locks shared variables shown above */
/*
@@ -690,6 +698,14 @@ static ChecksumStateType LocalDataChecksumState = 0;
*/
int data_checksums = 0;
+/*
+ * Whether replay of the next checkpoint-family record must adopt the data
+ * checksum state it carries. Set when recovery starts from a base backup,
+ * where the state at the redo point takes precedence over the control file
+ * copied with the backup later.
+ */
+static bool adoptChecksumStateFromNextCheckpoint = false;
+
/* For WALInsertLockAcquire/Release functions */
static int MyLockNo = 0;
static bool holdingAllLocks = false;
@@ -730,6 +746,8 @@ static void ValidateXLOGDirectoryStructure(void);
static void CleanupBackupHistory(void);
static void UpdateMinRecoveryPoint(XLogRecPtr lsn, bool force);
static bool PerformRecoveryXLogAction(void);
+static void CheckReplayedDataChecksumState(uint32 replayed_version);
+static void AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn);
static void InitControlFile(uint64 sysidentifier, uint32 data_checksum_version);
static void WriteControlFile(void);
static void ReadControlFile(void);
@@ -757,7 +775,7 @@ static void WALInsertLockAcquireExclusive(void);
static void WALInsertLockRelease(void);
static void WALInsertLockUpdateInsertingAt(XLogRecPtr insertingAt);
-static void XLogChecksums(uint32 new_type);
+static XLogRecPtr XLogChecksums(uint32 new_type);
/*
* Insert an XLOG record represented by an already-constructed chain of data
@@ -4774,6 +4792,7 @@ void
SetDataChecksumsOnInProgress(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
/*
* The state transition is performed in a critical section with
@@ -4782,14 +4801,12 @@ SetDataChecksumsOnInProgress(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4827,6 +4844,8 @@ void
SetDataChecksumsOn(void)
{
uint64 barrier;
+ bool persist;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4847,23 +4866,11 @@ SetDataChecksumsOn(void)
SpinLockRelease(&XLogCtl->info_lck);
INJECTION_POINT("datachecksums-enable-checksums-delay", NULL);
+ INJECTION_POINT_LOAD("datachecksums-on-before-publish");
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_VERSION);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- SpinLockRelease(&XLogCtl->info_lck);
-
- /*
- * Update the controlfile before waiting since if we have an immediate
- * shutdown while waiting we want to come back up with checksums enabled.
- */
- LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- UpdateControlFile();
- LWLockRelease(ControlFileLock);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_VERSION);
barrier = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_CHECKSUM_ON);
@@ -4873,6 +4880,39 @@ SetDataChecksumsOn(void)
INJECTION_POINT("datachecksums-on-before-checkpoint", NULL);
RequestCheckpoint(CHECKPOINT_FORCE | CHECKPOINT_WAIT | CHECKPOINT_FAST);
+
+ INJECTION_POINT("datachecksums-on-after-checkpoint", NULL);
+
+ /*
+ * Persist "on" only now that the checkpoint has flushed the pages the
+ * transition rewrote. Crash recovery initializes verification from the
+ * control file but resumes from a checkpoint that can predate the
+ * transition, so an "on" persisted earlier would have replay verify pages
+ * whose rewrite never reached disk: the pages the worker found in shared
+ * buffers are not written back by its ring buffer.
+ *
+ * The checkpoint above normally persists the state itself; this covers the
+ * case where it started before the record written above and left the field
+ * alone. Crashing before this point is safe, as replay then re-establishes
+ * "on" from the full page images of the rewrite. Skip the write if the
+ * state moved on meanwhile, since whatever moved it persists its own.
+ * Compare the watermark rather than the state: a state comparison could
+ * not tell our transition from a later round trip back to "on".
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ persist = (XLogCtl->data_checksum_lsn == recptr);
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (persist)
+ {
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+ }
+
WaitForProcSignalBarrier(barrier);
}
@@ -4893,6 +4933,7 @@ void
SetDataChecksumsOff(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4918,14 +4959,12 @@ SetDataChecksumsOff(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4956,14 +4995,12 @@ SetDataChecksumsOff(void)
/* Ensure that we don't incur a checkpoint during disabling checksums */
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -5001,6 +5038,132 @@ SetLocalDataChecksumState(uint32 data_checksum_version)
data_checksums = data_checksum_version;
}
+/*
+ * Cross-check the data checksum state carried by a replayed checkpoint record
+ * against the state of this node.
+ *
+ * The state in the record belongs to whichever node wrote the WAL, and must
+ * not be adopted: an offline state change made with pg_checksums on one node
+ * of a replication set generates no WAL, and adopting would leak it into the
+ * other nodes through replay. Backup label recovery is the one exception,
+ * see AdoptReplayedDataChecksumState(). XLOG_CHECKPOINT_ONLINE needs no call
+ * here, as an online checkpoint's state already traveled in the preceding
+ * XLOG_CHECKPOINT_REDO record.
+ *
+ * Only archive recovery can see a lasting mismatch, as only there can the WAL
+ * and the control file come from different nodes or different times. In
+ * crash recovery a mismatch means replay resumed from a restartpoint
+ * predating an already-applied XLOG2_CHECKSUMS record, and replaying forward
+ * re-establishes the same state.
+ */
+static void
+CheckReplayedDataChecksumState(uint32 replayed_version)
+{
+ /*
+ * Warn once per remote value, so a lasting mismatch does not flood the
+ * log. Matching states re-arm the warning. Backend-local state is
+ * enough: replay only runs in the startup process, and restarting it
+ * re-arms as well.
+ */
+ static uint32 last_warned_version = PG_UINT32_MAX;
+ uint32 local_version;
+
+ if (!ArchiveRecoveryRequested)
+ return;
+
+ /*
+ * Re-replayed WAL below the consistency point was already cross-checked
+ * before minRecoveryPoint was last persisted, and the persisted state can
+ * legitimately be newer than what checkpoint records there carry:
+ * XLOG2_CHECKSUMS replay persists most states ahead of the restartpoint
+ * horizon. In particular the checkpoint record recovery restarts from is
+ * such a re-replay.
+ */
+ if (!reachedConsistency)
+ return;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ local_version = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (replayed_version == local_version)
+ {
+ /*
+ * Report convergence if this process warned before. Nothing else
+ * tells the operator that running pg_checksums on the other nodes, or
+ * a rebuild, took effect. A restart in between loses the context,
+ * but it re-arms the warning too.
+ */
+ if (last_warned_version != PG_UINT32_MAX)
+ ereport(LOG,
+ errmsg("data checksum state \"%s\" of this node now agrees with the replayed WAL",
+ get_checksum_state_string(local_version)));
+
+ last_warned_version = PG_UINT32_MAX;
+ return;
+ }
+
+ /* the nodes legitimately differ while an online transition runs */
+ if (replayed_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ replayed_version == PG_DATA_CHECKSUM_INPROGRESS_OFF ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ return;
+
+ if (replayed_version == last_warned_version)
+ return;
+ last_warned_version = replayed_version;
+
+ ereport(WARNING,
+ errmsg("data checksum state \"%s\" of this node does not match the state \"%s\" in the replayed WAL",
+ get_checksum_state_string(local_version),
+ get_checksum_state_string(replayed_version)),
+ errdetail("The data checksum state was most likely changed with pg_checksums on another node."),
+ errhint("Apply the same change with pg_checksums on the primary and all standby servers, or rebuild this server from a base backup."));
+}
+
+/*
+ * Adopt the data checksum state found at the redo point of backup label
+ * recovery. Persist it immediately so that a crash before the first
+ * restartpoint does not resurrect the state copied with the backup; a crash
+ * at this point restarts from the same redo point, so the control file does
+ * not run ahead of the replay position. If the value is unchanged the
+ * control file already carries it, so both the barrier and the persist are
+ * skipped.
+ *
+ * lsn is the location of the checkpoint-family record the state was taken
+ * from and becomes the new watermark: the adopted state covers everything
+ * below the redo point, which replay never revisits.
+ */
+static void
+AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn)
+{
+ bool changed = false;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (XLogCtl->data_checksum_version != new_version)
+ {
+ XLogCtl->data_checksum_version = new_version;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(new_version);
+ changed = true;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (!changed)
+ return;
+
+ EmitAndWaitDataChecksumsBarrier(new_version);
+
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = new_version;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+}
+
/* guc hook */
const char *
show_data_checksums(void)
@@ -5454,6 +5617,8 @@ XLOGShmemInit(void *arg)
/* Use the checksum info from control file */
XLogCtl->data_checksum_version = ControlFile->data_checksum_version;
+ XLogCtl->data_checksum_lsn = ControlFile->data_checksum_lsn;
+ XLogCtl->data_checksum_is_local = ControlFile->data_checksum_is_local;
SetLocalDataChecksumState(XLogCtl->data_checksum_version);
SpinLockInit(&XLogCtl->Insert.insertpos_lck);
@@ -6031,6 +6196,53 @@ StartupXLOG(void)
SetCommitTsLimit(checkPoint.oldestCommitTsXid,
checkPoint.newestCommitTsXid);
+ /*
+ * When recovery starts from a base backup, the control file was copied at
+ * an arbitrary moment and its data checksum state may differ from the
+ * state at the redo point, which is what the WAL from there on was
+ * written under. Adopt the state of the starting checkpoint: a shutdown
+ * checkpoint is not replayed, so take it from the record read above; the
+ * redo point of an online checkpoint is its CHECKPOINT_REDO record, so
+ * let the replay of that record adopt it. Check backupStartPoint in
+ * addition to the label: on a crash restart during backup recovery the
+ * label file is already renamed away, but the start point persists until
+ * the backup end record.
+ *
+ * Not for a base backup taken from a standby, though. Its starting
+ * checkpoint is the standby's last restartpoint, a record written by the
+ * upstream primary, whose state is not the one the copied files were
+ * written under. The copied control file is right there, as a standby
+ * persists its state only at restartpoint horizons and so never claims
+ * more than what reached disk. Such backups are recognized by
+ * backupEndPoint together with backupEndRequired; backupEndPoint is only
+ * set for "BACKUP FROM: standby" labels and persists across a crash
+ * restart. pg_rewind writes a standby label as well, but no
+ * backupEndPoint, and its recovery keeps adopting: the control file it
+ * installs carries the target's own checksum state, which can lag the
+ * redo point of the last common checkpoint the same way a restartpoint
+ * horizon can.
+ *
+ * Never adopt over a state the control file's watermark or local flag
+ * marks as newer than the starting checkpoint. A pg_checksums change is
+ * local to the node and generates no WAL, so nothing in the replayed WAL
+ * could ever restore it once overwritten; and a control file whose
+ * watermark lies above the redo point already contains the effect of
+ * every transition record up to there, including the state the starting
+ * checkpoint carries.
+ */
+ if ((haveBackupLabel || XLogRecPtrIsValid(ControlFile->backupStartPoint)) &&
+ !(XLogRecPtrIsValid(ControlFile->backupEndPoint) &&
+ ControlFile->backupEndRequired) &&
+ !ControlFile->data_checksum_is_local &&
+ checkPoint.redo > ControlFile->data_checksum_lsn)
+ {
+ if (wasShutdown)
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ checkPoint.redo);
+ else
+ adoptChecksumStateFromNextCheckpoint = true;
+ }
+
/*
* Clear out any old relcache cache files. This is *necessary* if we do
* any WAL replay, since that would probably result in the cache files
@@ -6633,11 +6845,7 @@ StartupXLOG(void)
if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
ereport(WARNING,
@@ -6654,11 +6862,7 @@ StartupXLOG(void)
else if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
}
@@ -6814,6 +7018,23 @@ static bool
PerformRecoveryXLogAction(void)
{
bool promoted = false;
+ bool flushForChecksums;
+ uint32 checksum_state;
+
+ /*
+ * The end-of-recovery record persists the data checksum state without
+ * flushing the buffer pool, but the control file may only claim "on" once
+ * every page on disk carries a checksum. If replay entered that state
+ * without a restartpoint following it, the pages rewritten by the
+ * transition are still only in the buffer pool, so take the full
+ * checkpoint below instead of the lightweight record.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ flushForChecksums = (checksum_state == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile->data_checksum_version != checksum_state);
/*
* Perform a checkpoint to update all our recovery activity to disk.
@@ -6829,7 +7050,7 @@ PerformRecoveryXLogAction(void)
* fully out of recovery mode and already accepting queries.
*/
if (ArchiveRecoveryRequested && IsUnderPostmaster &&
- PromoteIsTriggered())
+ PromoteIsTriggered() && !flushForChecksums)
{
promoted = true;
@@ -7436,6 +7657,7 @@ CreateCheckPoint(int flags)
uint32 freespace;
XLogRecPtr PriorRedoPtr;
XLogRecPtr last_important_lsn;
+ XLogRecPtr checksumLsn;
VirtualTransactionId *vxids;
int nvxids;
int oldXLogAllowed = 0;
@@ -7548,11 +7770,14 @@ CreateCheckPoint(int flags)
checkPoint.wal_level = wal_level;
/*
- * Get the current data_checksum_version value from xlogctl, valid at the
- * time of the checkpoint.
+ * Get the current data_checksum_version value from xlogctl. This is
+ * final only for a shutdown checkpoint, where no concurrent transition
+ * is possible; an online checkpoint resamples it together with the redo
+ * record below.
*/
SpinLockAcquire(&XLogCtl->info_lck);
checkPoint.dataChecksumState = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
if (shutdown)
@@ -7609,10 +7834,21 @@ CreateCheckPoint(int flags)
{
xl_checkpoint_redo redo_rec;
+ /*
+ * Sample the data checksum state and insert the redo record under
+ * DataChecksumTransitionLock, so that a concurrent transition cannot
+ * insert its XLOG2_CHECKSUMS record between the sampling and the
+ * insertion below. Without this, the redo record could follow the
+ * transition record in WAL while carrying the pre-transition state,
+ * and recovery resuming here would never learn about the transition.
+ * See XLogChecksums().
+ */
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
WALInsertLockAcquire();
redo_rec.wal_level = wal_level;
SpinLockAcquire(&XLogCtl->info_lck);
redo_rec.data_checksum_version = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
WALInsertLockRelease();
@@ -7620,6 +7856,14 @@ CreateCheckPoint(int flags)
XLogBeginInsert();
XLogRegisterData(&redo_rec, sizeof(xl_checkpoint_redo));
(void) XLogInsert(RM_XLOG_ID, XLOG_CHECKPOINT_REDO);
+ LWLockRelease(DataChecksumTransitionLock);
+
+ /*
+ * The checkpoint record must carry the same state as the redo record
+ * just inserted: the sample taken before redo determination can be
+ * stale by now, and the pair would otherwise disagree.
+ */
+ checkPoint.dataChecksumState = redo_rec.data_checksum_version;
/*
* XLogInsertRecord will have updated XLogCtl->Insert.RedoRecPtr in
@@ -7823,6 +8067,40 @@ CreateCheckPoint(int flags)
ControlFile->minRecoveryPoint = InvalidXLogRecPtr;
ControlFile->minRecoveryPointTLI = 0;
+ /*
+ * Persist the data checksum state this node runs under. Only the
+ * top-level field tracks this node; ControlFile->checkPointCopy above is
+ * a historical record used to resume replay.
+ *
+ * checkPoint.dataChecksumState was sampled under
+ * DataChecksumTransitionLock together with the redo record, so it is the
+ * state in effect at the redo point. If it was "on",
+ * the XLOG2_CHECKSUMS record announcing that precedes the redo point and
+ * every page the transition rewrote was dirtied before it, so
+ * CheckPointGuts() has just written all of them out. Recording the state
+ * here is what keeps a finished transition from being resolved as
+ * interrupted when this checkpoint is the one crash recovery resumes
+ * from: replay never sees the record announcing it.
+ *
+ * Persist it only if the state did not change while the flush was in
+ * progress. If it changed in between, the pages written out straddle two
+ * states, and the newer one could claim checksums that pages already on
+ * disk do not carry; leave the field to the next checkpoint then.
+ * SetDataChecksumsOff() persists the states that are safe to enter
+ * without a flush already. Compare the watermark rather than the state:
+ * a state comparison could not tell a full round trip back to the
+ * sampled value apart from no change at all, and the flushed pages
+ * straddle the intermediate states all the same.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksumLsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checkPoint.dataChecksumState;
+ ControlFile->data_checksum_lsn = checksumLsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
/*
* Persist unloggedLSN value. It's reset on crash recovery, so this goes
* unused on non-shutdown checkpoints, but seems useful to store it always
@@ -7967,9 +8245,11 @@ CreateEndOfRecoveryRecord(void)
ControlFile->minRecoveryPoint = recptr;
ControlFile->minRecoveryPointTLI = xlrec.ThisTimeLineID;
- /* start with the latest checksum version (as of the end of recovery) */
+ /* persist the data checksum state this node ended recovery with */
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
@@ -8174,6 +8454,9 @@ CreateRestartPoint(int flags)
XLogRecPtr endptr;
XLogSegNo _logSegNo;
TimestampTz xtime;
+ uint32 checksum_state;
+ XLogRecPtr checksum_lsn;
+ bool checksum_is_local;
/* Concurrent checkpoint/restartpoint cannot happen */
Assert(!IsUnderPostmaster || MyBackendType == B_CHECKPOINTER);
@@ -8220,8 +8503,45 @@ CreateRestartPoint(int flags)
UpdateMinRecoveryPoint(InvalidXLogRecPtr, true);
if (flags & CHECKPOINT_IS_SHUTDOWN)
{
+ bool catchUpChecksums;
+
+ /*
+ * There is no new restartpoint to persist the data checksum state
+ * with, but a cleanly stopped node should not leave the control
+ * file behind the state replay reached: pg_checksums and
+ * pg_rewind read it, and an in-progress state there makes them
+ * refuse to run. Catching it up needs the same guarantee a
+ * restartpoint gives: that every page on disk carries a checksum,
+ * so flush the buffer pool first. Only a transition to
+ * "on" that no restartpoint followed can get here; the other
+ * states are already persisted by XLOG2_CHECKSUMS replay. Replay
+ * has ended by now, so the state cannot change under us.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ checksum_is_local = XLogCtl->data_checksum_is_local;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ catchUpChecksums =
+ (checksum_lsn != ControlFile->data_checksum_lsn &&
+ XLogRecPtrIsValid(lastCheckPointRecPtr));
+
+ if (catchUpChecksums)
+ {
+ MemSet(&CheckpointStats, 0, sizeof(CheckpointStats));
+ CheckpointStats.ckpt_start_t = GetCurrentTimestamp();
+ CheckPointGuts(lastCheckPoint.redo, flags);
+ }
+
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
+ if (catchUpChecksums)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = checksum_is_local;
+ }
UpdateControlFile();
LWLockRelease(ControlFileLock);
}
@@ -8262,6 +8582,17 @@ CreateRestartPoint(int flags)
/* Update the process title */
update_checkpoint_display(flags, true, false);
+ /*
+ * Note the data checksum state the flush below starts under. Replay runs
+ * concurrently and can change the state while the flush is in progress,
+ * in which case the flush covers pages written under both states; see
+ * where the state is persisted further down.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
CheckPointGuts(lastCheckPoint.redo, flags);
/*
@@ -8321,8 +8652,26 @@ CreateRestartPoint(int flags)
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
}
- /* we shall start with the latest checksum version */
- ControlFile->data_checksum_version = lastCheckPoint.dataChecksumState;
+ /*
+ * Persist the data checksum state of this node. Not the state of the
+ * replayed checkpoint: that one belongs to the node that wrote it and
+ * may differ after an offline change on either side.
+ * ControlFile->checkPointCopy above keeps the replayed value on
+ * purpose, being a historical record used to resume replay rather
+ * than a tracker of node state.
+ *
+ * Persist only if the flush above ran under one state throughout; see
+ * CreateCheckPoint() for why, including why this compares the
+ * watermark and not the state.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksum_lsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
}
@@ -8763,9 +9112,21 @@ XLogReportParameters(void)
}
/*
- * Log the new state of checksums
+ * Log and publish the new state of checksums
+ *
+ * Inserting the record and publishing the new state must be atomic with
+ * respect to a checkpoint sampling the state for its XLOG_CHECKPOINT_REDO
+ * record: without that, a checkpoint could read the old state after the
+ * record is already in WAL and insert a redo record that both precedes the
+ * transition in WAL order and carries the pre-transition state. Recovery
+ * resuming from such a redo point would never replay the transition record
+ * and resolve the finished transition as interrupted.
+ * DataChecksumTransitionLock serializes the two; see CreateCheckPoint().
+ *
+ * Returns the end LSN of the inserted record, which the caller persists
+ * together with the new state as the data checksum watermark.
*/
-static void
+static XLogRecPtr
XLogChecksums(uint32 new_type)
{
xl_checksum_state xlrec;
@@ -8773,12 +9134,28 @@ XLogChecksums(uint32 new_type)
xlrec.new_checksum_state = new_type;
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
+
XLogBeginInsert();
XLogRegisterData((char *) &xlrec, sizeof(xl_checksum_state));
recptr = XLogInsert(RM_XLOG2_ID, XLOG2_CHECKSUMS);
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, recptr);
+
+ /* only loaded by SetDataChecksumsOn(), a no-op for the other callers */
+ INJECTION_POINT_CACHED("datachecksums-on-before-publish", NULL);
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ XLogCtl->data_checksum_version = new_type;
+ XLogCtl->data_checksum_lsn = recptr;
+ XLogCtl->data_checksum_is_local = false;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockRelease(DataChecksumTransitionLock);
+
XLogFlush(recptr);
+
+ return recptr;
}
/*
@@ -8966,11 +9343,19 @@ xlog_redo(XLogReaderState *record)
/* ControlFile->checkPointCopy always tracks the latest ckpt XID */
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->checkPointCopy.nextXid = checkPoint.nextXid;
- ControlFile->data_checksum_version = checkPoint.dataChecksumState;
UpdateControlFile();
LWLockRelease(ControlFileLock);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(checkPoint.dataChecksumState);
+
/*
* We should've already switched to the new TLI before replaying this
* record.
@@ -9206,19 +9591,17 @@ xlog_redo(XLogReaderState *record)
else if (info == XLOG_CHECKPOINT_REDO)
{
xl_checkpoint_redo redo_rec;
- bool new_state = false;
memcpy(&redo_rec, XLogRecGetData(record), sizeof(xl_checkpoint_redo));
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = redo_rec.data_checksum_version;
- SetLocalDataChecksumState(redo_rec.data_checksum_version);
- if (redo_rec.data_checksum_version != ControlFile->data_checksum_version)
- new_state = true;
- SpinLockRelease(&XLogCtl->info_lck);
-
- if (new_state)
- EmitAndWaitDataChecksumsBarrier(redo_rec.data_checksum_version);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(redo_rec.data_checksum_version,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(redo_rec.data_checksum_version);
}
else if (info == XLOG_LOGICAL_DECODING_STATUS_CHANGE)
{
@@ -9280,25 +9663,63 @@ xlog2_redo(XLogReaderState *record)
{
xl_checksum_state state;
XLogRecPtr lsn = record->EndRecPtr;
+ XLogRecPtr watermark;
memcpy(&state, XLogRecGetData(record), sizeof(xl_checksum_state));
+ SpinLockAcquire(&XLogCtl->info_lck);
+ watermark = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ /*
+ * Skip records this node has already applied. The control file
+ * carries the watermark, so this holds across restarts: recovery
+ * resuming below a record whose effect the control file already
+ * contains must not re-apply it, or it would revert a state change
+ * made with pg_checksums in between, which moves the state without
+ * writing any record of its own.
+ */
+ if (lsn <= watermark)
+ return;
+
/* advertise the location before the new state becomes visible */
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, lsn);
SpinLockAcquire(&XLogCtl->info_lck);
XLogCtl->data_checksum_version = state.new_checksum_state;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(state.new_checksum_state);
SpinLockRelease(&XLogCtl->info_lck);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = state.new_checksum_state;
+
+ /*
+ * Persist the new state, except when it is "on". Only "on" verifies
+ * checksums during reads, and between the last restartpoint and this
+ * record there may be pages on disk flushed under the old state; a
+ * crash-restart initializes verification from the control file and
+ * replay reads those pages back, so the control file may only say
+ * "on" once everything written under the transition has been flushed,
+ * as restartpoints and the end of recovery do.
+ * The opposite direction cannot wait for the restartpoint: once this
+ * record is replayed, evicted pages are written without checksums,
+ * and a control file still saying "on" would fail verification on
+ * exactly those pages after a crash.
+ */
+ if (state.new_checksum_state != PG_DATA_CHECKSUM_VERSION)
+ {
+ ControlFile->data_checksum_version = state.new_checksum_state;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ }
/*
* Update minRecoveryPoint to ensure that if recovery is aborted, we
* recover back up to this point before allowing hot standby again.
- * The new state is durable in pg_control while its location is only
- * tracked in shared memory; a standby becoming consistent below this
- * record would let base backups resume checksum verification with the
+ * The change location is only tracked in shared memory and is lost
+ * over a restart; a standby becoming consistent below this record
+ * would let base backups resume checksum verification with the
* location unknown. The local copies cannot be updated as long as
* crash recovery is happening and we expect all the WAL to be
* replayed.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 256b3a3c02e..77fa543f9d8 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -371,6 +371,7 @@ WaitLSN "Waiting to read or update shared Wait-for-LSN state."
LogicalDecodingControl "Waiting to read or update logical decoding status information."
DataChecksumsWorker "Waiting for data checksums worker."
AioWorkerControl "Waiting to update AIO worker information."
+DataChecksumTransition "Waiting for a data checksum state transition to be written to WAL."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 3b3ae23f1a6..20a99112838 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -648,6 +648,16 @@ main(int argc, char *argv[])
ControlFile->data_checksum_version =
(mode == PG_MODE_ENABLE) ? PG_DATA_CHECKSUM_VERSION : PG_DATA_CHECKSUM_OFF;
+ /*
+ * Mark the state as changed locally, without a WAL record. Recovery
+ * then knows the state is newer than anything the WAL carries and
+ * does not let a replayed checkpoint overwrite it. The watermark is
+ * left alone: any XLOG2_CHECKSUMS record this node had applied stays
+ * covered, and only records above it, written after this change, take
+ * effect again.
+ */
+ ControlFile->data_checksum_is_local = true;
+
if (do_sync)
{
pg_log_info("syncing data directory");
diff --git a/src/bin/pg_controldata/pg_controldata.c b/src/bin/pg_controldata/pg_controldata.c
index 6fc87ed114d..c363ce3dbb7 100644
--- a/src/bin/pg_controldata/pg_controldata.c
+++ b/src/bin/pg_controldata/pg_controldata.c
@@ -349,6 +349,10 @@ main(int argc, char *argv[])
(ControlFile->float8ByVal ? _("by value") : _("by reference")));
printf(_("Data page checksum version: %u\n"),
ControlFile->data_checksum_version);
+ printf(_("Data checksum watermark: %X/%08X\n"),
+ LSN_FORMAT_ARGS(ControlFile->data_checksum_lsn));
+ printf(_("Data checksum state is node-local: %s\n"),
+ (ControlFile->data_checksum_is_local ? _("yes") : _("no")));
printf(_("Default char data signedness: %s\n"),
(ControlFile->default_char_signedness ? _("signed") : _("unsigned")));
printf(_("Mock authentication nonce: %s\n"),
diff --git a/src/bin/pg_resetwal/pg_resetwal.c b/src/bin/pg_resetwal/pg_resetwal.c
index 1542a56ca4b..cdfb9898066 100644
--- a/src/bin/pg_resetwal/pg_resetwal.c
+++ b/src/bin/pg_resetwal/pg_resetwal.c
@@ -923,6 +923,13 @@ RewriteControlFile(void)
ControlFile.backupEndPoint = InvalidXLogRecPtr;
ControlFile.backupEndRequired = false;
+ /*
+ * The old WAL is gone and the new position may lie below the old
+ * watermark, which would make replay ignore future checksum transition
+ * records. The state itself is kept.
+ */
+ ControlFile.data_checksum_lsn = InvalidXLogRecPtr;
+
/*
* Force the defaults for max_* settings. The values don't really matter
* as long as wal_level='minimal'; the postmaster will reset these fields
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 2e86fd158d0..936469ea5f8 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -738,6 +738,20 @@ perform_rewind(filemap_t *filemap, rewind_source *source,
ControlFile_new.minRecoveryPoint = endrec;
ControlFile_new.minRecoveryPointTLI = endtli;
ControlFile_new.state = DB_IN_ARCHIVE_RECOVERY;
+
+ /*
+ * Keep the target's own data checksum state. Most of the data directory
+ * is still the target's: only blocks it changed since the divergence
+ * were copied from the source, so the source's state says nothing about
+ * the pages that stay. Replay from the last common checkpoint applies
+ * any WAL-logged transition the target has not seen (the watermark tells
+ * them apart), which converges the rewound server to the source's state
+ * whenever the WAL carries it.
+ */
+ ControlFile_new.data_checksum_version = ControlFile_target.data_checksum_version;
+ ControlFile_new.data_checksum_lsn = ControlFile_target.data_checksum_lsn;
+ ControlFile_new.data_checksum_is_local = ControlFile_target.data_checksum_is_local;
+
if (!dry_run)
update_controlfile(datadir_target, &ControlFile_new, do_sync);
}
diff --git a/src/bin/pg_upgrade/controldata.c b/src/bin/pg_upgrade/controldata.c
index b3bd4ccde83..de539479bdc 100644
--- a/src/bin/pg_upgrade/controldata.c
+++ b/src/bin/pg_upgrade/controldata.c
@@ -431,7 +431,7 @@ get_control_data(ClusterInfo *cluster)
cluster->controldata.date_is_int = strstr(p, "64-bit integers") != NULL;
got_date_is_int = true;
}
- else if ((p = strstr(bufin, "checksum")) != NULL)
+ else if ((p = strstr(bufin, "Data page checksum version:")) != NULL)
{
p = strchr(p, ':');
diff --git a/src/include/catalog/pg_control.h b/src/include/catalog/pg_control.h
index 7b5404460ec..f898447f195 100644
--- a/src/include/catalog/pg_control.h
+++ b/src/include/catalog/pg_control.h
@@ -22,7 +22,7 @@
/* Version identifier for this pg_control format */
-#define PG_CONTROL_VERSION 1903
+#define PG_CONTROL_VERSION 1904
/* Nonce key length, see below */
#define MOCK_AUTH_NONCE_LEN 32
@@ -237,6 +237,25 @@ typedef struct ControlFileData
/* Current data checksums state */
uint32 data_checksum_version;
+ /*
+ * End of the newest XLOG2_CHECKSUMS record this node has written or
+ * applied. Replay ignores XLOG2_CHECKSUMS records at or below this
+ * point: their effect is already contained in data_checksum_version, or
+ * an offline pg_checksums change made after they were first applied
+ * supersedes them. InvalidXLogRecPtr if the node has never written or
+ * applied such a record.
+ */
+ XLogRecPtr data_checksum_lsn;
+
+ /*
+ * True when data_checksum_version was last set by pg_checksums rather
+ * than by a WAL-logged transition. Such a state is local to this node
+ * and newer than anything the WAL carries, so recovery must not replace
+ * it with a state taken from a checkpoint record. Cleared by the next
+ * WAL-logged transition.
+ */
+ bool data_checksum_is_local;
+
/*
* True if the default signedness of char is "signed" on a platform where
* the cluster is initialized.
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index d7eb648bd27..8d858be9927 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -89,6 +89,7 @@ PG_LWLOCK(54, WaitLSN)
PG_LWLOCK(55, LogicalDecodingControl)
PG_LWLOCK(56, DataChecksumsWorker)
PG_LWLOCK(57, AioWorkerControl)
+PG_LWLOCK(58, DataChecksumTransition)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/test/modules/test_checksums/Makefile b/src/test/modules/test_checksums/Makefile
index 71455cd5577..80f54bf6d8c 100644
--- a/src/test/modules/test_checksums/Makefile
+++ b/src/test/modules/test_checksums/Makefile
@@ -9,7 +9,7 @@
#
#-------------------------------------------------------------------------
-EXTRA_INSTALL = src/test/modules/injection_points
+EXTRA_INSTALL = contrib/pg_buffercache src/test/modules/injection_points
export enable_injection_points
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index fb7129d796f..70870715976 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -35,6 +35,15 @@ tests += {
't/009_fpi.pl',
't/010_backup_straddle.pl',
't/011_standby_straddle.pl',
+ 't/012_offline_standby.pl',
+ 't/013_rewind.pl',
+ 't/014_lockstep.pl',
+ 't/015_standby_crash_after_disable.pl',
+ 't/016_promote_enable_crash.pl',
+ 't/017_restartpoint_race.pl',
+ 't/018_enable_crash_windows.pl',
+ 't/019_standby_shutdown_catchup.pl',
+ 't/020_cascade_divergence.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
new file mode 100644
index 00000000000..241a1282ae9
--- /dev/null
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -0,0 +1,289 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Offline checksum changes with pg_checksums are local to one node. A
+# standby must neither adopt the state of the primary from replayed
+# checkpoint records, nor lose its own offline change to them.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', 'autovacuum = off');
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Scenario 1: enable offline on the primary only. The standby must
+# stay off, warn about the mismatch, and remain readable.
+$standby->stop;
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+test_checksum_state($primary, 'on');
+test_checksum_state($standby, 'off');
+
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after offline enable on the primary');
+
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+my @warnings = $log =~ /(does not match the state)/g;
+is(scalar(@warnings), 1, 'mismatch warned once per remote value');
+
+# Matching states re-arm the warning: undo the divergence on the primary,
+# then diverge again to the same value, all without restarting the standby.
+$primary->stop;
+$primary->checksum_disable_offline;
+$logstart = -s $standby->logfile;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no warning while the states match again');
+like($log, qr/now agrees with the replayed WAL/,
+ 'convergence is reported once the states match again');
+
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+test_checksum_state($standby, 'off');
+
+# The local state survives both clean and immediate restarts.
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+# Still in scenario 1's divergence: take a base backup *from the
+# standby*. A backup taken on a standby uses the last restartpoint as
+# its starting checkpoint (do_pg_backup_start()), so the record at the
+# redo point was written by the upstream primary and carries the
+# primary's state. Recovery from such a backup must not adopt that
+# state: the files were copied from the standby, and the new node
+# would come up verifying checksums its files do not have.
+
+# Make sure the standby has written pages under its own "off" state, so
+# its files really do lack checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,50000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', "SELECT count(*) FROM t2;");
+
+$standby->backup('from_standby');
+my $newnode = PostgreSQL::Test::Cluster->new('newnode');
+$newnode->init_from_backup($standby, 'from_standby');
+
+# Stream from the primary, not from the standby the files were copied
+# from, so the new node replays the "on" primary's records.
+$newnode->enable_streaming($primary);
+$newnode->start;
+
+# The new node's files all came from a cluster running with checksums
+# off. Anything but "off" here means it adopted the primary's state
+# through the checkpoint record at the redo point.
+my ($rc, $stdout, $stderr) = $newnode->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'the node copied from the standby accepts connections')
+ or diag("stderr: $stderr");
+is($stdout, 'off', 'backup of an "off" standby comes up with checksums off');
+
+# A checkpoint record streamed from the "on" primary must not flip the
+# state either.
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($newnode);
+test_checksum_state($newnode, 'off');
+
+# And it must be able to read the pages the standby wrote without
+# checksums.
+(undef, $stdout, $stderr) =
+ $newnode->psql('postgres', "SELECT count(*) FROM t2;");
+is($stdout, '50000', 'pages copied from the standby are readable')
+ or diag("stderr: $stderr");
+
+my $newnode_log = PostgreSQL::Test::Utils::slurp_file($newnode->logfile);
+unlike(
+ $newnode_log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the node copied from the standby');
+
+$newnode->stop('immediate');
+
+# Converge the cluster: enable offline on the standby too.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+$primary->wait_for_catchup($standby);
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after converging');
+
+# Scenario 2: disable offline on the standby only. The replayed
+# checkpoint records of the still-enabled primary must not override it.
+$standby->stop;
+$standby->checksum_disable_offline;
+$standby->start;
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'off');
+
+# Restartpoints must persist the local state, not the replayed copy.
+$standby->safe_psql('postgres', "CHECKPOINT;");
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable with checksums disabled locally');
+
+# Scenario 3: crash-restart right after an online transition, before the
+# next restartpoint. Replay then resumes from an older restartpoint whose
+# checkpoint records still carry the pre-transition state. Those must
+# still match the state seeded from the control file, and the transition
+# itself must be re-established by re-replaying the XLOG2_CHECKSUMS record,
+# without a spurious mismatch warning along the way.
+
+# Converge first: bring the standby back to "on" offline.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+disable_data_checksums($primary, wait => 'off');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+# Crash-restart the standby immediately, before any restartpoint has had a
+# chance to persist the new state to its control file.
+$logstart = -s $standby->logfile;
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001',
+ 'standby readable after crash-restart across an online transition');
+
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no spurious mismatch warning after crash-restart across an online transition'
+);
+
+# Scenario 4: a standby stopped while replaying an interrupted online
+# transition keeps the interrupted state in its own control file. A
+# primary is never caught this way, as its checksums launcher resolves
+# inprogress-on back to off from its exit cleanup. A standby has no
+# launcher; it carries forward whatever the last replayed record left
+# it in.
+
+# Block an online enable on the primary at inprogress-on with a
+# blocking temp table, same trick as in 004_offline.pl.
+my $bsession = $primary->background_psql('postgres');
+$bsession->query_safe('CREATE TEMPORARY TABLE tt (a integer);');
+enable_data_checksums($primary, wait => 'inprogress-on');
+
+# The standby picks up the in-progress state from the XLOG2_CHECKSUMS record.
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# Stop the standby cleanly; its restartpoint persists inprogress-on to
+# its own control file, since nothing on a standby resolves it away.
+$standby->stop;
+$standby->start;
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# The primary still sits at inprogress-on, so a base backup taken now
+# copies a control file and a redo point that both carry the
+# in-progress state; replay of the XLOG2_CHECKSUMS records completes
+# the transition on the new standby.
+$primary->backup('inprogress_backup');
+my $standby2 = PostgreSQL::Test::Cluster->new('standby2');
+$standby2->init_from_backup($primary, 'inprogress_backup',
+ has_streaming => 1);
+$standby2->start;
+
+# Backup label recovery adopts the redo point's state immediately,
+# before any further WAL is replayed.
+wait_for_checksum_state($standby2, 'inprogress-on');
+
+$bsession->quit;
+wait_for_checksum_state($primary, 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable once the transition completes');
+
+# Scenario 4 continued: the backup taken mid-transition must also
+# complete the transition and stay readable.
+$primary->wait_for_catchup($standby2);
+wait_for_checksum_state($standby2, 'on');
+
+is($standby2->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'mid-transition backup standby readable after the transition');
+
+$standby2->stop;
+
+# Nothing along this standby's replay can legitimately disagree: it
+# started out at the same inprogress-on state as the primary.
+my $standby2_log = PostgreSQL::Test::Utils::slurp_file($standby2->logfile);
+unlike(
+ $standby2_log,
+ qr/does not match the state/,
+ 'no mismatch warning on a backup taken mid-transition');
+
+# No extra CHECKPOINT needed: the transition's completion checkpoint
+# wrote every page with a checksum, and the shutdown restartpoint
+# flushes the rest.
+command_ok([ 'pg_checksums', '--check', '-D', $standby2->data_dir ],
+ 'checksums valid on the mid-transition backup standby');
+
+$standby->stop;
+$primary->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/013_rewind.pl b/src/test/modules/test_checksums/t/013_rewind.pl
new file mode 100644
index 00000000000..a791e24317d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/013_rewind.pl
@@ -0,0 +1,201 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test pg_rewind across an online data checksum enable.
+#
+# A clean switchover leaves the shutdown checkpoint of the old primary
+# as the last common checkpoint between the two nodes. When data
+# checksums are enabled online on the new primary before the old one is
+# rewound, pg_rewind installs the control file of the new primary, which
+# already claims checksums are fully enabled, while replay begins at the
+# shutdown checkpoint whose record still carries the old state. Replay
+# of the WAL stretch from before the enable must run with checksums off,
+# as recorded in the checkpoint, else it would verify pages which never
+# had checksums written and fail recovery.
+#
+# The new primary requests a checkpoint right after promotion, and
+# replaying its CHECKPOINT_REDO record would repair the state before
+# any interesting WAL is reached. Hold the checkpointer on the standby
+# in a restartpoint over the promotion, like in the recovery test
+# 041_checkpoint_at_promote, so that the post-promotion writes end up
+# in WAL before the first checkpoint of the new timeline.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Old primary. full_page_writes is off so that the updates done on the
+# promoted node do not carry full page images, forcing replay on the
+# rewound node to read the pages from disk. wal_log_hints is required
+# by pg_rewind on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+full_page_writes = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+log_checkpoints = on
+]);
+$node_a->start;
+
+if (!$node_a->check_extension('injection_points'))
+{
+ plan skip_all => 'Extension injection_points not installed';
+}
+
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+$node_a->safe_psql('postgres', "CREATE EXTENSION injection_points;");
+
+# Set the hint bits on t before taking the backup, so that reads on the
+# promoted node do not emit full page images for its pages later.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->safe_psql('postgres', "SELECT count(*) FROM t;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_a, 'off');
+test_checksum_state($node_b, 'off');
+
+# Hold the next restartpoint on the standby.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point', 'wait');");
+
+# Give the restartpoint a checkpoint record to work on, then start it
+# in a background session; it will block on the injection point with
+# the checkpointer busy until released.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+my $bg_psql = $node_b->background_psql('postgres', on_error_stop => 0);
+$bg_psql->query_until(
+ qr/starting_restartpoint/, q(
+ \echo starting_restartpoint
+ CHECKPOINT;
+));
+$node_b->wait_for_event('checkpointer', 'create-restart-point');
+
+# Clean switchover: the shutdown checkpoint of A streams to B and
+# becomes the last common checkpoint.
+$node_a->stop('fast');
+
+my ($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+$stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+my $shutdown_ckpt = $1;
+$node_b->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_ckpt'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+
+my $logstart = -s $node_b->logfile;
+$node_b->promote;
+
+# Accidental restart of the old primary, diverging its timeline.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->stop('fast');
+
+# Updates on the new primary before its first checkpoint; replay of
+# these on the rewound node has to read the pages from disk.
+$node_b->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 25 = 0;");
+ok( !$node_b->log_contains("checkpoint complete", $logstart),
+ "no checkpoint on the new timeline before the updates");
+
+# Release the checkpointer; the queued post-promotion checkpoint runs
+# after the updates.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+$bg_psql->quit;
+
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+# pg_rewind refuses to run with full_page_writes disabled on the
+# source; the updates it was disabled for are already in WAL.
+$node_b->safe_psql('postgres', "ALTER SYSTEM SET full_page_writes = on;");
+$node_b->safe_psql('postgres', "SELECT pg_reload_conf();");
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the new primary');
+
+# Replay on the rewound node must start at the shutdown checkpoint of
+# the switchover, with the control file of the new primary.
+my $backup_label = slurp_file($node_a->data_dir . '/backup_label');
+$backup_label =~ /^CHECKPOINT LOCATION: ([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from backup_label";
+is($1, $shutdown_ckpt, 'replay starts at the switchover checkpoint');
+
+($stdout, $stderr) = run_command(
+ [
+ 'pg_waldump',
+ '-p' => $node_a->data_dir . '/pg_wal',
+ '-t' => 1,
+ '-s' => $shutdown_ckpt,
+ '-n' => 1,
+ ]);
+like($stdout, qr/CHECKPOINT_SHUTDOWN/,
+ 'last common checkpoint is a shutdown checkpoint');
+
+# pg_rewind keeps the target's own checksum state in the control file it
+# writes; the online enable reaches the rewound node through WAL replay
+# below, not through the copied control file.
+($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+like(
+ $stdout,
+ qr/^Data page checksum version:\s*0$/m,
+ 'rewound node keeps its own checksum state in the control file');
+
+# Start the rewound node as a standby of the new primary. Replay runs
+# through the pre-enable WAL stretch and the online enable.
+#
+# The rewind replaced the configuration files with those of the new
+# primary, so put the port back.
+my $connstr = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_b->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/014_lockstep.pl b/src/test/modules/test_checksums/t/014_lockstep.pl
new file mode 100644
index 00000000000..713f37fb22c
--- /dev/null
+++ b/src/test/modules/test_checksums/t/014_lockstep.pl
@@ -0,0 +1,182 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# The lockstep procedure for offline checksum changes in a replication
+# setup: stop all nodes, run pg_checksums on all of them, restart.
+# Replay of checkpoint records written before the change must not
+# revert the state of the standby.
+#
+# The same pair then runs the procedure in the disable direction, and
+# finally across a divergence: after a failover both nodes get their
+# checksums enabled offline, while the last common checkpoint still
+# carries "off". pg_rewind must keep the offline "on" on the rewound
+# node, since no record in the replayed WAL could ever restore it.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# wal_log_hints keeps the pair eligible for pg_rewind without checksums.
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Part 1: the lockstep procedure in the enable direction. Stop the
+# standby first: the WAL written after this point is replayed only
+# after the offline switch, and every checkpoint record in it still
+# carries the old state.
+$standby->stop;
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 0;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 1;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+# The lockstep procedure.
+$primary->checksum_enable_offline;
+$standby->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to off.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "off" in the replayed WAL/,
+ 0);
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep enable');
+
+# Crash the standby and replay the same stretch again.
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'on');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after crash restart');
+
+# Once a post-switch checkpoint has been replayed the states match and
+# no warning may be logged.
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no mismatch warning once the states match');
+
+# Every page the standby wrote in this window must carry a checksum.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$standby->stop;
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'checksums valid on the standby');
+
+# Part 2: the lockstep procedure in the other direction. The standby
+# is already stopped; give it a pre-switch WAL stretch to replay, with
+# every checkpoint record in it still carrying "on" -- an explicit
+# checkpoint plus the shutdown checkpoint written by stop().
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 2;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+$primary->checksum_disable_offline;
+$standby->checksum_disable_offline;
+
+my $disable_logstart = -s $standby->logfile;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to on.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $disable_logstart);
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep disable');
+
+# Part 3: failover, divergence and pg_rewind across offline enables.
+# The roles swap from here on: the standby becomes the new primary and
+# the old primary is rewound to follow it, but the variables keep the
+# names they had above.
+#
+# Failover: promote the standby, then diverge the old primary.
+$standby->promote;
+$standby->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$primary->stop;
+
+# The lockstep procedure across the divergence: both nodes get their
+# checksums enabled offline.
+$primary->checksum_enable_offline;
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+
+# The states match, so the rewind proceeds without complaint.
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $primary->data_dir,
+ '--source-server' => $standby->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled offline on both nodes');
+
+# The rewind replaced the configuration files with those of the source,
+# so put the port back. The copied file also carries the source's own
+# stale primary_conninfo; enable_streaming() below appends ours, which
+# wins as the later entry.
+$primary->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$primary->port]}
+]);
+$primary->enable_streaming($standby);
+
+my $rewind_logstart = -s $primary->logfile;
+$primary->start;
+$standby->wait_for_catchup($primary);
+
+is($primary->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+
+# The offline enable survives: replay from the common checkpoint must
+# not resurrect the pre-divergence "off".
+test_checksum_state($primary, 'on');
+
+$log =
+ PostgreSQL::Test::Utils::slurp_file($primary->logfile, $rewind_logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no divergence reported between the rewound server and its source');
+
+$primary->stop;
+$standby->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
new file mode 100644
index 00000000000..315b27b93d4
--- /dev/null
+++ b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
@@ -0,0 +1,125 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that a standby crashing after a replayed online disable, but before
+# its next restartpoint, restarts cleanly.
+#
+# Pages dirtied before the XLOG2_CHECKSUMS record and evicted after it are
+# written out under the new "off" state, without checksums. The control
+# file must follow the record immediately in this direction: a crash-restart
+# initializes checksum verification from the control file, and one still
+# saying "on" would fail verification on exactly those pages while replaying
+# records older than the state change.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$primary->start;
+
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,1000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+
+# Small buffer pool so that replaying a bulk load evicts the pages we care
+# about, and no restartpoints so the control file stays where it is.
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 1MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_delay = 10000
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'on');
+
+# No full page images, so that replay has to read the pages of "t" back from
+# disk instead of overwriting them from the WAL.
+$primary->append_conf('postgresql.conf', 'full_page_writes = off');
+$primary->reload;
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" on the standby, *before* the state change. They are
+# not flushed: the standby has no restartpoint from here on.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+# Online disable. The standby replays it and switches to "off", but its
+# control file is not updated.
+$primary->safe_psql('postgres', 'SELECT pg_disable_data_checksums();');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+my ($ctl_before) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($ctl_state) = $ctl_before =~ /Data page checksum version:\s+(\d+)/;
+note(
+ "standby control file data_checksum_version after the disable: "
+ . "$ctl_state (0 = off, 1 = on)");
+
+# Force the standby to evict the dirty pages of "t" now that it is "off":
+# they are written back without checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE filler AS SELECT generate_series(1,300000) AS a;");
+$primary->wait_for_catchup($standby);
+
+# Crash the standby before it gets a chance to run a restartpoint.
+$standby->stop('immediate');
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'standby restarts after crashing between a checksum state '
+ . 'change and the next restartpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'the pages written while "off" are readable')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("standby log tail:\n" . join("\n", @lines[ -12 .. -1 ]))
+ if @lines >= 12;
+ fail('the pages written while "off" are readable');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/016_promote_enable_crash.pl b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
new file mode 100644
index 00000000000..97f84da5685
--- /dev/null
+++ b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a standby promoted right after replaying an online enable, before any
+# restartpoint has flushed the rewritten pages.
+#
+# The end-of-recovery record persists the checksum state without flushing the
+# buffer pool, while the control file still points at a restartpoint older than
+# the transition. Persisting "on" there would make a crash before the
+# post-promotion checkpoint verify checksums over the pages that were flushed
+# while the state was still "off"; promotion must take the full checkpoint
+# instead.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Dirty the pages of "t" without full page images, then push them out to disk
+# on the standby while checksums are still off.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Online enable on the primary; the standby replays it into its own buffers,
+# where the rewritten pages stay dirty (no restartpoint, no bgwriter).
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+note("standby control file before the promotion: $before");
+
+# Promote. The end-of-recovery record persists the live state without any
+# buffer flush; the checkpoint requested afterwards is not immediate.
+$standby->promote;
+$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery();')
+ or die 'timed out waiting for the promotion';
+$standby->stop('immediate');
+
+($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($ckpt) = $ctl =~ /Latest checkpoint location:\s+(\S+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note(
+ "standby control file after the promotion: $after, checkpoint $ckpt, redo $redo"
+);
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'promoted node restarts after crashing before its first checkpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/017_restartpoint_race.pl b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
new file mode 100644
index 00000000000..65aa834f366
--- /dev/null
+++ b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
@@ -0,0 +1,138 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a restartpoint whose flush races the replay of an online enable.
+#
+# Replay keeps running while CheckPointGuts() writes out the buffer pool, so
+# the state at the end of the flush can be newer than the one the flush ran
+# under. Persisting it would claim checksums for pages the flush wrote out
+# before the transition, against a redo pointer that predates it.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# C1: the checkpoint the pending restartpoint will target.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" without full page images and push them to disk on
+# the standby while it is still "off".
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Start a restartpoint for C1 and hold it right after CheckPointGuts().
+$standby->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point','wait');");
+my $bg = $standby->background_psql('postgres');
+$bg->query_until(qr//, "\\echo restartpoint\nCHECKPOINT;\n");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint injection point';
+
+# The startup process keeps replaying while the checkpointer is held: the
+# whole online enable lands, and the rewritten pages stay dirty.
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+# Let the restartpoint finish; it persists the state it samples now.
+$standby->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint to finish';
+$standby->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+
+$standby->stop('immediate');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note("standby control file after the restartpoint: $after, redo $redo");
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started, 'standby restarts after a restartpoint that raced the enable');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/018_enable_crash_windows.pl b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
new file mode 100644
index 00000000000..17f96ff6272
--- /dev/null
+++ b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
@@ -0,0 +1,496 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Crash and checkpoint windows inside an online data checksum enable on
+# a single node. Each scenario starts from "off" on the same cluster,
+# runs the enable into a held injection point, and checks what a crash
+# or a concurrent checkpoint in that window may and may not do to the
+# control file and the pages on disk.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+use IPC::Run;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Bring the node back to "off" with no leftover table between
+# scenarios. The crash restarts already resolve an interrupted enable
+# to "off", so only disable when something is left to disable. Those
+# restarts also clear any attached injection points. The checkpoint
+# bounds the WAL the next scenario's crash recovery has to replay.
+sub reset_scenario
+{
+ my ($node) = @_;
+
+ BAIL_OUT('node is not running after the previous scenario')
+ unless $node->is_alive;
+
+ my $state = $node->safe_psql('postgres', 'SHOW data_checksums;');
+ disable_data_checksums($node, wait => 'off') if $state ne 'off';
+ $node->safe_psql('postgres', 'DROP TABLE IF EXISTS t;');
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+ test_checksum_state($node, 'off');
+}
+
+my $node = PostgreSQL::Test::Cluster->new('crash_windows');
+$node->init(no_data_checksums => 1);
+
+# Scenario 1 needs the pages of "t" to reach disk without a full page
+# image, hence full_page_writes and wal_log_hints off. Scenario 2
+# needs them to stay dirty in shared buffers until a checkpoint, hence
+# no bgwriter and a large shared_buffers. wal_level is pinned because
+# init() writes "minimal" for a non-streaming node. The rest merely
+# tolerate the settings.
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+wal_level = replica
+max_wal_size = 10GB
+full_page_writes = off
+wal_log_hints = off
+shared_buffers = 512MB
+bgwriter_lru_maxpages = 0
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$node->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+
+test_checksum_state($node, 'off');
+
+# Scenario 1: a primary crashes inside SetDataChecksumsOn(), just before the
+# forced checkpoint that flushes the rewritten pages, with crash recovery
+# resuming from a checkpoint older than the transition. The control file
+# must still say "inprogress-on" there: replay does not adopt the state of
+# the checkpoint record it resumes from, so an "on" written before the flush
+# would stay in effect while replay reads pages whose rewrite never reached
+# disk. Here the pages went out through the rewriting worker's ring buffer,
+# so only the control file state discriminates; see the next scenario for
+# the case where they did not.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Establish the checkpoint that crash recovery will resume from.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty the pages of "t" without full page images, then push them out
+ # to disk while checksums are still off.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+ my $evicted = $node->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+ note("evict_relation on t: $evicted, relpath $relpath");
+
+ # Stop the enable right after the control file has been updated to
+ # "on" but before the checkpoint that flushes the rewritten pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+ $node->safe_psql('postgres', 'SELECT pg_enable_data_checksums();');
+
+ $node->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'datachecksums-on-before-checkpoint';"
+ ) or die 'timed out waiting for the injection point';
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ $node->stop('immediate');
+
+ # Show the on-disk checksum field of the first page of "t".
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 2: the same crash as above, but with the rewritten pages resident
+# in shared buffers instead of gone out through the ring. The rewriting
+# worker reads through a BAS_VACUUM ring, which writes pages back as the
+# ring recycles, but a page already resident in shared buffers is not read
+# through the ring: ReadBufferExtended() hands back the existing buffer, and
+# it stays dirty until a checkpoint. The control file may therefore not
+# say "on" before that checkpoint has run, or crash recovery would resume
+# from a checkpoint older than the transition with verification already
+# enabled, and replay records that read those still-unchecksummed pages.
+# full_page_writes is off so that the records do not simply overwrite the
+# pages with a full page image.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,100000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Establish the checkpoint that crash recovery will resume from, with the
+ # pages of "t" written out while checksums are still off.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty those pages again without emitting full page images, and leave
+ # them in shared buffers. Replay of these records has to read the
+ # pages from disk.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+
+ my $dirty_before = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_before, '>', 0,
+ 'pages of t are resident and dirty before enabling checksums');
+
+ # Hold the enabling right before the checkpoint that flushes the rewritten
+ # pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ # The rewritten pages must still be sitting dirty in shared buffers, or
+ # the window this test is about does not exist.
+ my $dirty_after = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_after, '>', 0,
+ 'rewritten pages of t are still dirty before the checkpoint');
+
+ $node->stop('immediate');
+
+ # The on-disk copy is the one written before the transition, without a
+ # checksum.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ is($pd_checksum, 0, 'block 0 of t on disk carries no checksum');
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 3: a checkpoint that runs between the XLOG2_CHECKSUMS("on")
+# record and the control file write at the end of SetDataChecksumsOn() must
+# not make a crash throw the completed transition away. SetDataChecksumsOn()
+# writes the record, flips shared memory to "on", emits the barrier and only
+# then requests the checkpoint that flushes the rewritten pages; the control
+# file is written after that checkpoint returns. A crash in that window is
+# harmless only as long as recovery still starts before the record. Any
+# checkpoint completing in the window moves the redo point past the record,
+# so recovery would never see it, would come up with the control file's
+# "inprogress-on" and StartupXLOG() would demote that to "off", even though
+# every page on disk carries a checksum by then. CreateCheckPoint()
+# therefore persists the state the checkpoint ran under, the same way
+# CreateRestartPoint() does. Here the window is made deterministic by
+# holding the launcher at the datachecksums-on-before-checkpoint injection
+# point and checkpointing from another session.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Hold the launcher after the record, the shared memory flip and the
+ # barrier, but before the checkpoint SetDataChecksumsOn() requests
+ # itself.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ # Every backend already sees "on" and writes checksums.
+ test_checksum_state($node, 'on');
+
+ # ... while the control file still says "inprogress-on".
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($before, '3', 'control file says "inprogress-on" inside the window');
+
+ # A concurrent checkpoint. It flushes the rewritten pages and moves
+ # the redo point past the XLOG2_CHECKSUMS("on") record, so it has to
+ # record the "on" state in the control file as well.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($after, '1',
+ 'the concurrent checkpoint records "on" in the control file');
+
+ $node->stop('immediate');
+
+ # The checkpoint flushed the rewritten pages, so they carry a checksum on
+ # disk: the transition really did complete.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ isnt($pd_checksum, 0, 'block 0 of t on disk carries a checksum');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition is complete on disk, so the cluster has to come back
+ # "on".
+ test_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+}
+reset_scenario($node);
+
+# Scenario 4: a crash between the checkpoint an online enable requests and
+# the control file write that follows it must not lose the transition. The
+# last steps of SetDataChecksumsOn() are
+#
+# WAL record -> shmem -> barrier -> checkpoint -> persist
+#
+# The checkpoint is what makes "on" safe to persist, but it also moves the
+# redo point above the XLOG2_CHECKSUMS record that carries the new state.
+# Crash recovery started from that checkpoint therefore never replays the
+# record, so without the state the checkpoint itself persists, a crash in
+# the remaining window would bring the cluster back at "inprogress-on",
+# which StartupXLOG() resolves to "off", discarding a transition whose
+# pages are all on disk with a checksum. The previous scenario exercises
+# the same window through a checkpoint requested by another session; this
+# one closes it from the other side: the transition's own checkpoint is the
+# one that moves the redo point, so persisting the state from
+# CreateCheckPoint() is what has to save it, not the write below the
+# injection point.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Hold the launcher after the checkpoint that licenses "on" has
+ # completed and before the state reaches the control file.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-after-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-after-checkpoint');
+
+ # The transition is complete as far as the running cluster is concerned.
+ test_checksum_state($node, 'on');
+
+ # The checkpoint has already recorded it, so the pending write below the
+ # injection point has nothing left to do.
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($version) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($version, '1',
+ 'the requested checkpoint recorded "on" in the control file');
+
+ # Crash before the control file write that follows the checkpoint.
+ $node->stop('immediate');
+
+ $node->start;
+
+ # Every page on disk carries a checksum and the checkpoint that
+ # flushed them completed, so the cluster has to come back verifying
+ # them.
+ test_checksum_state($node, 'on');
+
+ is($node->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '10000', 'relation readable after the crash');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ # The data directory must be one the offline tools accept.
+ $node->stop;
+ $node->command_ok([ 'pg_checksums', '--check', '-D', $node->data_dir ],
+ 'pg_checksums accepts the data directory');
+
+ # Leave the node running for the reset.
+ $node->start;
+}
+reset_scenario($node);
+
+# Scenario 5: a checkpoint racing SetDataChecksumsOn() between the insertion
+# of the XLOG2_CHECKSUMS("on") record and the shared memory update must not
+# insert an XLOG_CHECKPOINT_REDO record that follows the transition in WAL
+# order while still carrying "inprogress-on". Recovery resuming from such a
+# redo point never replays the preceding transition record, comes up in
+# "inprogress-on" and resolves the finished transition as interrupted.
+# XLogChecksums() closes the window by inserting the record and publishing
+# the new state under DataChecksumTransitionLock, which CreateCheckPoint()
+# takes around sampling the state and inserting the redo record. Here the
+# launcher is held between the two steps at the
+# datachecksums-on-before-publish injection point, and a concurrent
+# CHECKPOINT has to block on the lock instead of completing inside the
+# window.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # The datachecksums-on-before-publish point fires inside a critical
+ # section, where the wait machinery must not allocate. Waiting once
+ # at the datachecksums-enable-checksums-delay point, which the
+ # launcher runs outside the critical section, initializes it; see
+ # 050_redo_segment_missing.pl for the same recipe around
+ # create-checkpoint-run.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-enable-checksums-delay','wait');"
+ );
+
+ # Hold the launcher after the XLOG2_CHECKSUMS("on") record is in WAL but
+ # before the new state is published in shared memory.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-publish','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-enable-checksums-delay');
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-enable-checksums-delay');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-enable-checksums-delay');");
+
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-publish');
+
+ # The record is in WAL, the published state is still the old one.
+ test_checksum_state($node, 'inprogress-on');
+
+ # A concurrent checkpoint. It must block on DataChecksumTransitionLock
+ # before inserting its redo record rather than complete inside the window.
+ my $checkpointer = IPC::Run::start(
+ [ 'psql', '-XAtq', '-d', $node->connstr('postgres'), '-c', 'CHECKPOINT;' ],
+ '<' => '/dev/null',
+ '>' => '/dev/null',
+ '2>' => '/dev/null',
+ IPC::Run::timer($PostgreSQL::Test::Utils::timeout_default));
+
+ ok( $node->poll_query_until(
+ 'postgres',
+ "SELECT wait_event = 'DataChecksumTransition' "
+ . "FROM pg_stat_activity WHERE backend_type = 'checkpointer';"),
+ 'concurrent checkpoint blocks on DataChecksumTransitionLock');
+
+ # Release the launcher; the checkpoint then samples the published "on" and
+ # its redo record follows the transition record.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-publish');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-publish');");
+
+ $checkpointer->finish;
+
+ # Crash while the launcher's own checkpoint may still be in flight.
+ # Recovery resumes from the concurrent checkpoint's redo point, which
+ # now lies above the transition record and carries "on".
+ $node->stop('immediate');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition completed, so the cluster has to come back "on".
+ wait_for_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ $node->stop;
+}
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
new file mode 100644
index 00000000000..aaa35f8f7dd
--- /dev/null
+++ b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
@@ -0,0 +1,126 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A standby stopped after replaying an online enable, but before the
+# checkpoint record that follows it on the primary.
+#
+# The shutdown restartpoint has no new checkpoint record to work from
+# and is skipped, so the control file would keep the in-progress state
+# the transition passed through, even though replay left the node at
+# "on" and rewrote every page. pg_checksums reads that field and
+# refuses to run on an in-progress state, so the shutdown has to flush
+# and catch it up instead.
+#
+# The caught-up control file then carries the watermark of the "on"
+# record, which the second half of this test relies on: an offline
+# disable made while the standby is down must survive the re-replay of
+# that record on the next startup, or the change would be silently
+# reverted.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+wal_keep_size = 1GB
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Anchor the standby's last restartpoint here, so that nothing replayed
+# from now on gives the shutdown restartpoint a newer checkpoint record
+# to use.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Hold the enable right after the state change record, before the
+# checkpoint it requests once the transition is complete.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+enable_data_checksums($primary);
+$primary->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+wait_for_checksum_state($primary, 'on');
+
+# The state change record is already flushed and streams on its own;
+# no checkpoint record follows while the launcher is held.
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+$standby->stop('fast');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($state, '1',
+ 'shutdown catches the control file up with the replayed state');
+
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'pg_checksums verifies the stopped standby');
+
+# Disable checksums offline while the standby is down. The restart
+# below resumes replay under the last restartpoint, re-reading the "on"
+# record; the watermark persisted with the catch-up above marks it as
+# already applied, so the offline change must survive.
+$standby->checksum_disable_offline;
+
+# Let the enable finish on the primary.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-checkpoint');");
+$primary->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-checkpoint');");
+
+# The launcher exits only after the checkpoint it requests completes,
+# so a checkpoint record now follows the transition in WAL.
+$primary->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_catalog.pg_stat_activity "
+ . "WHERE backend_type = 'datachecksums launcher';");
+
+my $log_offset = -s $standby->logfile;
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+
+# The nodes legitimately diverged, which replayed checkpoints report.
+$standby->wait_for_log(
+ qr/data checksum state "off" of this node does not match the state "on" in the replayed WAL/,
+ $log_offset);
+
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/020_cascade_divergence.pl b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
new file mode 100644
index 00000000000..b4b474fa583
--- /dev/null
+++ b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
@@ -0,0 +1,115 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# An offline data checksum state change made on one node of a cascading setup
+# is reported all the way down the chain.
+#
+# CheckReplayedDataChecksumState() is reached from xlog_redo() when a
+# checkpoint record (XLOG_CHECKPOINT_SHUTDOWN or XLOG_CHECKPOINT_REDO) is
+# replayed after consistency has been reached. Those records are written by
+# the root primary only and are relayed verbatim by every intermediate
+# standby, so a cascaded standby that was changed offline learns about the
+# divergence too. An intermediate standby's own restartpoints are not
+# WAL-logged and cannot surface it.
+#
+# Note that this makes detection checkpoint-driven: an idle or read-only
+# primary surfaces the divergence no sooner than its next checkpoint, so up to
+# checkpoint_timeout may pass before the operator sees anything. Closing that
+# window would require comparing the state when a walreceiver connects.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# checkpoint_timeout is deliberately long so that the only checkpoint record in
+# this test is the one requested explicitly below.
+my $conf = qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+);
+
+my $primary = PostgreSQL::Test::Cluster->new('cascade_primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', $conf);
+$primary->start;
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,1000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('cascade_standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf('postgresql.conf', $conf);
+$standby->start;
+
+$standby->backup('backup');
+my $cascade = PostgreSQL::Test::Cluster->new('cascade_cascade');
+$cascade->init_from_backup($standby, 'backup', has_streaming => 1);
+$cascade->append_conf('postgresql.conf', $conf);
+$cascade->start;
+
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+test_checksum_state($cascade, 'off');
+
+# Change the cascaded standby offline, and only it.
+$cascade->stop;
+$cascade->command_ok([ 'pg_checksums', '--enable', '-D', $cascade->data_dir ],
+ 'pg_checksums enables checksums on the cascaded standby only');
+
+my $logstart = -s $cascade->logfile;
+$cascade->start;
+
+test_checksum_state($cascade, 'on');
+
+# Ordinary WAL traffic carries no state, so the cascaded standby keeps
+# streaming from a chain whose state is "off" without noticing.
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (1);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+is($cascade->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '1001', 'cascaded standby keeps streaming from a divergent chain');
+
+# Restartpoints on the intermediate standby are not WAL-logged, so this does
+# not surface anything on the cascaded standby either.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (2);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+my $log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+unlike(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'an intermediate restartpoint does not surface the divergence');
+
+# A checkpoint on the root primary does: the record is relayed down the whole
+# chain, so both the direct and the cascaded standby compare it against their
+# own state. wait_for_catchup() below waits for replay, not just receipt, so
+# the record has been through xlog_redo() on both nodes once it returns.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+$log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+like(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'the cascaded standby reports the divergence on the primary checkpoint');
+
+$cascade->stop;
+$standby->stop;
+$primary->stop;
+
+done_testing();
--
2.55.0
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-31 15:28 Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
parent: Zsolt Parragi <zsolt.parragi@percona.com>
0 siblings, 1 reply; 43+ messages in thread
From: Bertrand Drouvot @ 2026-08-31 15:28 UTC (permalink / raw)
To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; pgsql-hackers@lists.postgresql.org
Hi,
On Mon, Aug 31, 2026 at 02:10:15PM +0100, Zsolt Parragi wrote:
> > Thanks! I did not look in details but it looks like that pg_upgrade --check against
> > a running source cluster with checksums enabled will still fail (finding === 2 in
> > [0], tested with the test shared in [1]).
>
> Yes, I missed that in the previous version, v7 fixes it.
Thanks, yeah it fixes it.
=== 1
The v7-0001 commit message says:
"
pg_rewind keeps the target's own state, watermark and flag in the
control file it installs, since most of the data directory remains the
target's; replay from the last common checkpoint still adopts a state
its watermark does not cover, and applies any online transition the
target has not seen.
"
I found a case where the source's online enable occurs after divergence and was
never seen by the target, but replay skips it instead of applying it:
1. Start a primary and standby with checksums off.
2. Stop the primary and promote the standby.
3. Enable checksums online on the promoted source.
4. Restart the old primary on the old timeline, advance its WAL beyond the source’s
enable watermark, then enable and disable checksums online.
5. Rewind the old primary from the promoted source.
The target is now off, with a watermark on the old timeline numerically greater
than the source's enable records on the new timeline. pg_rewind preserves that
watermark, so recovery treats those source records as already applied and skips
them.
Maybe the watermark needs timeline context, or pg_rewind needs to adjust it
when it comes from the target's divergent history?
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-08-31 23:31 Zsolt Parragi <zsolt.parragi@percona.com>
parent: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 1 reply; 43+ messages in thread
From: Zsolt Parragi @ 2026-08-31 23:31 UTC (permalink / raw)
To: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; pgsql-hackers@lists.postgresql.org
> I found a case where the source's online enable occurs after divergence and was
> never seen by the target, but replay skips it instead of applying it:
> ....
> Maybe the watermark needs timeline context, or pg_rewind needs to adjust it
> when it comes from the target's divergent history?
Thanks! v8 adds the latter, with a new test case verifying this scenario.
Attachments:
[application/octet-stream] v8-0002-pg_checksums-Refuse-interrupted-transitions-note-.patch (7.7K, ../../CAN4CZFN5sOQvWTXr7Dm1J1HXwczikBfMjkbFQ7QmHsr+VHd5Mg@mail.gmail.com/2-v8-0002-pg_checksums-Refuse-interrupted-transitions-note-.patch)
download | inline diff:
From 4cb3ca4332e8cae5973311b4d3fb0c089337a817 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:19 +0000
Subject: [PATCH v8 2/5] pg_checksums: Refuse interrupted transitions, note
change is local
An inprogress-on or inprogress-off control file means an online
transition was cut short; running the offline tool on top of it mixes
two procedures. Refuse it in every mode, with a hint pointing at the
start-stop cycle (or, on a standby, at letting replication finish)
that resets the state. Also print that the change applies to one
data directory only, with standby-aware wording, since in a
replication setup the same change must be made on every node.
On a primary, inprogress-on cannot survive a graceful stop: the
checksums launcher resolves it from its exit cleanup, and a crashed
primary is already rejected by the existing "cluster must be shut
down" check. inprogress-off can, however: it is set by the backend
running pg_disable_data_checksums(), and a fast shutdown arriving
between its two barriers leaves it behind in a cleanly shut down
control file, where the next start-stop cycle resolves it at end of
recovery. A standby can be stopped with either state, since it has
no launcher and only carries forward whatever state the last replayed
WAL record left it in, with a restartpoint persisting that as-is.
The standby is also the deterministic way to reach the new guard,
which is why the test coverage uses one; the primary window would
need an injection point between the two barriers.
The pg_checksums page said the tool still processes all relation files
regardless of an interrupted online transition; describe the refusal
instead.
---
doc/src/sgml/ref/pg_checksums.sgml | 11 ++++---
src/bin/pg_checksums/pg_checksums.c | 21 +++++++++++++
.../test_checksums/t/012_offline_standby.pl | 31 ++++++++++++++-----
3 files changed, 51 insertions(+), 12 deletions(-)
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index bf07da09754..000940e7cb7 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -49,11 +49,12 @@ PostgreSQL documentation
</para>
<para>
- When enabling checksums with <application>pg_checksums</application>, if
- checksums were in the process of being enabled using
- <xref linkend="checksums-online-enable-disable"/> when the cluster was shut
- down, <application>pg_checksums</application> will still process all
- relation files regardless of the progress of online checksum processing.
+ If checksums were in the process of being enabled or disabled using
+ <xref linkend="checksums-online-enable-disable"/> when the cluster was
+ shut down, the control file still records that interrupted state, and
+ <application>pg_checksums</application> refuses to run in any mode.
+ Start the cluster and shut it down cleanly to reset the state, then
+ retry; on a standby, let replication complete the transition first.
</para>
<para>
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 20a99112838..a49b6f368b7 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -586,6 +586,21 @@ main(int argc, char *argv[])
ControlFile->state != DB_SHUTDOWNED_IN_RECOVERY)
pg_fatal("cluster must be shut down");
+ /*
+ * An inprogress state means an online transition was cut short. A
+ * standby stopped mid-transition carries either state; a cleanly shut
+ * down primary can still carry inprogress-off, which a fast shutdown
+ * during pg_disable_data_checksums() leaves behind, while inprogress-on
+ * is always resolved by the launcher's exit cleanup.
+ */
+ if (ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted");
+ pg_log_error_hint("Start and cleanly shut down the cluster once to reset the data checksum state, then retry. On a standby, let replication complete the transition first.");
+ exit(1);
+ }
+
if (ControlFile->data_checksum_version != PG_DATA_CHECKSUM_VERSION &&
mode == PG_MODE_CHECK)
pg_fatal("data checksums are not enabled in cluster");
@@ -673,6 +688,12 @@ main(int argc, char *argv[])
printf(_("Checksums enabled in cluster\n"));
else
printf(_("Checksums disabled in cluster\n"));
+
+ printf(_("This change applies to this data directory only.\n"));
+ if (ControlFile->state == DB_SHUTDOWNED_IN_RECOVERY)
+ printf(_("This node appears to be a standby; apply the same change to the primary and all other standbys.\n"));
+ else
+ printf(_("In a replication setup, apply the same change to every node while all are stopped, before restarting any of them.\n"));
}
return 0;
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
index 241a1282ae9..5efe11f993e 100644
--- a/src/test/modules/test_checksums/t/012_offline_standby.pl
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -149,7 +149,10 @@ $newnode->stop('immediate');
# Converge the cluster: enable offline on the standby too.
$standby->stop;
-$standby->checksum_enable_offline;
+command_checks_all(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ 0, [qr/appears to be a standby/],
+ [], 'standby-role notice on offline enable');
$standby->start;
test_checksum_state($standby, 'on');
$primary->wait_for_catchup($standby);
@@ -217,11 +220,11 @@ unlike(
);
# Scenario 4: a standby stopped while replaying an interrupted online
-# transition keeps the interrupted state in its own control file. A
-# primary is never caught this way, as its checksums launcher resolves
-# inprogress-on back to off from its exit cleanup. A standby has no
-# launcher; it carries forward whatever the last replayed record left
-# it in.
+# transition keeps the interrupted state in its own control file, and
+# pg_checksums must refuse to touch it. A primary is never caught this
+# way, as its checksums launcher resolves inprogress-on back to off from
+# its exit cleanup. A standby has no launcher; it carries forward
+# whatever the last replayed record left it in.
# Block an online enable on the primary at inprogress-on with a
# blocking temp table, same trick as in 004_offline.pl.
@@ -235,6 +238,20 @@ wait_for_checksum_state($standby, 'inprogress-on');
# Stop the standby cleanly; its restartpoint persists inprogress-on to
# its own control file, since nothing on a standby resolves it away.
$standby->stop;
+
+command_fails_like(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --enable refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --check refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--disable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --disable refuses a standby stopped mid-transition');
+
$standby->start;
wait_for_checksum_state($standby, 'inprogress-on');
@@ -257,7 +274,7 @@ wait_for_checksum_state($primary, 'on');
$primary->wait_for_catchup($standby);
wait_for_checksum_state($standby, 'on');
-is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
'10001', 'standby readable once the transition completes');
# Scenario 4 continued: the backup taken mid-transition must also
--
2.55.0
[application/octet-stream] v8-0004-pg_combinebackup-Refuse-mixed-data-checksum-state.patch (9.5K, ../../CAN4CZFN5sOQvWTXr7Dm1J1HXwczikBfMjkbFQ7QmHsr+VHd5Mg@mail.gmail.com/3-v8-0004-pg_combinebackup-Refuse-mixed-data-checksum-state.patch)
download | inline diff:
From 9fd27ce2386a164230e7194a9b5ae2abfedd745f Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Tue, 25 Aug 2026 15:06:58 +0000
Subject: [PATCH v8 4/5] pg_combinebackup: Refuse mixed data checksum states in
a backup chain
check_control_files() detected a chain whose backups were taken under
different data checksum states, warned, and proceeded. The output
directory keeps the last backup's control file, so a full backup taken
with checksums off combined with an incremental taken after an offline
enable produces a cluster whose control file says "on" while most of
its blocks carry no checksums; it starts, and then every connection
dies on the first unchecksummed catalog page. An offline enable
between two backups of a chain is all it takes, since it rewrites
every page without logging anything, so the incremental backup does
not re-ship the pages.
Turn the warning into an error, matching what pg_rewind does for the
equivalent combinations. The check stays asymmetric on purpose: when
the last backup was taken with checksums off, stale checksums from an
earlier backup are never verified, and that chain remains usable.
---
doc/src/sgml/ref/pg_combinebackup.sgml | 19 +--
src/bin/pg_combinebackup/pg_combinebackup.c | 14 +-
src/test/modules/test_checksums/meson.build | 1 +
.../t/024_combinebackup_mixed.pl | 134 ++++++++++++++++++
4 files changed, 155 insertions(+), 13 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
diff --git a/doc/src/sgml/ref/pg_combinebackup.sgml b/doc/src/sgml/ref/pg_combinebackup.sgml
index 9a6d201e0b8..f5d7e177b36 100644
--- a/doc/src/sgml/ref/pg_combinebackup.sgml
+++ b/doc/src/sgml/ref/pg_combinebackup.sgml
@@ -306,18 +306,19 @@ PostgreSQL documentation
<para>
<literal>pg_combinebackup</literal> does not recompute page checksums when
- writing the output directory. Therefore, if any of the backups used for
- reconstruction were taken with checksums disabled, but the final backup was
- taken with checksums enabled, the resulting directory may contain pages
- with invalid checksums.
+ writing the output directory. It therefore refuses a chain in which some
+ of the backups used for reconstruction were taken with checksums disabled
+ but the final backup was taken with checksums enabled: the resulting
+ directory would contain pages with invalid checksums, and the cluster
+ would fail checksum verification as soon as it reads them. Take a new
+ full backup after enabling data checksums with
+ <xref linkend="app-pgchecksums"/>.
</para>
<para>
- To avoid this problem, taking a new full backup after changing the checksum
- state of the cluster using <xref linkend="app-pgchecksums"/> is
- recommended. Otherwise, you can disable and then optionally reenable
- checksums on the directory produced by <literal>pg_combinebackup</literal>
- in order to correct the problem.
+ The reverse case is accepted: when the final backup was taken with
+ checksums disabled, stale checksums copied from an earlier backup are
+ never verified.
</para>
</refsect1>
diff --git a/src/bin/pg_combinebackup/pg_combinebackup.c b/src/bin/pg_combinebackup/pg_combinebackup.c
index 86e2ee37c40..254a27b125b 100644
--- a/src/bin/pg_combinebackup/pg_combinebackup.c
+++ b/src/bin/pg_combinebackup/pg_combinebackup.c
@@ -670,13 +670,19 @@ check_control_files(int n_backups, char **backup_dirs)
pg_log_debug("system identifier is %" PRIu64, system_identifier);
/*
- * Warn the user if not all backups are in the same state with regards to
- * checksums.
+ * Reject a chain whose backups were not all in the same state with
+ * regards to checksums. The loop above only flags that when the last
+ * backup has checksums enabled: the blocks taken from the older backups
+ * have no checksums, and pg_combinebackup does not recompute them, so the
+ * combined cluster would fail verification as soon as it reads them. The
+ * other direction is harmless, since stale checksums are never verified
+ * while checksums are disabled.
*/
if (data_checksum_mismatch)
{
- pg_log_warning("only some backups have checksums enabled");
- pg_log_warning_hint("Disable, and optionally reenable, checksums on the output directory to avoid failures.");
+ pg_log_error("only some backups have checksums enabled");
+ pg_log_error_hint("Take a new full backup after changing the data checksum state with pg_checksums.");
+ exit(1);
}
return system_identifier;
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index 7d07c757052..7eccd5156b9 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -47,6 +47,7 @@ tests += {
't/021_rewind_divergent_transitions.pl',
't/022_rewind_state.pl',
't/023_rewind_standby_target.pl',
+ 't/024_combinebackup_mixed.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
new file mode 100644
index 00000000000..ce9dee9413b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_combinebackup refuses a chain whose backups were taken under
+# different data checksum states when the final state has them enabled.
+#
+# The output directory keeps the last backup's control file, so a full
+# backup taken with checksums off plus an incremental taken after an
+# offline enable would produce a directory whose control file says "on"
+# while most of its blocks have no checksums. An offline enable between
+# the two backups is enough to get there: it rewrites every page but logs
+# nothing, so the incremental backup does not re-ship the pages.
+#
+# The reverse order stays allowed: with the final backup taken with
+# checksums off, stale checksums from an earlier backup are never verified.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $node = PostgreSQL::Test::Cluster->new('node');
+$node->init(
+ no_data_checksums => 1,
+ has_archiving => 1,
+ allows_streaming => 1);
+$node->append_conf('postgresql.conf', 'summarize_wal = on');
+$node->append_conf('postgresql.conf', 'autovacuum = off');
+$node->start;
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+test_checksum_state($node, 'off');
+
+$node->backup('full');
+
+# Offline enable: rewrites every page, logs nothing.
+$node->stop;
+system_or_bail('pg_checksums', '--enable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'on');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full/backup_manifest',
+ ],
+ 'incremental backup with checksums on');
+
+# Combining the chain must be refused: the result would say "on" while the
+# blocks inherited from the full backup have no checksums.
+command_fails_like(
+ [
+ 'pg_combinebackup',
+ $node->backup_dir . '/full',
+ $node->backup_dir . '/incr',
+ '--output' => $node->backup_dir . '/combined',
+ ],
+ qr/only some backups have checksums enabled/,
+ 'pg_combinebackup refuses a chain crossing an offline enable');
+
+# The other direction: full backup with checksums on, offline disable, then
+# an incremental. The last backup wins, so the combined cluster comes up off.
+$node->backup('full2');
+
+$node->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr2',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full2/backup_manifest',
+ ],
+ 'incremental backup with checksums off');
+
+my $restored = PostgreSQL::Test::Cluster->new('restored');
+$restored->init_from_backup(
+ $node, 'incr2',
+ combine_with_prior => ['full2'],
+ has_restoring => 1,
+ standby => 0);
+
+$restored->start;
+
+my ($rc, $stdout, $stderr) = $restored->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'combined backup accepts connections') or diag("stderr: $stderr");
+is($stdout, 'off', 'combined backup follows the last backup state');
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t3;');
+is($rc, 0, 'blocks from the incremental backup are readable')
+ or diag("stderr: $stderr");
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t1;');
+is($rc, 0, 'blocks inherited from the full backup are readable')
+ or diag("stderr: $stderr");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($restored->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures in the combined backup');
+
+$restored->stop('immediate');
+$node->stop;
+
+done_testing();
--
2.55.0
[application/octet-stream] v8-0005-doc-Explain-how-offline-and-online-checksum-chang.patch (4.2K, ../../CAN4CZFN5sOQvWTXr7Dm1J1HXwczikBfMjkbFQ7QmHsr+VHd5Mg@mail.gmail.com/4-v8-0005-doc-Explain-how-offline-and-online-checksum-chang.patch)
download | inline diff:
From 1427a826fbd45d6efd955a94a664430b56f524a0 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Mon, 31 Aug 2026 08:40:44 +0000
Subject: [PATCH v8 5/5] doc: Explain how offline and online checksum changes
interact
The lockstep procedure for offline checksum changes in a replication
setup was documented without preconditions. It is not sufficient on
its own: an offline change is recorded only in the control file and
has no ordering against WAL the node has not replayed yet, so a node
stopped before replaying an online state transition applies it on
restart and overrides the offline change. The nodes then silently
diverge even though the change was applied to all of them while
stopped.
Document the two durability models next to each other in wal.sgml,
and extend the pg_checksums notes: require standbys to have replayed
all WAL of their upstream node before stopping them, show how to
check that, and recommend not mixing online and offline changes.
---
doc/src/sgml/ref/pg_checksums.sgml | 25 ++++++++++++++++++++++---
doc/src/sgml/wal.sgml | 13 +++++++++++++
2 files changed, 35 insertions(+), 3 deletions(-)
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index 000940e7cb7..b337950be3d 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -248,9 +248,28 @@ PostgreSQL documentation
directory; the new state does not propagate over replication. In a
replication setup the same change must be applied to every node: stop all
nodes, run <application>pg_checksums</application> on each of them, and
- only then restart them. Tools that copy relation file blocks directly
- between nodes, such as <xref linkend="app-pgrewind"/>, likewise require
- both nodes to be in the same data checksum state.
+ only then restart them. Before stopping a standby, make sure it has
+ replayed all WAL of its upstream node, for example by stopping the
+ primary first and comparing
+ <function>pg_last_wal_replay_lsn()</function> with
+ <function>pg_last_wal_receive_lsn()</function> on the standby. Tools
+ that copy relation file blocks directly between nodes, such as
+ <xref linkend="app-pgrewind"/>, likewise require both nodes to be in
+ the same data checksum state.
+ </para>
+ <para>
+ The replay requirement exists because an offline change is recorded
+ only in the control file and has no defined ordering against WAL the
+ node has not replayed yet; see
+ <xref linkend="checksums-offline-enable-disable"/>. A node stopped
+ before replaying an online checksum state change applies that change
+ when it is restarted, overriding the offline change, and the states of
+ the nodes silently diverge until a later checkpoint record triggers
+ the warning described below. Because of this it is best not to mix
+ the two mechanisms: change the state of a replication setup either
+ with the offline procedure above or with an online transition, and
+ make sure the previous change has reached every node before starting
+ the next one.
</para>
<para>
If the change is applied inconsistently, each node keeps its own state,
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index db18a8c516e..d88a832c2a8 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -317,6 +317,19 @@
verify checksums, on an offline cluster.
</para>
+ <para>
+ An offline change provides durability differently from an
+ <link linkend="checksums-online-enable-disable">online change</link>.
+ An online transition is WAL-logged: it is ordered against all other
+ WAL records, it is replayed after a crash, and it propagates to
+ standbys. An offline change is recorded only in the cluster's
+ control file: it writes no WAL, it is invisible to replication, and
+ it has no defined ordering against WAL the node has not replayed
+ yet. When a node later replays WAL that contains an online checksum
+ state change, that change takes effect on the node even if it was
+ written before the offline change was made.
+ </para>
+
<para>
An offline change only affects the data directory it is run on; the
new state does not propagate over replication. In a replication setup
--
2.55.0
[application/octet-stream] v8-0001-Do-not-adopt-data-checksum-state-from-another-nod.patch (128.5K, ../../CAN4CZFN5sOQvWTXr7Dm1J1HXwczikBfMjkbFQ7QmHsr+VHd5Mg@mail.gmail.com/5-v8-0001-Do-not-adopt-data-checksum-state-from-another-nod.patch)
download | inline diff:
From 6f15c59638a09f7c32b9447c33f033b779e60a51 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:02 +0000
Subject: [PATCH v8 1/5] Do not adopt data checksum state from another node
during replay
Offline pg_checksums changes are local to one node, but replay adopted
the data checksum state carried by checkpoint records unconditionally.
After an offline enable on the primary, a standby whose pages were
never rewritten started verifying checksums it does not have; after an
offline change on a standby, the next replayed checkpoint silently
reverted it.
Instead, cross-check the replayed state against the local one and warn
once per divergent value. When the states match again, say so in the
log and re-arm the warning.
The control file now tracks this node's state alone, which makes the
moment it is written out part of the design: it may only claim "on"
once every page on disk carries a checksum, or a crash-restart would
verify pages that were flushed before the transition rewrote them.
XLOG2_CHECKSUMS replay therefore persists every state but "on" as soon
as it is replayed, none of them verifying anything, and leaves "on" to
the flush of the next restartpoint. A restartpoint only persists a
state its flush ran under from beginning to end, a promotion writes a
full checkpoint instead of an end-of-recovery record while the state is
still unpersisted, and a shutdown restartpoint with no new checkpoint
record to work from flushes before catching the control file up rather
than leaving it behind. Enabling checksums on the primary defers the
same write to after the checkpoint that flushes the rewritten pages:
the ones the worker found in shared buffers do not go out through its
ring buffer. A checkpoint persists the state its own redo point ran
under, under the same rule a restartpoint follows, since the checkpoint
that licenses "on" also moves the redo point past the record announcing
it: without that, a crash before the deferred write would resume above
the record and resolve a finished transition as interrupted. Checkpoint
records replayed below the consistency point are not cross-checked, the
persisted state being legitimately newer than what they carry.
For the redo point to be a reliable anchor, the states carried by WAL
records must match their WAL order. Transitions insert their record
and publish the new state under the new DataChecksumTransitionLock,
which a checkpoint holds across sampling the state and inserting its
XLOG_CHECKPOINT_REDO record; without that, a redo record could follow
a transition record in WAL while still carrying the pre-transition
state, and recovery resuming there would resolve the finished
transition as interrupted all the same.
pg_control also gains a watermark, the end LSN of the newest
XLOG2_CHECKSUMS record the node has written or applied, persisted
together with the state it produced, and replay skips records at or
below it. Without it, a standby that replayed a transition and stopped
cleanly before any restartpoint moved past the record would re-apply it
on the next startup, overriding a pg_checksums change made while it was
down; the offline change writes no WAL, so nothing would restore it. A
flag next to the watermark marks a state last written by pg_checksums
as local to this node, and recovery never adopts a checkpoint-borne
state over a local one, nor over a control file whose watermark already
covers the starting checkpoint. The watermark also stands in for the
state comparisons around the flushes above: record positions are
unique, so a full round trip back to the sampled state cannot alias.
A base backup copies the control file at an arbitrary moment, so its
state can be newer than the redo point replay starts from. Adopt the
state carried by the starting checkpoint record under backup-label
recovery; for a shutdown checkpoint, which replay does not see, take
it in StartupXLOG. Backups taken from a standby are the exception:
their starting checkpoint was written by the upstream primary, so they
keep the state of the control file that was copied with them.
pg_rewind keeps the target's own state, watermark and flag in the
control file it installs, since most of the data directory remains the
target's; replay from the last common checkpoint still adopts a state
its watermark does not cover, and applies any online transition the
target has not seen. The kept watermark is clamped to the divergence
point: the skip comparison has no timeline context, so a watermark set
by a transition record on the target's abandoned fork can numerically
cover records the source wrote after the divergence, and replay would
skip them as already applied, leaving the rewound node permanently out
of step with its new primary. Records at or below the divergence
point are common history and stay covered, still protecting an offline
change made after they were applied. No other path needs this: a node
keeps replaying the history its watermark came from unless pg_rewind
moves it, since a standby re-pointed at a promoted peer without a
rewind cannot have replayed past the fork, and pg_resetwal already
resets the watermark it may have orphaned.
Bump PG_CONTROL_VERSION.
The documented procedure for offline changes in a replication setup
becomes the lockstep one: stop all nodes, run pg_checksums on each of
them, then restart. Tests cover the lockstep procedure, divergent
offline changes on either node and down a cascading chain,
crash-restarts and promotions around online transitions, checkpoints
racing an online enable and the transition record itself, a base backup
taken during an online enable, an offline disable on a standby
surviving the re-replay of the enable that preceded it, and pg_rewind
across an online enable as well as across offline enables on both nodes
after a divergence, from a source whose online enable the target's own
abandoned transitions had outrun, and with checksums enabled online on
both nodes independently.
---
doc/src/sgml/ref/pg_checksums.sgml | 32 +-
doc/src/sgml/wal.sgml | 8 +
src/backend/access/transam/xlog.c | 549 ++++++++++++++++--
.../utils/activity/wait_event_names.txt | 1 +
src/bin/pg_checksums/pg_checksums.c | 10 +
src/bin/pg_controldata/pg_controldata.c | 4 +
src/bin/pg_resetwal/pg_resetwal.c | 7 +
src/bin/pg_rewind/pg_rewind.c | 35 +-
src/bin/pg_upgrade/controldata.c | 2 +-
src/include/catalog/pg_control.h | 26 +-
src/include/storage/lwlocklist.h | 1 +
src/test/modules/test_checksums/Makefile | 2 +-
src/test/modules/test_checksums/meson.build | 10 +
.../test_checksums/t/012_offline_standby.pl | 289 +++++++++
.../modules/test_checksums/t/013_rewind.pl | 201 +++++++
.../modules/test_checksums/t/014_lockstep.pl | 182 ++++++
.../t/015_standby_crash_after_disable.pl | 125 ++++
.../t/016_promote_enable_crash.pl | 134 +++++
.../test_checksums/t/017_restartpoint_race.pl | 138 +++++
.../t/018_enable_crash_windows.pl | 496 ++++++++++++++++
.../t/019_standby_shutdown_catchup.pl | 126 ++++
.../t/020_cascade_divergence.pl | 115 ++++
.../t/021_rewind_divergent_transitions.pl | 172 ++++++
23 files changed, 2586 insertions(+), 79 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/012_offline_standby.pl
create mode 100644 src/test/modules/test_checksums/t/013_rewind.pl
create mode 100644 src/test/modules/test_checksums/t/014_lockstep.pl
create mode 100644 src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
create mode 100644 src/test/modules/test_checksums/t/016_promote_enable_crash.pl
create mode 100644 src/test/modules/test_checksums/t/017_restartpoint_race.pl
create mode 100644 src/test/modules/test_checksums/t/018_enable_crash_windows.pl
create mode 100644 src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
create mode 100644 src/test/modules/test_checksums/t/020_cascade_divergence.pl
create mode 100644 src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index ae66fad3f0f..bf07da09754 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -242,15 +242,29 @@ PostgreSQL documentation
data directory must not be started or else data loss may occur.
</para>
<para>
- When using a replication setup with tools which perform direct copies
- of relation file blocks (for example <xref linkend="app-pgrewind"/>),
- enabling or disabling checksums can lead to page corruptions in the
- shape of incorrect checksums if the operation is not done consistently
- across all nodes. When enabling or disabling checksums in a replication
- setup, it is thus recommended to stop all the clusters before switching
- them all consistently. Destroying all standbys, performing the operation
- on the primary and finally recreating the standbys from scratch is also
- safe.
+ Enabling or disabling checksums with
+ <application>pg_checksums</application> changes only the local data
+ directory; the new state does not propagate over replication. In a
+ replication setup the same change must be applied to every node: stop all
+ nodes, run <application>pg_checksums</application> on each of them, and
+ only then restart them. Tools that copy relation file blocks directly
+ between nodes, such as <xref linkend="app-pgrewind"/>, likewise require
+ both nodes to be in the same data checksum state.
+ </para>
+ <para>
+ If the change is applied inconsistently, each node keeps its own state,
+ and a standby logs a warning when the state recorded in the replayed WAL
+ differs from its own. The same warning can appear transiently while a
+ standby catches up over WAL written before a consistent change; it stops
+ once a checkpoint record carrying the new state has been replayed.
+ </para>
+ <para>
+ A standby whose data directory was never checksummed must not have
+ checksums enabled by catching up this way. Converge the cluster by
+ running <application>pg_checksums</application> on it while stopped, by
+ enabling checksums online, or by recreating it from a base backup. Note
+ that an online enable only starts from a primary whose checksums are off,
+ so if they are already enabled there, disable them online first.
</para>
<para>
If <application>pg_checksums</application> is aborted or killed while
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index ec62d17fbcc..db18a8c516e 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -317,6 +317,14 @@
verify checksums, on an offline cluster.
</para>
+ <para>
+ An offline change only affects the data directory it is run on; the
+ new state does not propagate over replication. In a replication setup
+ the same change must be applied to all nodes while all of them are
+ stopped, as described in <xref linkend="app-pgchecksums"/>. A standby
+ whose state diverges logs a warning but keeps its local setting.
+ </para>
+
</sect2>
<sect2 id="checksums-online-enable-disable" xreflabel="Online Enabling of Checksums">
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index de4c96e135f..20428ad68cf 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -556,9 +556,17 @@ typedef struct XLogCtlData
*/
XLogRecPtr lastFpwDisableRecPtr;
- /* last data_checksum_version we've seen */
+ /* current data checksum state of this node */
uint32 data_checksum_version;
+ /*
+ * In-memory copies of ControlFile->data_checksum_lsn and
+ * ControlFile->data_checksum_is_local, see there. Updated together with
+ * data_checksum_version under info_lck.
+ */
+ XLogRecPtr data_checksum_lsn;
+ bool data_checksum_is_local;
+
slock_t info_lck; /* locks shared variables shown above */
/*
@@ -690,6 +698,14 @@ static ChecksumStateType LocalDataChecksumState = 0;
*/
int data_checksums = 0;
+/*
+ * Whether replay of the next checkpoint-family record must adopt the data
+ * checksum state it carries. Set when recovery starts from a base backup,
+ * where the state at the redo point takes precedence over the control file
+ * copied with the backup later.
+ */
+static bool adoptChecksumStateFromNextCheckpoint = false;
+
/* For WALInsertLockAcquire/Release functions */
static int MyLockNo = 0;
static bool holdingAllLocks = false;
@@ -730,6 +746,8 @@ static void ValidateXLOGDirectoryStructure(void);
static void CleanupBackupHistory(void);
static void UpdateMinRecoveryPoint(XLogRecPtr lsn, bool force);
static bool PerformRecoveryXLogAction(void);
+static void CheckReplayedDataChecksumState(uint32 replayed_version);
+static void AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn);
static void InitControlFile(uint64 sysidentifier, uint32 data_checksum_version);
static void WriteControlFile(void);
static void ReadControlFile(void);
@@ -757,7 +775,7 @@ static void WALInsertLockAcquireExclusive(void);
static void WALInsertLockRelease(void);
static void WALInsertLockUpdateInsertingAt(XLogRecPtr insertingAt);
-static void XLogChecksums(uint32 new_type);
+static XLogRecPtr XLogChecksums(uint32 new_type);
/*
* Insert an XLOG record represented by an already-constructed chain of data
@@ -4774,6 +4792,7 @@ void
SetDataChecksumsOnInProgress(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
/*
* The state transition is performed in a critical section with
@@ -4782,14 +4801,12 @@ SetDataChecksumsOnInProgress(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4827,6 +4844,8 @@ void
SetDataChecksumsOn(void)
{
uint64 barrier;
+ bool persist;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4847,23 +4866,11 @@ SetDataChecksumsOn(void)
SpinLockRelease(&XLogCtl->info_lck);
INJECTION_POINT("datachecksums-enable-checksums-delay", NULL);
+ INJECTION_POINT_LOAD("datachecksums-on-before-publish");
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_VERSION);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- SpinLockRelease(&XLogCtl->info_lck);
-
- /*
- * Update the controlfile before waiting since if we have an immediate
- * shutdown while waiting we want to come back up with checksums enabled.
- */
- LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- UpdateControlFile();
- LWLockRelease(ControlFileLock);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_VERSION);
barrier = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_CHECKSUM_ON);
@@ -4873,6 +4880,39 @@ SetDataChecksumsOn(void)
INJECTION_POINT("datachecksums-on-before-checkpoint", NULL);
RequestCheckpoint(CHECKPOINT_FORCE | CHECKPOINT_WAIT | CHECKPOINT_FAST);
+
+ INJECTION_POINT("datachecksums-on-after-checkpoint", NULL);
+
+ /*
+ * Persist "on" only now that the checkpoint has flushed the pages the
+ * transition rewrote. Crash recovery initializes verification from the
+ * control file but resumes from a checkpoint that can predate the
+ * transition, so an "on" persisted earlier would have replay verify pages
+ * whose rewrite never reached disk: the pages the worker found in shared
+ * buffers are not written back by its ring buffer.
+ *
+ * The checkpoint above normally persists the state itself; this covers the
+ * case where it started before the record written above and left the field
+ * alone. Crashing before this point is safe, as replay then re-establishes
+ * "on" from the full page images of the rewrite. Skip the write if the
+ * state moved on meanwhile, since whatever moved it persists its own.
+ * Compare the watermark rather than the state: a state comparison could
+ * not tell our transition from a later round trip back to "on".
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ persist = (XLogCtl->data_checksum_lsn == recptr);
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (persist)
+ {
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+ }
+
WaitForProcSignalBarrier(barrier);
}
@@ -4893,6 +4933,7 @@ void
SetDataChecksumsOff(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4918,14 +4959,12 @@ SetDataChecksumsOff(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4956,14 +4995,12 @@ SetDataChecksumsOff(void)
/* Ensure that we don't incur a checkpoint during disabling checksums */
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -5001,6 +5038,132 @@ SetLocalDataChecksumState(uint32 data_checksum_version)
data_checksums = data_checksum_version;
}
+/*
+ * Cross-check the data checksum state carried by a replayed checkpoint record
+ * against the state of this node.
+ *
+ * The state in the record belongs to whichever node wrote the WAL, and must
+ * not be adopted: an offline state change made with pg_checksums on one node
+ * of a replication set generates no WAL, and adopting would leak it into the
+ * other nodes through replay. Backup label recovery is the one exception,
+ * see AdoptReplayedDataChecksumState(). XLOG_CHECKPOINT_ONLINE needs no call
+ * here, as an online checkpoint's state already traveled in the preceding
+ * XLOG_CHECKPOINT_REDO record.
+ *
+ * Only archive recovery can see a lasting mismatch, as only there can the WAL
+ * and the control file come from different nodes or different times. In
+ * crash recovery a mismatch means replay resumed from a restartpoint
+ * predating an already-applied XLOG2_CHECKSUMS record, and replaying forward
+ * re-establishes the same state.
+ */
+static void
+CheckReplayedDataChecksumState(uint32 replayed_version)
+{
+ /*
+ * Warn once per remote value, so a lasting mismatch does not flood the
+ * log. Matching states re-arm the warning. Backend-local state is
+ * enough: replay only runs in the startup process, and restarting it
+ * re-arms as well.
+ */
+ static uint32 last_warned_version = PG_UINT32_MAX;
+ uint32 local_version;
+
+ if (!ArchiveRecoveryRequested)
+ return;
+
+ /*
+ * Re-replayed WAL below the consistency point was already cross-checked
+ * before minRecoveryPoint was last persisted, and the persisted state can
+ * legitimately be newer than what checkpoint records there carry:
+ * XLOG2_CHECKSUMS replay persists most states ahead of the restartpoint
+ * horizon. In particular the checkpoint record recovery restarts from is
+ * such a re-replay.
+ */
+ if (!reachedConsistency)
+ return;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ local_version = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (replayed_version == local_version)
+ {
+ /*
+ * Report convergence if this process warned before. Nothing else
+ * tells the operator that running pg_checksums on the other nodes, or
+ * a rebuild, took effect. A restart in between loses the context,
+ * but it re-arms the warning too.
+ */
+ if (last_warned_version != PG_UINT32_MAX)
+ ereport(LOG,
+ errmsg("data checksum state \"%s\" of this node now agrees with the replayed WAL",
+ get_checksum_state_string(local_version)));
+
+ last_warned_version = PG_UINT32_MAX;
+ return;
+ }
+
+ /* the nodes legitimately differ while an online transition runs */
+ if (replayed_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ replayed_version == PG_DATA_CHECKSUM_INPROGRESS_OFF ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ return;
+
+ if (replayed_version == last_warned_version)
+ return;
+ last_warned_version = replayed_version;
+
+ ereport(WARNING,
+ errmsg("data checksum state \"%s\" of this node does not match the state \"%s\" in the replayed WAL",
+ get_checksum_state_string(local_version),
+ get_checksum_state_string(replayed_version)),
+ errdetail("The data checksum state was most likely changed with pg_checksums on another node."),
+ errhint("Apply the same change with pg_checksums on the primary and all standby servers, or rebuild this server from a base backup."));
+}
+
+/*
+ * Adopt the data checksum state found at the redo point of backup label
+ * recovery. Persist it immediately so that a crash before the first
+ * restartpoint does not resurrect the state copied with the backup; a crash
+ * at this point restarts from the same redo point, so the control file does
+ * not run ahead of the replay position. If the value is unchanged the
+ * control file already carries it, so both the barrier and the persist are
+ * skipped.
+ *
+ * lsn is the location of the checkpoint-family record the state was taken
+ * from and becomes the new watermark: the adopted state covers everything
+ * below the redo point, which replay never revisits.
+ */
+static void
+AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn)
+{
+ bool changed = false;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (XLogCtl->data_checksum_version != new_version)
+ {
+ XLogCtl->data_checksum_version = new_version;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(new_version);
+ changed = true;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (!changed)
+ return;
+
+ EmitAndWaitDataChecksumsBarrier(new_version);
+
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = new_version;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+}
+
/* guc hook */
const char *
show_data_checksums(void)
@@ -5454,6 +5617,8 @@ XLOGShmemInit(void *arg)
/* Use the checksum info from control file */
XLogCtl->data_checksum_version = ControlFile->data_checksum_version;
+ XLogCtl->data_checksum_lsn = ControlFile->data_checksum_lsn;
+ XLogCtl->data_checksum_is_local = ControlFile->data_checksum_is_local;
SetLocalDataChecksumState(XLogCtl->data_checksum_version);
SpinLockInit(&XLogCtl->Insert.insertpos_lck);
@@ -6031,6 +6196,53 @@ StartupXLOG(void)
SetCommitTsLimit(checkPoint.oldestCommitTsXid,
checkPoint.newestCommitTsXid);
+ /*
+ * When recovery starts from a base backup, the control file was copied at
+ * an arbitrary moment and its data checksum state may differ from the
+ * state at the redo point, which is what the WAL from there on was
+ * written under. Adopt the state of the starting checkpoint: a shutdown
+ * checkpoint is not replayed, so take it from the record read above; the
+ * redo point of an online checkpoint is its CHECKPOINT_REDO record, so
+ * let the replay of that record adopt it. Check backupStartPoint in
+ * addition to the label: on a crash restart during backup recovery the
+ * label file is already renamed away, but the start point persists until
+ * the backup end record.
+ *
+ * Not for a base backup taken from a standby, though. Its starting
+ * checkpoint is the standby's last restartpoint, a record written by the
+ * upstream primary, whose state is not the one the copied files were
+ * written under. The copied control file is right there, as a standby
+ * persists its state only at restartpoint horizons and so never claims
+ * more than what reached disk. Such backups are recognized by
+ * backupEndPoint together with backupEndRequired; backupEndPoint is only
+ * set for "BACKUP FROM: standby" labels and persists across a crash
+ * restart. pg_rewind writes a standby label as well, but no
+ * backupEndPoint, and its recovery keeps adopting: the control file it
+ * installs carries the target's own checksum state, which can lag the
+ * redo point of the last common checkpoint the same way a restartpoint
+ * horizon can.
+ *
+ * Never adopt over a state the control file's watermark or local flag
+ * marks as newer than the starting checkpoint. A pg_checksums change is
+ * local to the node and generates no WAL, so nothing in the replayed WAL
+ * could ever restore it once overwritten; and a control file whose
+ * watermark lies above the redo point already contains the effect of
+ * every transition record up to there, including the state the starting
+ * checkpoint carries.
+ */
+ if ((haveBackupLabel || XLogRecPtrIsValid(ControlFile->backupStartPoint)) &&
+ !(XLogRecPtrIsValid(ControlFile->backupEndPoint) &&
+ ControlFile->backupEndRequired) &&
+ !ControlFile->data_checksum_is_local &&
+ checkPoint.redo > ControlFile->data_checksum_lsn)
+ {
+ if (wasShutdown)
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ checkPoint.redo);
+ else
+ adoptChecksumStateFromNextCheckpoint = true;
+ }
+
/*
* Clear out any old relcache cache files. This is *necessary* if we do
* any WAL replay, since that would probably result in the cache files
@@ -6633,11 +6845,7 @@ StartupXLOG(void)
if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
ereport(WARNING,
@@ -6654,11 +6862,7 @@ StartupXLOG(void)
else if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
}
@@ -6814,6 +7018,23 @@ static bool
PerformRecoveryXLogAction(void)
{
bool promoted = false;
+ bool flushForChecksums;
+ uint32 checksum_state;
+
+ /*
+ * The end-of-recovery record persists the data checksum state without
+ * flushing the buffer pool, but the control file may only claim "on" once
+ * every page on disk carries a checksum. If replay entered that state
+ * without a restartpoint following it, the pages rewritten by the
+ * transition are still only in the buffer pool, so take the full
+ * checkpoint below instead of the lightweight record.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ flushForChecksums = (checksum_state == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile->data_checksum_version != checksum_state);
/*
* Perform a checkpoint to update all our recovery activity to disk.
@@ -6829,7 +7050,7 @@ PerformRecoveryXLogAction(void)
* fully out of recovery mode and already accepting queries.
*/
if (ArchiveRecoveryRequested && IsUnderPostmaster &&
- PromoteIsTriggered())
+ PromoteIsTriggered() && !flushForChecksums)
{
promoted = true;
@@ -7436,6 +7657,7 @@ CreateCheckPoint(int flags)
uint32 freespace;
XLogRecPtr PriorRedoPtr;
XLogRecPtr last_important_lsn;
+ XLogRecPtr checksumLsn;
VirtualTransactionId *vxids;
int nvxids;
int oldXLogAllowed = 0;
@@ -7548,11 +7770,14 @@ CreateCheckPoint(int flags)
checkPoint.wal_level = wal_level;
/*
- * Get the current data_checksum_version value from xlogctl, valid at the
- * time of the checkpoint.
+ * Get the current data_checksum_version value from xlogctl. This is
+ * final only for a shutdown checkpoint, where no concurrent transition
+ * is possible; an online checkpoint resamples it together with the redo
+ * record below.
*/
SpinLockAcquire(&XLogCtl->info_lck);
checkPoint.dataChecksumState = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
if (shutdown)
@@ -7609,10 +7834,21 @@ CreateCheckPoint(int flags)
{
xl_checkpoint_redo redo_rec;
+ /*
+ * Sample the data checksum state and insert the redo record under
+ * DataChecksumTransitionLock, so that a concurrent transition cannot
+ * insert its XLOG2_CHECKSUMS record between the sampling and the
+ * insertion below. Without this, the redo record could follow the
+ * transition record in WAL while carrying the pre-transition state,
+ * and recovery resuming here would never learn about the transition.
+ * See XLogChecksums().
+ */
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
WALInsertLockAcquire();
redo_rec.wal_level = wal_level;
SpinLockAcquire(&XLogCtl->info_lck);
redo_rec.data_checksum_version = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
WALInsertLockRelease();
@@ -7620,6 +7856,14 @@ CreateCheckPoint(int flags)
XLogBeginInsert();
XLogRegisterData(&redo_rec, sizeof(xl_checkpoint_redo));
(void) XLogInsert(RM_XLOG_ID, XLOG_CHECKPOINT_REDO);
+ LWLockRelease(DataChecksumTransitionLock);
+
+ /*
+ * The checkpoint record must carry the same state as the redo record
+ * just inserted: the sample taken before redo determination can be
+ * stale by now, and the pair would otherwise disagree.
+ */
+ checkPoint.dataChecksumState = redo_rec.data_checksum_version;
/*
* XLogInsertRecord will have updated XLogCtl->Insert.RedoRecPtr in
@@ -7823,6 +8067,40 @@ CreateCheckPoint(int flags)
ControlFile->minRecoveryPoint = InvalidXLogRecPtr;
ControlFile->minRecoveryPointTLI = 0;
+ /*
+ * Persist the data checksum state this node runs under. Only the
+ * top-level field tracks this node; ControlFile->checkPointCopy above is
+ * a historical record used to resume replay.
+ *
+ * checkPoint.dataChecksumState was sampled under
+ * DataChecksumTransitionLock together with the redo record, so it is the
+ * state in effect at the redo point. If it was "on",
+ * the XLOG2_CHECKSUMS record announcing that precedes the redo point and
+ * every page the transition rewrote was dirtied before it, so
+ * CheckPointGuts() has just written all of them out. Recording the state
+ * here is what keeps a finished transition from being resolved as
+ * interrupted when this checkpoint is the one crash recovery resumes
+ * from: replay never sees the record announcing it.
+ *
+ * Persist it only if the state did not change while the flush was in
+ * progress. If it changed in between, the pages written out straddle two
+ * states, and the newer one could claim checksums that pages already on
+ * disk do not carry; leave the field to the next checkpoint then.
+ * SetDataChecksumsOff() persists the states that are safe to enter
+ * without a flush already. Compare the watermark rather than the state:
+ * a state comparison could not tell a full round trip back to the
+ * sampled value apart from no change at all, and the flushed pages
+ * straddle the intermediate states all the same.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksumLsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checkPoint.dataChecksumState;
+ ControlFile->data_checksum_lsn = checksumLsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
/*
* Persist unloggedLSN value. It's reset on crash recovery, so this goes
* unused on non-shutdown checkpoints, but seems useful to store it always
@@ -7967,9 +8245,11 @@ CreateEndOfRecoveryRecord(void)
ControlFile->minRecoveryPoint = recptr;
ControlFile->minRecoveryPointTLI = xlrec.ThisTimeLineID;
- /* start with the latest checksum version (as of the end of recovery) */
+ /* persist the data checksum state this node ended recovery with */
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
@@ -8174,6 +8454,9 @@ CreateRestartPoint(int flags)
XLogRecPtr endptr;
XLogSegNo _logSegNo;
TimestampTz xtime;
+ uint32 checksum_state;
+ XLogRecPtr checksum_lsn;
+ bool checksum_is_local;
/* Concurrent checkpoint/restartpoint cannot happen */
Assert(!IsUnderPostmaster || MyBackendType == B_CHECKPOINTER);
@@ -8220,8 +8503,45 @@ CreateRestartPoint(int flags)
UpdateMinRecoveryPoint(InvalidXLogRecPtr, true);
if (flags & CHECKPOINT_IS_SHUTDOWN)
{
+ bool catchUpChecksums;
+
+ /*
+ * There is no new restartpoint to persist the data checksum state
+ * with, but a cleanly stopped node should not leave the control
+ * file behind the state replay reached: pg_checksums and
+ * pg_rewind read it, and an in-progress state there makes them
+ * refuse to run. Catching it up needs the same guarantee a
+ * restartpoint gives: that every page on disk carries a checksum,
+ * so flush the buffer pool first. Only a transition to
+ * "on" that no restartpoint followed can get here; the other
+ * states are already persisted by XLOG2_CHECKSUMS replay. Replay
+ * has ended by now, so the state cannot change under us.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ checksum_is_local = XLogCtl->data_checksum_is_local;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ catchUpChecksums =
+ (checksum_lsn != ControlFile->data_checksum_lsn &&
+ XLogRecPtrIsValid(lastCheckPointRecPtr));
+
+ if (catchUpChecksums)
+ {
+ MemSet(&CheckpointStats, 0, sizeof(CheckpointStats));
+ CheckpointStats.ckpt_start_t = GetCurrentTimestamp();
+ CheckPointGuts(lastCheckPoint.redo, flags);
+ }
+
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
+ if (catchUpChecksums)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = checksum_is_local;
+ }
UpdateControlFile();
LWLockRelease(ControlFileLock);
}
@@ -8262,6 +8582,17 @@ CreateRestartPoint(int flags)
/* Update the process title */
update_checkpoint_display(flags, true, false);
+ /*
+ * Note the data checksum state the flush below starts under. Replay runs
+ * concurrently and can change the state while the flush is in progress,
+ * in which case the flush covers pages written under both states; see
+ * where the state is persisted further down.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
CheckPointGuts(lastCheckPoint.redo, flags);
/*
@@ -8321,8 +8652,26 @@ CreateRestartPoint(int flags)
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
}
- /* we shall start with the latest checksum version */
- ControlFile->data_checksum_version = lastCheckPoint.dataChecksumState;
+ /*
+ * Persist the data checksum state of this node. Not the state of the
+ * replayed checkpoint: that one belongs to the node that wrote it and
+ * may differ after an offline change on either side.
+ * ControlFile->checkPointCopy above keeps the replayed value on
+ * purpose, being a historical record used to resume replay rather
+ * than a tracker of node state.
+ *
+ * Persist only if the flush above ran under one state throughout; see
+ * CreateCheckPoint() for why, including why this compares the
+ * watermark and not the state.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksum_lsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
}
@@ -8763,9 +9112,21 @@ XLogReportParameters(void)
}
/*
- * Log the new state of checksums
+ * Log and publish the new state of checksums
+ *
+ * Inserting the record and publishing the new state must be atomic with
+ * respect to a checkpoint sampling the state for its XLOG_CHECKPOINT_REDO
+ * record: without that, a checkpoint could read the old state after the
+ * record is already in WAL and insert a redo record that both precedes the
+ * transition in WAL order and carries the pre-transition state. Recovery
+ * resuming from such a redo point would never replay the transition record
+ * and resolve the finished transition as interrupted.
+ * DataChecksumTransitionLock serializes the two; see CreateCheckPoint().
+ *
+ * Returns the end LSN of the inserted record, which the caller persists
+ * together with the new state as the data checksum watermark.
*/
-static void
+static XLogRecPtr
XLogChecksums(uint32 new_type)
{
xl_checksum_state xlrec;
@@ -8773,12 +9134,28 @@ XLogChecksums(uint32 new_type)
xlrec.new_checksum_state = new_type;
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
+
XLogBeginInsert();
XLogRegisterData((char *) &xlrec, sizeof(xl_checksum_state));
recptr = XLogInsert(RM_XLOG2_ID, XLOG2_CHECKSUMS);
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, recptr);
+
+ /* only loaded by SetDataChecksumsOn(), a no-op for the other callers */
+ INJECTION_POINT_CACHED("datachecksums-on-before-publish", NULL);
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ XLogCtl->data_checksum_version = new_type;
+ XLogCtl->data_checksum_lsn = recptr;
+ XLogCtl->data_checksum_is_local = false;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockRelease(DataChecksumTransitionLock);
+
XLogFlush(recptr);
+
+ return recptr;
}
/*
@@ -8966,11 +9343,19 @@ xlog_redo(XLogReaderState *record)
/* ControlFile->checkPointCopy always tracks the latest ckpt XID */
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->checkPointCopy.nextXid = checkPoint.nextXid;
- ControlFile->data_checksum_version = checkPoint.dataChecksumState;
UpdateControlFile();
LWLockRelease(ControlFileLock);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(checkPoint.dataChecksumState);
+
/*
* We should've already switched to the new TLI before replaying this
* record.
@@ -9206,19 +9591,17 @@ xlog_redo(XLogReaderState *record)
else if (info == XLOG_CHECKPOINT_REDO)
{
xl_checkpoint_redo redo_rec;
- bool new_state = false;
memcpy(&redo_rec, XLogRecGetData(record), sizeof(xl_checkpoint_redo));
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = redo_rec.data_checksum_version;
- SetLocalDataChecksumState(redo_rec.data_checksum_version);
- if (redo_rec.data_checksum_version != ControlFile->data_checksum_version)
- new_state = true;
- SpinLockRelease(&XLogCtl->info_lck);
-
- if (new_state)
- EmitAndWaitDataChecksumsBarrier(redo_rec.data_checksum_version);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(redo_rec.data_checksum_version,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(redo_rec.data_checksum_version);
}
else if (info == XLOG_LOGICAL_DECODING_STATUS_CHANGE)
{
@@ -9280,25 +9663,63 @@ xlog2_redo(XLogReaderState *record)
{
xl_checksum_state state;
XLogRecPtr lsn = record->EndRecPtr;
+ XLogRecPtr watermark;
memcpy(&state, XLogRecGetData(record), sizeof(xl_checksum_state));
+ SpinLockAcquire(&XLogCtl->info_lck);
+ watermark = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ /*
+ * Skip records this node has already applied. The control file
+ * carries the watermark, so this holds across restarts: recovery
+ * resuming below a record whose effect the control file already
+ * contains must not re-apply it, or it would revert a state change
+ * made with pg_checksums in between, which moves the state without
+ * writing any record of its own.
+ */
+ if (lsn <= watermark)
+ return;
+
/* advertise the location before the new state becomes visible */
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, lsn);
SpinLockAcquire(&XLogCtl->info_lck);
XLogCtl->data_checksum_version = state.new_checksum_state;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(state.new_checksum_state);
SpinLockRelease(&XLogCtl->info_lck);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = state.new_checksum_state;
+
+ /*
+ * Persist the new state, except when it is "on". Only "on" verifies
+ * checksums during reads, and between the last restartpoint and this
+ * record there may be pages on disk flushed under the old state; a
+ * crash-restart initializes verification from the control file and
+ * replay reads those pages back, so the control file may only say
+ * "on" once everything written under the transition has been flushed,
+ * as restartpoints and the end of recovery do.
+ * The opposite direction cannot wait for the restartpoint: once this
+ * record is replayed, evicted pages are written without checksums,
+ * and a control file still saying "on" would fail verification on
+ * exactly those pages after a crash.
+ */
+ if (state.new_checksum_state != PG_DATA_CHECKSUM_VERSION)
+ {
+ ControlFile->data_checksum_version = state.new_checksum_state;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ }
/*
* Update minRecoveryPoint to ensure that if recovery is aborted, we
* recover back up to this point before allowing hot standby again.
- * The new state is durable in pg_control while its location is only
- * tracked in shared memory; a standby becoming consistent below this
- * record would let base backups resume checksum verification with the
+ * The change location is only tracked in shared memory and is lost
+ * over a restart; a standby becoming consistent below this record
+ * would let base backups resume checksum verification with the
* location unknown. The local copies cannot be updated as long as
* crash recovery is happening and we expect all the WAL to be
* replayed.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 256b3a3c02e..77fa543f9d8 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -371,6 +371,7 @@ WaitLSN "Waiting to read or update shared Wait-for-LSN state."
LogicalDecodingControl "Waiting to read or update logical decoding status information."
DataChecksumsWorker "Waiting for data checksums worker."
AioWorkerControl "Waiting to update AIO worker information."
+DataChecksumTransition "Waiting for a data checksum state transition to be written to WAL."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 3b3ae23f1a6..20a99112838 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -648,6 +648,16 @@ main(int argc, char *argv[])
ControlFile->data_checksum_version =
(mode == PG_MODE_ENABLE) ? PG_DATA_CHECKSUM_VERSION : PG_DATA_CHECKSUM_OFF;
+ /*
+ * Mark the state as changed locally, without a WAL record. Recovery
+ * then knows the state is newer than anything the WAL carries and
+ * does not let a replayed checkpoint overwrite it. The watermark is
+ * left alone: any XLOG2_CHECKSUMS record this node had applied stays
+ * covered, and only records above it, written after this change, take
+ * effect again.
+ */
+ ControlFile->data_checksum_is_local = true;
+
if (do_sync)
{
pg_log_info("syncing data directory");
diff --git a/src/bin/pg_controldata/pg_controldata.c b/src/bin/pg_controldata/pg_controldata.c
index 6fc87ed114d..c363ce3dbb7 100644
--- a/src/bin/pg_controldata/pg_controldata.c
+++ b/src/bin/pg_controldata/pg_controldata.c
@@ -349,6 +349,10 @@ main(int argc, char *argv[])
(ControlFile->float8ByVal ? _("by value") : _("by reference")));
printf(_("Data page checksum version: %u\n"),
ControlFile->data_checksum_version);
+ printf(_("Data checksum watermark: %X/%08X\n"),
+ LSN_FORMAT_ARGS(ControlFile->data_checksum_lsn));
+ printf(_("Data checksum state is node-local: %s\n"),
+ (ControlFile->data_checksum_is_local ? _("yes") : _("no")));
printf(_("Default char data signedness: %s\n"),
(ControlFile->default_char_signedness ? _("signed") : _("unsigned")));
printf(_("Mock authentication nonce: %s\n"),
diff --git a/src/bin/pg_resetwal/pg_resetwal.c b/src/bin/pg_resetwal/pg_resetwal.c
index 1542a56ca4b..cdfb9898066 100644
--- a/src/bin/pg_resetwal/pg_resetwal.c
+++ b/src/bin/pg_resetwal/pg_resetwal.c
@@ -923,6 +923,13 @@ RewriteControlFile(void)
ControlFile.backupEndPoint = InvalidXLogRecPtr;
ControlFile.backupEndRequired = false;
+ /*
+ * The old WAL is gone and the new position may lie below the old
+ * watermark, which would make replay ignore future checksum transition
+ * records. The state itself is kept.
+ */
+ ControlFile.data_checksum_lsn = InvalidXLogRecPtr;
+
/*
* Force the defaults for max_* settings. The values don't really matter
* as long as wal_level='minimal'; the postmaster will reset these fields
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 2e86fd158d0..466f4223501 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -37,7 +37,8 @@ static void usage(const char *progname);
static void perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo);
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec);
static void createBackupLabel(XLogRecPtr startpoint, TimeLineID starttli,
XLogRecPtr checkpointloc);
@@ -531,7 +532,8 @@ main(int argc, char **argv)
* This is the point of no return. Once we start copying things, there is
* no turning back!
*/
- perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo);
+ perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo,
+ divergerec);
if (showprogress)
pg_log_info("syncing target data directory");
@@ -566,7 +568,8 @@ static void
perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo)
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec)
{
XLogRecPtr endrec;
TimeLineID endtli;
@@ -738,6 +741,32 @@ perform_rewind(filemap_t *filemap, rewind_source *source,
ControlFile_new.minRecoveryPoint = endrec;
ControlFile_new.minRecoveryPointTLI = endtli;
ControlFile_new.state = DB_IN_ARCHIVE_RECOVERY;
+
+ /*
+ * Keep the target's own data checksum state. Most of the data directory
+ * is still the target's: only blocks it changed since the divergence
+ * were copied from the source, so the source's state says nothing about
+ * the pages that stay. Replay from the last common checkpoint applies
+ * any WAL-logged transition the target has not seen (the watermark tells
+ * them apart), which converges the rewound server to the source's state
+ * whenever the WAL carries it.
+ */
+ ControlFile_new.data_checksum_version = ControlFile_target.data_checksum_version;
+ ControlFile_new.data_checksum_lsn = ControlFile_target.data_checksum_lsn;
+ ControlFile_new.data_checksum_is_local = ControlFile_target.data_checksum_is_local;
+
+ /*
+ * The watermark is only meaningful within the history the node replays.
+ * Records at or below the divergence point are common to both histories
+ * and stay covered, but a watermark above it was set by a transition
+ * record on the target's own abandoned fork: numerically it can cover
+ * transition records the source wrote after the divergence, and replay
+ * would skip them as already applied. Clamp it to the divergence point,
+ * so that every transition record on the source's history takes effect.
+ */
+ if (ControlFile_new.data_checksum_lsn > divergerec)
+ ControlFile_new.data_checksum_lsn = divergerec;
+
if (!dry_run)
update_controlfile(datadir_target, &ControlFile_new, do_sync);
}
diff --git a/src/bin/pg_upgrade/controldata.c b/src/bin/pg_upgrade/controldata.c
index b3bd4ccde83..de539479bdc 100644
--- a/src/bin/pg_upgrade/controldata.c
+++ b/src/bin/pg_upgrade/controldata.c
@@ -431,7 +431,7 @@ get_control_data(ClusterInfo *cluster)
cluster->controldata.date_is_int = strstr(p, "64-bit integers") != NULL;
got_date_is_int = true;
}
- else if ((p = strstr(bufin, "checksum")) != NULL)
+ else if ((p = strstr(bufin, "Data page checksum version:")) != NULL)
{
p = strchr(p, ':');
diff --git a/src/include/catalog/pg_control.h b/src/include/catalog/pg_control.h
index 7b5404460ec..86395d5191e 100644
--- a/src/include/catalog/pg_control.h
+++ b/src/include/catalog/pg_control.h
@@ -22,7 +22,7 @@
/* Version identifier for this pg_control format */
-#define PG_CONTROL_VERSION 1903
+#define PG_CONTROL_VERSION 1904
/* Nonce key length, see below */
#define MOCK_AUTH_NONCE_LEN 32
@@ -237,6 +237,30 @@ typedef struct ControlFileData
/* Current data checksums state */
uint32 data_checksum_version;
+ /*
+ * End of the newest XLOG2_CHECKSUMS record this node has written or
+ * applied. Replay ignores XLOG2_CHECKSUMS records at or below this
+ * point: their effect is already contained in data_checksum_version, or
+ * an offline pg_checksums change made after they were first applied
+ * supersedes them. InvalidXLogRecPtr if the node has never written or
+ * applied such a record.
+ *
+ * The comparison has no timeline context, so the value is only valid
+ * within the WAL history this node replays. A tool that moves the node
+ * to another history must clamp the watermark to the point where the
+ * histories fork, as pg_rewind does, or reset it, as pg_resetwal does.
+ */
+ XLogRecPtr data_checksum_lsn;
+
+ /*
+ * True when data_checksum_version was last set by pg_checksums rather
+ * than by a WAL-logged transition. Such a state is local to this node
+ * and newer than anything the WAL carries, so recovery must not replace
+ * it with a state taken from a checkpoint record. Cleared by the next
+ * WAL-logged transition.
+ */
+ bool data_checksum_is_local;
+
/*
* True if the default signedness of char is "signed" on a platform where
* the cluster is initialized.
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index d7eb648bd27..8d858be9927 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -89,6 +89,7 @@ PG_LWLOCK(54, WaitLSN)
PG_LWLOCK(55, LogicalDecodingControl)
PG_LWLOCK(56, DataChecksumsWorker)
PG_LWLOCK(57, AioWorkerControl)
+PG_LWLOCK(58, DataChecksumTransition)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/test/modules/test_checksums/Makefile b/src/test/modules/test_checksums/Makefile
index 71455cd5577..80f54bf6d8c 100644
--- a/src/test/modules/test_checksums/Makefile
+++ b/src/test/modules/test_checksums/Makefile
@@ -9,7 +9,7 @@
#
#-------------------------------------------------------------------------
-EXTRA_INSTALL = src/test/modules/injection_points
+EXTRA_INSTALL = contrib/pg_buffercache src/test/modules/injection_points
export enable_injection_points
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index fb7129d796f..e5d38fafb7d 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -35,6 +35,16 @@ tests += {
't/009_fpi.pl',
't/010_backup_straddle.pl',
't/011_standby_straddle.pl',
+ 't/012_offline_standby.pl',
+ 't/013_rewind.pl',
+ 't/014_lockstep.pl',
+ 't/015_standby_crash_after_disable.pl',
+ 't/016_promote_enable_crash.pl',
+ 't/017_restartpoint_race.pl',
+ 't/018_enable_crash_windows.pl',
+ 't/019_standby_shutdown_catchup.pl',
+ 't/020_cascade_divergence.pl',
+ 't/021_rewind_divergent_transitions.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
new file mode 100644
index 00000000000..241a1282ae9
--- /dev/null
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -0,0 +1,289 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Offline checksum changes with pg_checksums are local to one node. A
+# standby must neither adopt the state of the primary from replayed
+# checkpoint records, nor lose its own offline change to them.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', 'autovacuum = off');
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Scenario 1: enable offline on the primary only. The standby must
+# stay off, warn about the mismatch, and remain readable.
+$standby->stop;
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+test_checksum_state($primary, 'on');
+test_checksum_state($standby, 'off');
+
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after offline enable on the primary');
+
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+my @warnings = $log =~ /(does not match the state)/g;
+is(scalar(@warnings), 1, 'mismatch warned once per remote value');
+
+# Matching states re-arm the warning: undo the divergence on the primary,
+# then diverge again to the same value, all without restarting the standby.
+$primary->stop;
+$primary->checksum_disable_offline;
+$logstart = -s $standby->logfile;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no warning while the states match again');
+like($log, qr/now agrees with the replayed WAL/,
+ 'convergence is reported once the states match again');
+
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+test_checksum_state($standby, 'off');
+
+# The local state survives both clean and immediate restarts.
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+# Still in scenario 1's divergence: take a base backup *from the
+# standby*. A backup taken on a standby uses the last restartpoint as
+# its starting checkpoint (do_pg_backup_start()), so the record at the
+# redo point was written by the upstream primary and carries the
+# primary's state. Recovery from such a backup must not adopt that
+# state: the files were copied from the standby, and the new node
+# would come up verifying checksums its files do not have.
+
+# Make sure the standby has written pages under its own "off" state, so
+# its files really do lack checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,50000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', "SELECT count(*) FROM t2;");
+
+$standby->backup('from_standby');
+my $newnode = PostgreSQL::Test::Cluster->new('newnode');
+$newnode->init_from_backup($standby, 'from_standby');
+
+# Stream from the primary, not from the standby the files were copied
+# from, so the new node replays the "on" primary's records.
+$newnode->enable_streaming($primary);
+$newnode->start;
+
+# The new node's files all came from a cluster running with checksums
+# off. Anything but "off" here means it adopted the primary's state
+# through the checkpoint record at the redo point.
+my ($rc, $stdout, $stderr) = $newnode->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'the node copied from the standby accepts connections')
+ or diag("stderr: $stderr");
+is($stdout, 'off', 'backup of an "off" standby comes up with checksums off');
+
+# A checkpoint record streamed from the "on" primary must not flip the
+# state either.
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($newnode);
+test_checksum_state($newnode, 'off');
+
+# And it must be able to read the pages the standby wrote without
+# checksums.
+(undef, $stdout, $stderr) =
+ $newnode->psql('postgres', "SELECT count(*) FROM t2;");
+is($stdout, '50000', 'pages copied from the standby are readable')
+ or diag("stderr: $stderr");
+
+my $newnode_log = PostgreSQL::Test::Utils::slurp_file($newnode->logfile);
+unlike(
+ $newnode_log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the node copied from the standby');
+
+$newnode->stop('immediate');
+
+# Converge the cluster: enable offline on the standby too.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+$primary->wait_for_catchup($standby);
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after converging');
+
+# Scenario 2: disable offline on the standby only. The replayed
+# checkpoint records of the still-enabled primary must not override it.
+$standby->stop;
+$standby->checksum_disable_offline;
+$standby->start;
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'off');
+
+# Restartpoints must persist the local state, not the replayed copy.
+$standby->safe_psql('postgres', "CHECKPOINT;");
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable with checksums disabled locally');
+
+# Scenario 3: crash-restart right after an online transition, before the
+# next restartpoint. Replay then resumes from an older restartpoint whose
+# checkpoint records still carry the pre-transition state. Those must
+# still match the state seeded from the control file, and the transition
+# itself must be re-established by re-replaying the XLOG2_CHECKSUMS record,
+# without a spurious mismatch warning along the way.
+
+# Converge first: bring the standby back to "on" offline.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+disable_data_checksums($primary, wait => 'off');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+# Crash-restart the standby immediately, before any restartpoint has had a
+# chance to persist the new state to its control file.
+$logstart = -s $standby->logfile;
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001',
+ 'standby readable after crash-restart across an online transition');
+
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no spurious mismatch warning after crash-restart across an online transition'
+);
+
+# Scenario 4: a standby stopped while replaying an interrupted online
+# transition keeps the interrupted state in its own control file. A
+# primary is never caught this way, as its checksums launcher resolves
+# inprogress-on back to off from its exit cleanup. A standby has no
+# launcher; it carries forward whatever the last replayed record left
+# it in.
+
+# Block an online enable on the primary at inprogress-on with a
+# blocking temp table, same trick as in 004_offline.pl.
+my $bsession = $primary->background_psql('postgres');
+$bsession->query_safe('CREATE TEMPORARY TABLE tt (a integer);');
+enable_data_checksums($primary, wait => 'inprogress-on');
+
+# The standby picks up the in-progress state from the XLOG2_CHECKSUMS record.
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# Stop the standby cleanly; its restartpoint persists inprogress-on to
+# its own control file, since nothing on a standby resolves it away.
+$standby->stop;
+$standby->start;
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# The primary still sits at inprogress-on, so a base backup taken now
+# copies a control file and a redo point that both carry the
+# in-progress state; replay of the XLOG2_CHECKSUMS records completes
+# the transition on the new standby.
+$primary->backup('inprogress_backup');
+my $standby2 = PostgreSQL::Test::Cluster->new('standby2');
+$standby2->init_from_backup($primary, 'inprogress_backup',
+ has_streaming => 1);
+$standby2->start;
+
+# Backup label recovery adopts the redo point's state immediately,
+# before any further WAL is replayed.
+wait_for_checksum_state($standby2, 'inprogress-on');
+
+$bsession->quit;
+wait_for_checksum_state($primary, 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable once the transition completes');
+
+# Scenario 4 continued: the backup taken mid-transition must also
+# complete the transition and stay readable.
+$primary->wait_for_catchup($standby2);
+wait_for_checksum_state($standby2, 'on');
+
+is($standby2->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'mid-transition backup standby readable after the transition');
+
+$standby2->stop;
+
+# Nothing along this standby's replay can legitimately disagree: it
+# started out at the same inprogress-on state as the primary.
+my $standby2_log = PostgreSQL::Test::Utils::slurp_file($standby2->logfile);
+unlike(
+ $standby2_log,
+ qr/does not match the state/,
+ 'no mismatch warning on a backup taken mid-transition');
+
+# No extra CHECKPOINT needed: the transition's completion checkpoint
+# wrote every page with a checksum, and the shutdown restartpoint
+# flushes the rest.
+command_ok([ 'pg_checksums', '--check', '-D', $standby2->data_dir ],
+ 'checksums valid on the mid-transition backup standby');
+
+$standby->stop;
+$primary->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/013_rewind.pl b/src/test/modules/test_checksums/t/013_rewind.pl
new file mode 100644
index 00000000000..a791e24317d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/013_rewind.pl
@@ -0,0 +1,201 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test pg_rewind across an online data checksum enable.
+#
+# A clean switchover leaves the shutdown checkpoint of the old primary
+# as the last common checkpoint between the two nodes. When data
+# checksums are enabled online on the new primary before the old one is
+# rewound, pg_rewind installs the control file of the new primary, which
+# already claims checksums are fully enabled, while replay begins at the
+# shutdown checkpoint whose record still carries the old state. Replay
+# of the WAL stretch from before the enable must run with checksums off,
+# as recorded in the checkpoint, else it would verify pages which never
+# had checksums written and fail recovery.
+#
+# The new primary requests a checkpoint right after promotion, and
+# replaying its CHECKPOINT_REDO record would repair the state before
+# any interesting WAL is reached. Hold the checkpointer on the standby
+# in a restartpoint over the promotion, like in the recovery test
+# 041_checkpoint_at_promote, so that the post-promotion writes end up
+# in WAL before the first checkpoint of the new timeline.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Old primary. full_page_writes is off so that the updates done on the
+# promoted node do not carry full page images, forcing replay on the
+# rewound node to read the pages from disk. wal_log_hints is required
+# by pg_rewind on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+full_page_writes = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+log_checkpoints = on
+]);
+$node_a->start;
+
+if (!$node_a->check_extension('injection_points'))
+{
+ plan skip_all => 'Extension injection_points not installed';
+}
+
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+$node_a->safe_psql('postgres', "CREATE EXTENSION injection_points;");
+
+# Set the hint bits on t before taking the backup, so that reads on the
+# promoted node do not emit full page images for its pages later.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->safe_psql('postgres', "SELECT count(*) FROM t;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_a, 'off');
+test_checksum_state($node_b, 'off');
+
+# Hold the next restartpoint on the standby.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point', 'wait');");
+
+# Give the restartpoint a checkpoint record to work on, then start it
+# in a background session; it will block on the injection point with
+# the checkpointer busy until released.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+my $bg_psql = $node_b->background_psql('postgres', on_error_stop => 0);
+$bg_psql->query_until(
+ qr/starting_restartpoint/, q(
+ \echo starting_restartpoint
+ CHECKPOINT;
+));
+$node_b->wait_for_event('checkpointer', 'create-restart-point');
+
+# Clean switchover: the shutdown checkpoint of A streams to B and
+# becomes the last common checkpoint.
+$node_a->stop('fast');
+
+my ($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+$stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+my $shutdown_ckpt = $1;
+$node_b->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_ckpt'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+
+my $logstart = -s $node_b->logfile;
+$node_b->promote;
+
+# Accidental restart of the old primary, diverging its timeline.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->stop('fast');
+
+# Updates on the new primary before its first checkpoint; replay of
+# these on the rewound node has to read the pages from disk.
+$node_b->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 25 = 0;");
+ok( !$node_b->log_contains("checkpoint complete", $logstart),
+ "no checkpoint on the new timeline before the updates");
+
+# Release the checkpointer; the queued post-promotion checkpoint runs
+# after the updates.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+$bg_psql->quit;
+
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+# pg_rewind refuses to run with full_page_writes disabled on the
+# source; the updates it was disabled for are already in WAL.
+$node_b->safe_psql('postgres', "ALTER SYSTEM SET full_page_writes = on;");
+$node_b->safe_psql('postgres', "SELECT pg_reload_conf();");
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the new primary');
+
+# Replay on the rewound node must start at the shutdown checkpoint of
+# the switchover, with the control file of the new primary.
+my $backup_label = slurp_file($node_a->data_dir . '/backup_label');
+$backup_label =~ /^CHECKPOINT LOCATION: ([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from backup_label";
+is($1, $shutdown_ckpt, 'replay starts at the switchover checkpoint');
+
+($stdout, $stderr) = run_command(
+ [
+ 'pg_waldump',
+ '-p' => $node_a->data_dir . '/pg_wal',
+ '-t' => 1,
+ '-s' => $shutdown_ckpt,
+ '-n' => 1,
+ ]);
+like($stdout, qr/CHECKPOINT_SHUTDOWN/,
+ 'last common checkpoint is a shutdown checkpoint');
+
+# pg_rewind keeps the target's own checksum state in the control file it
+# writes; the online enable reaches the rewound node through WAL replay
+# below, not through the copied control file.
+($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+like(
+ $stdout,
+ qr/^Data page checksum version:\s*0$/m,
+ 'rewound node keeps its own checksum state in the control file');
+
+# Start the rewound node as a standby of the new primary. Replay runs
+# through the pre-enable WAL stretch and the online enable.
+#
+# The rewind replaced the configuration files with those of the new
+# primary, so put the port back.
+my $connstr = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_b->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/014_lockstep.pl b/src/test/modules/test_checksums/t/014_lockstep.pl
new file mode 100644
index 00000000000..713f37fb22c
--- /dev/null
+++ b/src/test/modules/test_checksums/t/014_lockstep.pl
@@ -0,0 +1,182 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# The lockstep procedure for offline checksum changes in a replication
+# setup: stop all nodes, run pg_checksums on all of them, restart.
+# Replay of checkpoint records written before the change must not
+# revert the state of the standby.
+#
+# The same pair then runs the procedure in the disable direction, and
+# finally across a divergence: after a failover both nodes get their
+# checksums enabled offline, while the last common checkpoint still
+# carries "off". pg_rewind must keep the offline "on" on the rewound
+# node, since no record in the replayed WAL could ever restore it.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# wal_log_hints keeps the pair eligible for pg_rewind without checksums.
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Part 1: the lockstep procedure in the enable direction. Stop the
+# standby first: the WAL written after this point is replayed only
+# after the offline switch, and every checkpoint record in it still
+# carries the old state.
+$standby->stop;
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 0;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 1;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+# The lockstep procedure.
+$primary->checksum_enable_offline;
+$standby->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to off.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "off" in the replayed WAL/,
+ 0);
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep enable');
+
+# Crash the standby and replay the same stretch again.
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'on');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after crash restart');
+
+# Once a post-switch checkpoint has been replayed the states match and
+# no warning may be logged.
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no mismatch warning once the states match');
+
+# Every page the standby wrote in this window must carry a checksum.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$standby->stop;
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'checksums valid on the standby');
+
+# Part 2: the lockstep procedure in the other direction. The standby
+# is already stopped; give it a pre-switch WAL stretch to replay, with
+# every checkpoint record in it still carrying "on" -- an explicit
+# checkpoint plus the shutdown checkpoint written by stop().
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 2;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+$primary->checksum_disable_offline;
+$standby->checksum_disable_offline;
+
+my $disable_logstart = -s $standby->logfile;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to on.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $disable_logstart);
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep disable');
+
+# Part 3: failover, divergence and pg_rewind across offline enables.
+# The roles swap from here on: the standby becomes the new primary and
+# the old primary is rewound to follow it, but the variables keep the
+# names they had above.
+#
+# Failover: promote the standby, then diverge the old primary.
+$standby->promote;
+$standby->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$primary->stop;
+
+# The lockstep procedure across the divergence: both nodes get their
+# checksums enabled offline.
+$primary->checksum_enable_offline;
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+
+# The states match, so the rewind proceeds without complaint.
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $primary->data_dir,
+ '--source-server' => $standby->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled offline on both nodes');
+
+# The rewind replaced the configuration files with those of the source,
+# so put the port back. The copied file also carries the source's own
+# stale primary_conninfo; enable_streaming() below appends ours, which
+# wins as the later entry.
+$primary->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$primary->port]}
+]);
+$primary->enable_streaming($standby);
+
+my $rewind_logstart = -s $primary->logfile;
+$primary->start;
+$standby->wait_for_catchup($primary);
+
+is($primary->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+
+# The offline enable survives: replay from the common checkpoint must
+# not resurrect the pre-divergence "off".
+test_checksum_state($primary, 'on');
+
+$log =
+ PostgreSQL::Test::Utils::slurp_file($primary->logfile, $rewind_logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no divergence reported between the rewound server and its source');
+
+$primary->stop;
+$standby->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
new file mode 100644
index 00000000000..315b27b93d4
--- /dev/null
+++ b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
@@ -0,0 +1,125 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that a standby crashing after a replayed online disable, but before
+# its next restartpoint, restarts cleanly.
+#
+# Pages dirtied before the XLOG2_CHECKSUMS record and evicted after it are
+# written out under the new "off" state, without checksums. The control
+# file must follow the record immediately in this direction: a crash-restart
+# initializes checksum verification from the control file, and one still
+# saying "on" would fail verification on exactly those pages while replaying
+# records older than the state change.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$primary->start;
+
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,1000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+
+# Small buffer pool so that replaying a bulk load evicts the pages we care
+# about, and no restartpoints so the control file stays where it is.
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 1MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_delay = 10000
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'on');
+
+# No full page images, so that replay has to read the pages of "t" back from
+# disk instead of overwriting them from the WAL.
+$primary->append_conf('postgresql.conf', 'full_page_writes = off');
+$primary->reload;
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" on the standby, *before* the state change. They are
+# not flushed: the standby has no restartpoint from here on.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+# Online disable. The standby replays it and switches to "off", but its
+# control file is not updated.
+$primary->safe_psql('postgres', 'SELECT pg_disable_data_checksums();');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+my ($ctl_before) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($ctl_state) = $ctl_before =~ /Data page checksum version:\s+(\d+)/;
+note(
+ "standby control file data_checksum_version after the disable: "
+ . "$ctl_state (0 = off, 1 = on)");
+
+# Force the standby to evict the dirty pages of "t" now that it is "off":
+# they are written back without checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE filler AS SELECT generate_series(1,300000) AS a;");
+$primary->wait_for_catchup($standby);
+
+# Crash the standby before it gets a chance to run a restartpoint.
+$standby->stop('immediate');
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'standby restarts after crashing between a checksum state '
+ . 'change and the next restartpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'the pages written while "off" are readable')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("standby log tail:\n" . join("\n", @lines[ -12 .. -1 ]))
+ if @lines >= 12;
+ fail('the pages written while "off" are readable');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/016_promote_enable_crash.pl b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
new file mode 100644
index 00000000000..97f84da5685
--- /dev/null
+++ b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a standby promoted right after replaying an online enable, before any
+# restartpoint has flushed the rewritten pages.
+#
+# The end-of-recovery record persists the checksum state without flushing the
+# buffer pool, while the control file still points at a restartpoint older than
+# the transition. Persisting "on" there would make a crash before the
+# post-promotion checkpoint verify checksums over the pages that were flushed
+# while the state was still "off"; promotion must take the full checkpoint
+# instead.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Dirty the pages of "t" without full page images, then push them out to disk
+# on the standby while checksums are still off.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Online enable on the primary; the standby replays it into its own buffers,
+# where the rewritten pages stay dirty (no restartpoint, no bgwriter).
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+note("standby control file before the promotion: $before");
+
+# Promote. The end-of-recovery record persists the live state without any
+# buffer flush; the checkpoint requested afterwards is not immediate.
+$standby->promote;
+$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery();')
+ or die 'timed out waiting for the promotion';
+$standby->stop('immediate');
+
+($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($ckpt) = $ctl =~ /Latest checkpoint location:\s+(\S+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note(
+ "standby control file after the promotion: $after, checkpoint $ckpt, redo $redo"
+);
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'promoted node restarts after crashing before its first checkpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/017_restartpoint_race.pl b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
new file mode 100644
index 00000000000..65aa834f366
--- /dev/null
+++ b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
@@ -0,0 +1,138 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a restartpoint whose flush races the replay of an online enable.
+#
+# Replay keeps running while CheckPointGuts() writes out the buffer pool, so
+# the state at the end of the flush can be newer than the one the flush ran
+# under. Persisting it would claim checksums for pages the flush wrote out
+# before the transition, against a redo pointer that predates it.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# C1: the checkpoint the pending restartpoint will target.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" without full page images and push them to disk on
+# the standby while it is still "off".
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Start a restartpoint for C1 and hold it right after CheckPointGuts().
+$standby->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point','wait');");
+my $bg = $standby->background_psql('postgres');
+$bg->query_until(qr//, "\\echo restartpoint\nCHECKPOINT;\n");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint injection point';
+
+# The startup process keeps replaying while the checkpointer is held: the
+# whole online enable lands, and the rewritten pages stay dirty.
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+# Let the restartpoint finish; it persists the state it samples now.
+$standby->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint to finish';
+$standby->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+
+$standby->stop('immediate');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note("standby control file after the restartpoint: $after, redo $redo");
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started, 'standby restarts after a restartpoint that raced the enable');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/018_enable_crash_windows.pl b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
new file mode 100644
index 00000000000..17f96ff6272
--- /dev/null
+++ b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
@@ -0,0 +1,496 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Crash and checkpoint windows inside an online data checksum enable on
+# a single node. Each scenario starts from "off" on the same cluster,
+# runs the enable into a held injection point, and checks what a crash
+# or a concurrent checkpoint in that window may and may not do to the
+# control file and the pages on disk.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+use IPC::Run;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Bring the node back to "off" with no leftover table between
+# scenarios. The crash restarts already resolve an interrupted enable
+# to "off", so only disable when something is left to disable. Those
+# restarts also clear any attached injection points. The checkpoint
+# bounds the WAL the next scenario's crash recovery has to replay.
+sub reset_scenario
+{
+ my ($node) = @_;
+
+ BAIL_OUT('node is not running after the previous scenario')
+ unless $node->is_alive;
+
+ my $state = $node->safe_psql('postgres', 'SHOW data_checksums;');
+ disable_data_checksums($node, wait => 'off') if $state ne 'off';
+ $node->safe_psql('postgres', 'DROP TABLE IF EXISTS t;');
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+ test_checksum_state($node, 'off');
+}
+
+my $node = PostgreSQL::Test::Cluster->new('crash_windows');
+$node->init(no_data_checksums => 1);
+
+# Scenario 1 needs the pages of "t" to reach disk without a full page
+# image, hence full_page_writes and wal_log_hints off. Scenario 2
+# needs them to stay dirty in shared buffers until a checkpoint, hence
+# no bgwriter and a large shared_buffers. wal_level is pinned because
+# init() writes "minimal" for a non-streaming node. The rest merely
+# tolerate the settings.
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+wal_level = replica
+max_wal_size = 10GB
+full_page_writes = off
+wal_log_hints = off
+shared_buffers = 512MB
+bgwriter_lru_maxpages = 0
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$node->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+
+test_checksum_state($node, 'off');
+
+# Scenario 1: a primary crashes inside SetDataChecksumsOn(), just before the
+# forced checkpoint that flushes the rewritten pages, with crash recovery
+# resuming from a checkpoint older than the transition. The control file
+# must still say "inprogress-on" there: replay does not adopt the state of
+# the checkpoint record it resumes from, so an "on" written before the flush
+# would stay in effect while replay reads pages whose rewrite never reached
+# disk. Here the pages went out through the rewriting worker's ring buffer,
+# so only the control file state discriminates; see the next scenario for
+# the case where they did not.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Establish the checkpoint that crash recovery will resume from.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty the pages of "t" without full page images, then push them out
+ # to disk while checksums are still off.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+ my $evicted = $node->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+ note("evict_relation on t: $evicted, relpath $relpath");
+
+ # Stop the enable right after the control file has been updated to
+ # "on" but before the checkpoint that flushes the rewritten pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+ $node->safe_psql('postgres', 'SELECT pg_enable_data_checksums();');
+
+ $node->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'datachecksums-on-before-checkpoint';"
+ ) or die 'timed out waiting for the injection point';
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ $node->stop('immediate');
+
+ # Show the on-disk checksum field of the first page of "t".
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 2: the same crash as above, but with the rewritten pages resident
+# in shared buffers instead of gone out through the ring. The rewriting
+# worker reads through a BAS_VACUUM ring, which writes pages back as the
+# ring recycles, but a page already resident in shared buffers is not read
+# through the ring: ReadBufferExtended() hands back the existing buffer, and
+# it stays dirty until a checkpoint. The control file may therefore not
+# say "on" before that checkpoint has run, or crash recovery would resume
+# from a checkpoint older than the transition with verification already
+# enabled, and replay records that read those still-unchecksummed pages.
+# full_page_writes is off so that the records do not simply overwrite the
+# pages with a full page image.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,100000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Establish the checkpoint that crash recovery will resume from, with the
+ # pages of "t" written out while checksums are still off.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty those pages again without emitting full page images, and leave
+ # them in shared buffers. Replay of these records has to read the
+ # pages from disk.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+
+ my $dirty_before = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_before, '>', 0,
+ 'pages of t are resident and dirty before enabling checksums');
+
+ # Hold the enabling right before the checkpoint that flushes the rewritten
+ # pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ # The rewritten pages must still be sitting dirty in shared buffers, or
+ # the window this test is about does not exist.
+ my $dirty_after = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_after, '>', 0,
+ 'rewritten pages of t are still dirty before the checkpoint');
+
+ $node->stop('immediate');
+
+ # The on-disk copy is the one written before the transition, without a
+ # checksum.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ is($pd_checksum, 0, 'block 0 of t on disk carries no checksum');
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 3: a checkpoint that runs between the XLOG2_CHECKSUMS("on")
+# record and the control file write at the end of SetDataChecksumsOn() must
+# not make a crash throw the completed transition away. SetDataChecksumsOn()
+# writes the record, flips shared memory to "on", emits the barrier and only
+# then requests the checkpoint that flushes the rewritten pages; the control
+# file is written after that checkpoint returns. A crash in that window is
+# harmless only as long as recovery still starts before the record. Any
+# checkpoint completing in the window moves the redo point past the record,
+# so recovery would never see it, would come up with the control file's
+# "inprogress-on" and StartupXLOG() would demote that to "off", even though
+# every page on disk carries a checksum by then. CreateCheckPoint()
+# therefore persists the state the checkpoint ran under, the same way
+# CreateRestartPoint() does. Here the window is made deterministic by
+# holding the launcher at the datachecksums-on-before-checkpoint injection
+# point and checkpointing from another session.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Hold the launcher after the record, the shared memory flip and the
+ # barrier, but before the checkpoint SetDataChecksumsOn() requests
+ # itself.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ # Every backend already sees "on" and writes checksums.
+ test_checksum_state($node, 'on');
+
+ # ... while the control file still says "inprogress-on".
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($before, '3', 'control file says "inprogress-on" inside the window');
+
+ # A concurrent checkpoint. It flushes the rewritten pages and moves
+ # the redo point past the XLOG2_CHECKSUMS("on") record, so it has to
+ # record the "on" state in the control file as well.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($after, '1',
+ 'the concurrent checkpoint records "on" in the control file');
+
+ $node->stop('immediate');
+
+ # The checkpoint flushed the rewritten pages, so they carry a checksum on
+ # disk: the transition really did complete.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ isnt($pd_checksum, 0, 'block 0 of t on disk carries a checksum');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition is complete on disk, so the cluster has to come back
+ # "on".
+ test_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+}
+reset_scenario($node);
+
+# Scenario 4: a crash between the checkpoint an online enable requests and
+# the control file write that follows it must not lose the transition. The
+# last steps of SetDataChecksumsOn() are
+#
+# WAL record -> shmem -> barrier -> checkpoint -> persist
+#
+# The checkpoint is what makes "on" safe to persist, but it also moves the
+# redo point above the XLOG2_CHECKSUMS record that carries the new state.
+# Crash recovery started from that checkpoint therefore never replays the
+# record, so without the state the checkpoint itself persists, a crash in
+# the remaining window would bring the cluster back at "inprogress-on",
+# which StartupXLOG() resolves to "off", discarding a transition whose
+# pages are all on disk with a checksum. The previous scenario exercises
+# the same window through a checkpoint requested by another session; this
+# one closes it from the other side: the transition's own checkpoint is the
+# one that moves the redo point, so persisting the state from
+# CreateCheckPoint() is what has to save it, not the write below the
+# injection point.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Hold the launcher after the checkpoint that licenses "on" has
+ # completed and before the state reaches the control file.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-after-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-after-checkpoint');
+
+ # The transition is complete as far as the running cluster is concerned.
+ test_checksum_state($node, 'on');
+
+ # The checkpoint has already recorded it, so the pending write below the
+ # injection point has nothing left to do.
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($version) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($version, '1',
+ 'the requested checkpoint recorded "on" in the control file');
+
+ # Crash before the control file write that follows the checkpoint.
+ $node->stop('immediate');
+
+ $node->start;
+
+ # Every page on disk carries a checksum and the checkpoint that
+ # flushed them completed, so the cluster has to come back verifying
+ # them.
+ test_checksum_state($node, 'on');
+
+ is($node->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '10000', 'relation readable after the crash');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ # The data directory must be one the offline tools accept.
+ $node->stop;
+ $node->command_ok([ 'pg_checksums', '--check', '-D', $node->data_dir ],
+ 'pg_checksums accepts the data directory');
+
+ # Leave the node running for the reset.
+ $node->start;
+}
+reset_scenario($node);
+
+# Scenario 5: a checkpoint racing SetDataChecksumsOn() between the insertion
+# of the XLOG2_CHECKSUMS("on") record and the shared memory update must not
+# insert an XLOG_CHECKPOINT_REDO record that follows the transition in WAL
+# order while still carrying "inprogress-on". Recovery resuming from such a
+# redo point never replays the preceding transition record, comes up in
+# "inprogress-on" and resolves the finished transition as interrupted.
+# XLogChecksums() closes the window by inserting the record and publishing
+# the new state under DataChecksumTransitionLock, which CreateCheckPoint()
+# takes around sampling the state and inserting the redo record. Here the
+# launcher is held between the two steps at the
+# datachecksums-on-before-publish injection point, and a concurrent
+# CHECKPOINT has to block on the lock instead of completing inside the
+# window.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # The datachecksums-on-before-publish point fires inside a critical
+ # section, where the wait machinery must not allocate. Waiting once
+ # at the datachecksums-enable-checksums-delay point, which the
+ # launcher runs outside the critical section, initializes it; see
+ # 050_redo_segment_missing.pl for the same recipe around
+ # create-checkpoint-run.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-enable-checksums-delay','wait');"
+ );
+
+ # Hold the launcher after the XLOG2_CHECKSUMS("on") record is in WAL but
+ # before the new state is published in shared memory.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-publish','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-enable-checksums-delay');
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-enable-checksums-delay');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-enable-checksums-delay');");
+
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-publish');
+
+ # The record is in WAL, the published state is still the old one.
+ test_checksum_state($node, 'inprogress-on');
+
+ # A concurrent checkpoint. It must block on DataChecksumTransitionLock
+ # before inserting its redo record rather than complete inside the window.
+ my $checkpointer = IPC::Run::start(
+ [ 'psql', '-XAtq', '-d', $node->connstr('postgres'), '-c', 'CHECKPOINT;' ],
+ '<' => '/dev/null',
+ '>' => '/dev/null',
+ '2>' => '/dev/null',
+ IPC::Run::timer($PostgreSQL::Test::Utils::timeout_default));
+
+ ok( $node->poll_query_until(
+ 'postgres',
+ "SELECT wait_event = 'DataChecksumTransition' "
+ . "FROM pg_stat_activity WHERE backend_type = 'checkpointer';"),
+ 'concurrent checkpoint blocks on DataChecksumTransitionLock');
+
+ # Release the launcher; the checkpoint then samples the published "on" and
+ # its redo record follows the transition record.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-publish');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-publish');");
+
+ $checkpointer->finish;
+
+ # Crash while the launcher's own checkpoint may still be in flight.
+ # Recovery resumes from the concurrent checkpoint's redo point, which
+ # now lies above the transition record and carries "on".
+ $node->stop('immediate');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition completed, so the cluster has to come back "on".
+ wait_for_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ $node->stop;
+}
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
new file mode 100644
index 00000000000..aaa35f8f7dd
--- /dev/null
+++ b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
@@ -0,0 +1,126 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A standby stopped after replaying an online enable, but before the
+# checkpoint record that follows it on the primary.
+#
+# The shutdown restartpoint has no new checkpoint record to work from
+# and is skipped, so the control file would keep the in-progress state
+# the transition passed through, even though replay left the node at
+# "on" and rewrote every page. pg_checksums reads that field and
+# refuses to run on an in-progress state, so the shutdown has to flush
+# and catch it up instead.
+#
+# The caught-up control file then carries the watermark of the "on"
+# record, which the second half of this test relies on: an offline
+# disable made while the standby is down must survive the re-replay of
+# that record on the next startup, or the change would be silently
+# reverted.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+wal_keep_size = 1GB
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Anchor the standby's last restartpoint here, so that nothing replayed
+# from now on gives the shutdown restartpoint a newer checkpoint record
+# to use.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Hold the enable right after the state change record, before the
+# checkpoint it requests once the transition is complete.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+enable_data_checksums($primary);
+$primary->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+wait_for_checksum_state($primary, 'on');
+
+# The state change record is already flushed and streams on its own;
+# no checkpoint record follows while the launcher is held.
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+$standby->stop('fast');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($state, '1',
+ 'shutdown catches the control file up with the replayed state');
+
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'pg_checksums verifies the stopped standby');
+
+# Disable checksums offline while the standby is down. The restart
+# below resumes replay under the last restartpoint, re-reading the "on"
+# record; the watermark persisted with the catch-up above marks it as
+# already applied, so the offline change must survive.
+$standby->checksum_disable_offline;
+
+# Let the enable finish on the primary.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-checkpoint');");
+$primary->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-checkpoint');");
+
+# The launcher exits only after the checkpoint it requests completes,
+# so a checkpoint record now follows the transition in WAL.
+$primary->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_catalog.pg_stat_activity "
+ . "WHERE backend_type = 'datachecksums launcher';");
+
+my $log_offset = -s $standby->logfile;
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+
+# The nodes legitimately diverged, which replayed checkpoints report.
+$standby->wait_for_log(
+ qr/data checksum state "off" of this node does not match the state "on" in the replayed WAL/,
+ $log_offset);
+
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/020_cascade_divergence.pl b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
new file mode 100644
index 00000000000..b4b474fa583
--- /dev/null
+++ b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
@@ -0,0 +1,115 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# An offline data checksum state change made on one node of a cascading setup
+# is reported all the way down the chain.
+#
+# CheckReplayedDataChecksumState() is reached from xlog_redo() when a
+# checkpoint record (XLOG_CHECKPOINT_SHUTDOWN or XLOG_CHECKPOINT_REDO) is
+# replayed after consistency has been reached. Those records are written by
+# the root primary only and are relayed verbatim by every intermediate
+# standby, so a cascaded standby that was changed offline learns about the
+# divergence too. An intermediate standby's own restartpoints are not
+# WAL-logged and cannot surface it.
+#
+# Note that this makes detection checkpoint-driven: an idle or read-only
+# primary surfaces the divergence no sooner than its next checkpoint, so up to
+# checkpoint_timeout may pass before the operator sees anything. Closing that
+# window would require comparing the state when a walreceiver connects.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# checkpoint_timeout is deliberately long so that the only checkpoint record in
+# this test is the one requested explicitly below.
+my $conf = qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+);
+
+my $primary = PostgreSQL::Test::Cluster->new('cascade_primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', $conf);
+$primary->start;
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,1000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('cascade_standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf('postgresql.conf', $conf);
+$standby->start;
+
+$standby->backup('backup');
+my $cascade = PostgreSQL::Test::Cluster->new('cascade_cascade');
+$cascade->init_from_backup($standby, 'backup', has_streaming => 1);
+$cascade->append_conf('postgresql.conf', $conf);
+$cascade->start;
+
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+test_checksum_state($cascade, 'off');
+
+# Change the cascaded standby offline, and only it.
+$cascade->stop;
+$cascade->command_ok([ 'pg_checksums', '--enable', '-D', $cascade->data_dir ],
+ 'pg_checksums enables checksums on the cascaded standby only');
+
+my $logstart = -s $cascade->logfile;
+$cascade->start;
+
+test_checksum_state($cascade, 'on');
+
+# Ordinary WAL traffic carries no state, so the cascaded standby keeps
+# streaming from a chain whose state is "off" without noticing.
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (1);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+is($cascade->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '1001', 'cascaded standby keeps streaming from a divergent chain');
+
+# Restartpoints on the intermediate standby are not WAL-logged, so this does
+# not surface anything on the cascaded standby either.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (2);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+my $log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+unlike(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'an intermediate restartpoint does not surface the divergence');
+
+# A checkpoint on the root primary does: the record is relayed down the whole
+# chain, so both the direct and the cascaded standby compare it against their
+# own state. wait_for_catchup() below waits for replay, not just receipt, so
+# the record has been through xlog_redo() on both nodes once it returns.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+$log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+like(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'the cascaded standby reports the divergence on the primary checkpoint');
+
+$cascade->stop;
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
new file mode 100644
index 00000000000..26b34778181
--- /dev/null
+++ b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
@@ -0,0 +1,172 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind across online data checksum transitions on both sides of a
+# divergence, with the two nodes trading roles between the scenarios.
+#
+# Scenario 1: after a switchover the new primary enables checksums
+# online, while the old primary restarts on its old timeline, advances
+# its WAL beyond the enable records and runs an online enable/disable
+# cycle of its own. The old primary ends "off" with a checksum
+# watermark numerically above every checksum record the new primary has
+# written. pg_rewind clamps the watermark it keeps to the divergence
+# point; without the clamp, replay on the rewound node would skip the
+# source's enable as already applied and stay "off" under an "on"
+# primary.
+#
+# Scenario 2: checksums are disabled again, and after another
+# switchover both nodes enable them online independently, so the
+# divergence checkpoint carries "off" while both control files say
+# "on". The rewind is allowed, the target keeps its own "on" state,
+# and replay re-walks the source's enable onto the already enabled
+# node.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+sub controldata_watermark
+{
+ my ($node) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $node->data_dir ]);
+ $stdout =~ /^Data checksum watermark:\s*([0-9A-F]+)\/([0-9A-F]+)$/m
+ or die "watermark missing from pg_controldata output";
+ return (hex($1) << 32) + hex($2);
+}
+
+# Old primary, checksums off. wal_log_hints is required by pg_rewind
+# on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+# Clean switchover to B; enable checksums online on it.
+$node_a->stop('fast');
+$node_b->promote;
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+$node_b->stop('fast');
+my $watermark_b = controldata_watermark($node_b);
+$node_b->start;
+
+# Accidental restart of the old primary on the old timeline. Advance
+# its WAL beyond the enable watermark of B, then run an online enable
+# and disable cycle: the node ends "off" with a watermark above every
+# checksum record B has written.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t_pad AS SELECT generate_series(1,200000) AS a;");
+enable_data_checksums($node_a, wait => 'on');
+disable_data_checksums($node_a, wait => 'off');
+test_checksum_state($node_a, 'off');
+$node_a->stop('fast');
+
+my $watermark_a = controldata_watermark($node_a);
+die "test broken: target watermark not above the source's enable"
+ unless $watermark_a > $watermark_b;
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the promoted node');
+
+# Start the rewound node as a standby of B. Replay from the last
+# common checkpoint runs through B's online enable, which the target
+# never saw, so the rewound node must converge to "on".
+my $connstr_b = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr_b application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+# Scenario 2, reusing the pair with the roles reversed. Disable
+# checksums online so the next divergence point carries "off", and let
+# A replay the change.
+disable_data_checksums($node_b, wait => 'off');
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'off');
+
+# Clean switchover back to A; enable checksums online on it.
+$node_b->stop('fast');
+$node_a->promote;
+enable_data_checksums($node_a, wait => 'on');
+test_checksum_state($node_a, 'on');
+
+# The old primary restarts on its old timeline and enables checksums
+# online independently: both control files say "on", the divergence
+# checkpoint says "off".
+$node_b->start;
+$node_b->safe_psql('postgres', "INSERT INTO t_div VALUES (2);");
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+$node_b->stop('fast');
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_b->data_dir,
+ '--source-server' => $node_a->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled on both nodes');
+
+my $connstr_a = $node_a->connstr;
+$node_b->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_b->port]}
+primary_conninfo = '$connstr_a application_name=@{[$node_b->name]}'
+]);
+$node_b->set_standby_mode;
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_b, 'on');
+
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_b->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_b->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the twice-rewound node');
+
+done_testing();
--
2.55.0
[application/octet-stream] v8-0003-pg_rewind-Check-the-data-checksum-states-of-sourc.patch (21.6K, ../../CAN4CZFN5sOQvWTXr7Dm1J1HXwczikBfMjkbFQ7QmHsr+VHd5Mg@mail.gmail.com/6-v8-0003-pg_rewind-Check-the-data-checksum-states-of-sourc.patch)
download | inline diff:
From 6e637d83d6bc4245086a61247f26460be14637ca Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 18:40:22 +0000
Subject: [PATCH v8 3/5] pg_rewind: Check the data checksum states of source
and target
pg_rewind installs the source's control file on the target, but every
block it does not copy keeps the target's content. With checksums
enabled on the target and disabled on the source, the rewound server
claims enabled checksums while the blocks copied from the source have
none, and fails checksum verification as soon as it reads them; in the
worst case every connection attempt dies on an unverifiable catalog
page.
Comparing the control files is not enough. Replay on the rewound
server resumes from the last common checkpoint and adopts the data
checksum state recorded there, so an offline disable on the source
after the divergence leaves both control files saying "off" while the
rewound server still resumes with verification enabled. Check the
state carried by the divergence checkpoint as well, which pg_rewind
already reads.
Refuse both cases, and an interrupted online transition on either
side, same as pg_checksums does. The opposite mismatch, checksums on
the source only, is allowed with a warning: recovery under the backup
label adopts the divergence state, so the rewound server keeps
checksums disabled (or converges through the replayed WAL if the
source enabled them online), and no page can fail verification.
---
doc/src/sgml/ref/pg_rewind.sgml | 14 ++
src/bin/pg_rewind/parsexlog.c | 4 +-
src/bin/pg_rewind/pg_rewind.c | 93 +++++++++++-
src/bin/pg_rewind/pg_rewind.h | 1 +
src/test/modules/test_checksums/meson.build | 2 +
.../test_checksums/t/022_rewind_state.pl | 133 +++++++++++++++++
.../t/023_rewind_standby_target.pl | 141 ++++++++++++++++++
7 files changed, 386 insertions(+), 2 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/022_rewind_state.pl
create mode 100644 src/test/modules/test_checksums/t/023_rewind_standby_target.pl
diff --git a/doc/src/sgml/ref/pg_rewind.sgml b/doc/src/sgml/ref/pg_rewind.sgml
index b95e3868e9d..ac3d0c9328f 100644
--- a/doc/src/sgml/ref/pg_rewind.sgml
+++ b/doc/src/sgml/ref/pg_rewind.sgml
@@ -124,6 +124,20 @@ PostgreSQL documentation
<literal>on</literal>, but is enabled by default.
</para>
+ <para>
+ The data checksum states of the source and target server must be
+ compatible. <application>pg_rewind</application> refuses to run if an
+ online checksum state transition was interrupted on either server, or
+ if data checksums are enabled on the target server, or were enabled at
+ the point of divergence, while the source server runs without them: the
+ rewound server would fail checksum verification on the blocks copied
+ from the source. The opposite combination is allowed with a warning;
+ the rewound server keeps data checksums disabled unless the source
+ server enabled them online after the point of divergence. See
+ <xref linkend="app-pgchecksums"/> for how to change the state
+ consistently across a replication setup.
+ </para>
+
<warning>
<title>Warning: Failures While Rewinding</title>
<para>
diff --git a/src/bin/pg_rewind/parsexlog.c b/src/bin/pg_rewind/parsexlog.c
index 023e23b063c..6e87b00f8c2 100644
--- a/src/bin/pg_rewind/parsexlog.c
+++ b/src/bin/pg_rewind/parsexlog.c
@@ -167,7 +167,8 @@ readOneRecord(const char *datadir, XLogRecPtr ptr, int tliIndex,
void
findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
- XLogRecPtr *lastchkptredo, const char *restoreCommand)
+ XLogRecPtr *lastchkptredo, uint32 *lastchkptdatachecksums,
+ const char *restoreCommand)
{
/* Walk backwards, starting from the given record */
XLogRecord *record;
@@ -255,6 +256,7 @@ findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
*lastchkptrec = searchptr;
*lastchkpttli = checkPoint.ThisTimeLineID;
*lastchkptredo = checkPoint.redo;
+ *lastchkptdatachecksums = checkPoint.dataChecksumState;
break;
}
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 466f4223501..9b99e628d4e 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -145,6 +145,7 @@ main(int argc, char **argv)
XLogRecPtr chkptrec;
TimeLineID chkpttli;
XLogRecPtr chkptredo;
+ uint32 chkptdatachecksums;
TimeLineID source_tli;
TimeLineID target_tli;
XLogRecPtr target_wal_endrec;
@@ -472,10 +473,51 @@ main(int argc, char **argv)
keepwal_init();
findLastCheckpoint(datadir_target, divergerec, lastcommontliIndex,
- &chkptrec, &chkpttli, &chkptredo, restore_command);
+ &chkptrec, &chkpttli, &chkptredo, &chkptdatachecksums,
+ restore_command);
pg_log_info("rewinding from last common checkpoint at %X/%08X on timeline %u",
LSN_FORMAT_ARGS(chkptrec), chkpttli);
+ /*
+ * Replay on the rewound server resumes from the last common checkpoint
+ * and adopts the data checksum state recorded there, not the state in the
+ * control file installed from the source. If checksums were enabled at
+ * the divergence point but the source runs without them, the rewound
+ * server would verify checksums while the blocks copied from the source
+ * have none. The control files cannot reveal this: an offline disable on
+ * the source after the divergence leaves both of them saying "off".
+ *
+ * Only the fully enabled state needs checking. An in-progress state at
+ * the divergence point means a transition was still running there, and
+ * all of its page rewrites are logged after that point, so replay brings
+ * the rewound server to whatever state the copied WAL ends in.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the source server or recreate the target server from a base backup.");
+ exit(1);
+ }
+
+ /*
+ * The same comparison is needed against the target: every block the
+ * rewind does not copy keeps the target's content. The control files
+ * cannot reveal this case either, in the other direction: a standby's
+ * checkpoints are written by its upstream primary, so a standby whose
+ * checksums were disabled offline still has "on" checkpoints in its WAL,
+ * while both control files may agree.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_target.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the target server");
+ pg_log_error_detail("Blocks kept from the target would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the target server with pg_checksums or recreate the target server from a base backup.");
+ exit(1);
+ }
+
/* Initialize the hash table to track the status of each file */
filehash_init();
@@ -799,6 +841,55 @@ sanityChecks(void)
pg_fatal("target server needs to use either data checksums or \"wal_log_hints = on\"");
}
+ /*
+ * The rewound target keeps its control file fields from the source, but
+ * every block the rewind does not copy keeps the target's content. If
+ * checksums are enabled on the target and disabled on the source, the
+ * result would claim enabled checksums while the blocks copied from the
+ * source have none, and the target would fail checksum verification as
+ * soon as it reads them. Refuse that combination, and refuse an
+ * interrupted online transition on either side, same as pg_checksums.
+ *
+ * The opposite mismatch is allowed: recovery under the backup label
+ * written by pg_rewind adopts the checksum state as of the divergence
+ * point, so a target whose own checkpoints carry no checksums stays
+ * without them, or converges through the replayed WAL if the source
+ * enabled them online. Only warn about it, so that an offline change on
+ * the source is not overlooked.
+ *
+ * That reasoning does not hold when the target is a standby, whose
+ * checkpoints were written by its upstream primary; see the checks after
+ * findLastCheckpoint().
+ */
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted on the target server");
+ pg_log_error_hint("Start the server and shut it down cleanly to reset the state, then retry; on a standby, let replication complete the transition first.");
+ exit(1);
+ }
+ if (ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition is incomplete on the source server");
+ pg_log_error_hint("Let the transition complete, or reset the state with a clean restart, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_OFF)
+ {
+ pg_log_error("data checksums are enabled on the target server but disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, and the target would fail checksum verification after the rewind.");
+ pg_log_error_hint("Either disable data checksums on the target server or enable them on the source server, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_OFF &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_warning("data checksums are disabled on the target server but enabled on the source server");
+ pg_log_warning_detail("The rewound server will keep data checksums disabled unless the source server enabled them online after the point of divergence.");
+ }
+
/*
* Target cluster better not be running. This doesn't guard against
* someone starting the cluster concurrently. Also, this is probably more
diff --git a/src/bin/pg_rewind/pg_rewind.h b/src/bin/pg_rewind/pg_rewind.h
index 9a981f7f246..9187c3bf72f 100644
--- a/src/bin/pg_rewind/pg_rewind.h
+++ b/src/bin/pg_rewind/pg_rewind.h
@@ -39,6 +39,7 @@ extern void findLastCheckpoint(const char *datadir, XLogRecPtr forkptr,
int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
XLogRecPtr *lastchkptredo,
+ uint32 *lastchkptdatachecksums,
const char *restoreCommand);
extern XLogRecPtr readOneRecord(const char *datadir, XLogRecPtr ptr,
int tliIndex, const char *restoreCommand);
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index e5d38fafb7d..7d07c757052 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -45,6 +45,8 @@ tests += {
't/019_standby_shutdown_catchup.pl',
't/020_cascade_divergence.pl',
't/021_rewind_divergent_transitions.pl',
+ 't/022_rewind_state.pl',
+ 't/023_rewind_standby_target.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/022_rewind_state.pl b/src/test/modules/test_checksums/t/022_rewind_state.pl
new file mode 100644
index 00000000000..b7adc968c4d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/022_rewind_state.pl
@@ -0,0 +1,133 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind checks the data checksum states of source and target.
+# Checksums enabled on the target, or at the point of divergence, with
+# a source running without them are refused: the rewound server would
+# verify checksums on blocks copied from a source that has none. The
+# opposite mismatch only warns; replay keeps the target's own state.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# Scenarios 1 and 2: checksums on from initdb. wal_log_hints keeps the
+# target eligible for pg_rewind after checksums are disabled on it.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b);
+
+# Failover to B, and divergence on A.
+$node_b->promote;
+$node_b->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_a->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_a->stop('fast');
+
+# Offline disable on the source only: enabled target, disabled source.
+$node_b->stop;
+$node_b->checksum_disable_offline;
+$node_b->start;
+test_checksum_state($node_b, 'off');
+
+my @rewind_cmd = (
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'));
+
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums are enabled on the target server but disabled on the source server/,
+ 'refuses an enabled target with a disabled source');
+
+# The refusal happens before any modification: the target still starts
+# on its own timeline.
+$node_a->start;
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'target untouched by the refused rewind');
+$node_a->stop('fast');
+
+# Scenario 2: disabling the target too makes the control files match,
+# but checksums were still enabled at the point of divergence, which is
+# the state replay on the rewound server would resume with.
+$node_a->checksum_disable_offline;
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums were enabled at the point of divergence but are disabled on the source server/,
+ 'refuses when checksums were enabled at the point of divergence');
+
+# Scenario 3: fresh pair without checksums; offline enable on the
+# source only. Allowed with a warning, and the rewound server keeps
+# checksums disabled.
+my $node_c = PostgreSQL::Test::Cluster->new('node_c');
+$node_c->init(allows_streaming => 1, no_data_checksums => 1);
+$node_c->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_c->start;
+$node_c->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_c->backup('backup');
+my $node_d = PostgreSQL::Test::Cluster->new('node_d');
+$node_d->init_from_backup($node_c, 'backup', has_streaming => 1);
+$node_d->start;
+$node_c->wait_for_catchup($node_d);
+
+$node_d->promote;
+$node_d->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_c->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_c->stop('fast');
+
+$node_d->stop;
+$node_d->checksum_enable_offline;
+$node_d->start;
+test_checksum_state($node_d, 'on');
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_c->data_dir,
+ '--source-server' => $node_d->connstr('postgres'),
+ ]);
+like(
+ $stderr,
+ qr/data checksums are disabled on the target server but enabled on the source server/,
+ 'warns for a disabled target with an enabled source');
+like($stderr, qr/Done!/, 'rewind completed despite the warning');
+
+# The rewound server follows D and keeps its own state.
+$node_c->append_conf('postgresql.conf', 'port = ' . $node_c->port);
+$node_c->enable_streaming($node_d);
+$node_c->set_standby_mode;
+$node_c->start;
+$node_d->wait_for_catchup($node_c);
+is($node_c->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+test_checksum_state($node_c, 'off');
+
+$node_c->stop;
+$node_d->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/023_rewind_standby_target.pl b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
new file mode 100644
index 00000000000..3f5a3c6be8b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
@@ -0,0 +1,141 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_rewind compares the data checksum state at the point of
+# divergence against the target as well as the source.
+#
+# A standby's checkpoints are written by its upstream primary, so a standby
+# whose checksums were disabled offline still has "on" checkpoints in its
+# WAL. Rewinding it onto a promoted sibling passes every control file
+# check (target off + source on only warns), yet recovery after the rewind
+# adopts the divergence checkpoint's state and the rewound server would
+# verify checksums over the pages it wrote while it was locally "off".
+# pg_rewind must refuse; after an offline enable on the target, which
+# rewrites all of its pages, the rewind goes through.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t0 AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+
+my $lagging = PostgreSQL::Test::Cluster->new('lagging');
+$lagging->init_from_backup($primary, 'backup', has_streaming => 1);
+$lagging->start;
+
+my $failover = PostgreSQL::Test::Cluster->new('failover');
+$failover->init_from_backup($primary, 'backup', has_streaming => 1);
+$failover->start;
+
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+
+test_checksum_state($primary, 'on');
+test_checksum_state($lagging, 'on');
+test_checksum_state($failover, 'on');
+
+# Offline-disable checksums on the lagging standby only. This is a
+# divergence 012_offline_standby.pl declares survivable: the standby keeps
+# its own state, warns, and stays readable.
+$lagging->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $lagging->data_dir);
+$lagging->start;
+test_checksum_state($lagging, 'off');
+
+# Give the lagging standby plenty of pages to write while it is locally
+# "off", and force a restartpoint so they reach disk without checksums.
+# The other standby replays the same WAL with checksums on.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+$lagging->safe_psql('postgres', "CHECKPOINT;");
+
+# Take the "failover" standby out of the picture. It is promoted from
+# this position later, so everything replayed from here on is the part of
+# the WAL that diverges and that pg_rewind will roll back. t1 is already
+# replayed everywhere, so its blocks are *not* rolled back.
+$failover->stop;
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$primary->wait_for_catchup($lagging);
+
+# Failover: promote the other standby, which forks the timeline behind the
+# replay position of the lagging standby.
+$primary->stop('fast');
+$failover->start;
+$failover->promote;
+$failover->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$failover->safe_psql('postgres', "CHECKPOINT;");
+
+# Re-attaching the lagging standby with pg_rewind must be refused: the
+# divergence checkpoint says "on" while the target is "off", so the rewound
+# server would verify checksums over the blocks it keeps.
+$lagging->stop;
+
+command_fails_like(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ],
+ qr/data checksums were enabled at the point of divergence but are disabled on the target server/,
+ 'pg_rewind refuses a target whose divergence checkpoint has checksums');
+
+# Enabling checksums offline on the target rewrites all of its pages, after
+# which the rewind is safe.
+system_or_bail('pg_checksums', '--enable', '--pgdata', $lagging->data_dir);
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ]);
+like($stderr, qr/Done!/, 'pg_rewind completes after the offline enable');
+
+$lagging->append_conf('postgresql.conf', 'port = ' . $lagging->port);
+$lagging->enable_streaming($failover);
+$lagging->set_standby_mode;
+$lagging->start;
+
+my ($rc, $out, $err) =
+ $lagging->psql('postgres',
+ "SELECT setting FROM pg_settings " . "WHERE name = 'data_checksums';");
+is($rc, 0, 'rewound standby accepts connections') or diag("stderr: $err");
+is($out, 'on', 'rewound standby resumes with checksums on');
+
+($rc, $out, $err) = $lagging->psql('postgres', "SELECT count(*) FROM t1;");
+is($rc, 0, 'blocks kept from the target stay readable')
+ or diag("stderr: $err");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($lagging->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the rewound standby');
+
+$lagging->stop('immediate');
+$failover->stop;
+done_testing();
--
2.55.0
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-01 08:38 Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
parent: Zsolt Parragi <zsolt.parragi@percona.com>
0 siblings, 1 reply; 43+ messages in thread
From: Bertrand Drouvot @ 2026-09-01 08:38 UTC (permalink / raw)
To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; pgsql-hackers@lists.postgresql.org
Hi,
On Tue, Sep 01, 2026 at 12:31:42AM +0100, Zsolt Parragi wrote:
> > I found a case where the source's online enable occurs after divergence and was
> > never seen by the target, but replay skips it instead of applying it:
> > ....
> > Maybe the watermark needs timeline context, or pg_rewind needs to adjust it
> > when it comes from the target's divergent history?
>
> Thanks! v8 adds the latter, with a new test case verifying this scenario.
Thanks!
As far the new test:
=== 1
+# Clean switchover back to A; enable checksums online on it.
+$node_b->stop('fast');
+$node_a->promote;
IIUC, the preceding wait_for_catchup() does not cover the shutdown checkpoint
written by stop(). Therefore, the divergence checkpoint in scenario 2 is not
guaranteed to carry off, as described.
=== 2
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
...
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_b, 'on');
The target is already "on" before pg_rewind, so the final assertion does not
prove that the source's enable record was replayed.
Please find attached a small patch addressing those two test comments to apply
on top of v8. What do you think?
=== 3
+ /*
+ * End of the newest XLOG2_CHECKSUMS record this node has written or
+ * applied.
and
+ * would skip them as already applied. Clamp it to the divergence point,
+ * so that every transition record on the source's history takes effect.
+ */
+ if (ControlFile_new.data_checksum_lsn > divergerec)
+ ControlFile_new.data_checksum_lsn = divergerec;
divergerec is not necessarily the end of an XLOG2_CHECKSUMS record, so the
comment no longer describes every value the field may contain. Maybe it should
describe it as the WAL position through which checksum transitions are covered?
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
diff --git a/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
--- a/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
+++ b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
@@ -40,6 +40,19 @@ sub controldata_watermark
return (hex($1) << 32) + hex($2);
}
+sub wait_for_shutdown_checkpoint_replay
+{
+ my ($primary, $standby) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $primary->data_dir ]);
+ $stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+ my $shutdown_checkpoint = $1;
+
+ $standby->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_checkpoint'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+}
+
# Old primary, checksums off. wal_log_hints is required by pg_rewind
# on a cluster without data checksums.
my $node_a = PostgreSQL::Test::Cluster->new('node_a');
@@ -63,6 +76,7 @@ $node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
# Clean switchover to B; enable checksums online on it.
$node_a->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_a, $node_b);
$node_b->promote;
enable_data_checksums($node_b, wait => 'on');
test_checksum_state($node_b, 'on');
@@ -123,9 +137,11 @@ test_checksum_state($node_a, 'off');
# Clean switchover back to A; enable checksums online on it.
$node_b->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_b, $node_a);
$node_a->promote;
enable_data_checksums($node_a, wait => 'on');
test_checksum_state($node_a, 'on');
+my $source_enable_watermark = controldata_watermark($node_a);
# The old primary restarts on its old timeline and enables checksums
# online independently: both control files say "on", the divergence
@@ -162,6 +178,8 @@ is($node_b->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
'0', 'divergent insert was rewound');
$node_b->stop('fast');
+is(controldata_watermark($node_b), $source_enable_watermark,
+ 'rewound node replayed the source checksum transition');
command_ok([ 'pg_checksums', '--check', '-D', $node_b->data_dir ],
'checksums valid on the rewound node');
Attachments:
[text/plain] v8-021-test-fixes.txt (2.2K, ../../apaPDmrlhtXgKR+E@bdtpg/2-v8-021-test-fixes.txt)
download | inline diff:
diff --git a/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
--- a/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
+++ b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
@@ -40,6 +40,19 @@ sub controldata_watermark
return (hex($1) << 32) + hex($2);
}
+sub wait_for_shutdown_checkpoint_replay
+{
+ my ($primary, $standby) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $primary->data_dir ]);
+ $stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+ my $shutdown_checkpoint = $1;
+
+ $standby->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_checkpoint'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+}
+
# Old primary, checksums off. wal_log_hints is required by pg_rewind
# on a cluster without data checksums.
my $node_a = PostgreSQL::Test::Cluster->new('node_a');
@@ -63,6 +76,7 @@ $node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
# Clean switchover to B; enable checksums online on it.
$node_a->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_a, $node_b);
$node_b->promote;
enable_data_checksums($node_b, wait => 'on');
test_checksum_state($node_b, 'on');
@@ -123,9 +137,11 @@ test_checksum_state($node_a, 'off');
# Clean switchover back to A; enable checksums online on it.
$node_b->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_b, $node_a);
$node_a->promote;
enable_data_checksums($node_a, wait => 'on');
test_checksum_state($node_a, 'on');
+my $source_enable_watermark = controldata_watermark($node_a);
# The old primary restarts on its old timeline and enables checksums
# online independently: both control files say "on", the divergence
@@ -162,6 +178,8 @@ is($node_b->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
'0', 'divergent insert was rewound');
$node_b->stop('fast');
+is(controldata_watermark($node_b), $source_enable_watermark,
+ 'rewound node replayed the source checksum transition');
command_ok([ 'pg_checksums', '--check', '-D', $node_b->data_dir ],
'checksums valid on the rewound node');
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-01 13:50 Zsolt Parragi <zsolt.parragi@percona.com>
parent: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 1 reply; 43+ messages in thread
From: Zsolt Parragi @ 2026-09-01 13:50 UTC (permalink / raw)
To: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; pgsql-hackers@lists.postgresql.org
Thanks!
I applied these changes to v9 with some additional comment editing. I
also squashed 0005 into 0001 because it describes what's implemented
there, and I also tried to significantly reduce the commit message of
0001. Otherwise everything else is unchanged.
Attachments:
[application/octet-stream] v9-0003-pg_rewind-Check-the-data-checksum-states-of-sourc.patch (21.6K, ../../CAN4CZFNdfb-yFRW7Sh3FkZ3Qc91Lz-JDQN0b-070-6G35g5Ssg@mail.gmail.com/2-v9-0003-pg_rewind-Check-the-data-checksum-states-of-sourc.patch)
download | inline diff:
From 1ace96adabf25e0c877cf26a568945e6bc50d49b Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 18:40:22 +0000
Subject: [PATCH v9 3/4] pg_rewind: Check the data checksum states of source
and target
pg_rewind installs the source's control file on the target, but every
block it does not copy keeps the target's content. With checksums
enabled on the target and disabled on the source, the rewound server
claims enabled checksums while the blocks copied from the source have
none, and fails checksum verification as soon as it reads them; in the
worst case every connection attempt dies on an unverifiable catalog
page.
Comparing the control files is not enough. Replay on the rewound
server resumes from the last common checkpoint and adopts the data
checksum state recorded there, so an offline disable on the source
after the divergence leaves both control files saying "off" while the
rewound server still resumes with verification enabled. Check the
state carried by the divergence checkpoint as well, which pg_rewind
already reads.
Refuse both cases, and an interrupted online transition on either
side, same as pg_checksums does. The opposite mismatch, checksums on
the source only, is allowed with a warning: recovery under the backup
label adopts the divergence state, so the rewound server keeps
checksums disabled (or converges through the replayed WAL if the
source enabled them online), and no page can fail verification.
---
doc/src/sgml/ref/pg_rewind.sgml | 14 ++
src/bin/pg_rewind/parsexlog.c | 4 +-
src/bin/pg_rewind/pg_rewind.c | 93 +++++++++++-
src/bin/pg_rewind/pg_rewind.h | 1 +
src/test/modules/test_checksums/meson.build | 2 +
.../test_checksums/t/022_rewind_state.pl | 133 +++++++++++++++++
.../t/023_rewind_standby_target.pl | 141 ++++++++++++++++++
7 files changed, 386 insertions(+), 2 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/022_rewind_state.pl
create mode 100644 src/test/modules/test_checksums/t/023_rewind_standby_target.pl
diff --git a/doc/src/sgml/ref/pg_rewind.sgml b/doc/src/sgml/ref/pg_rewind.sgml
index b95e3868e9d..ac3d0c9328f 100644
--- a/doc/src/sgml/ref/pg_rewind.sgml
+++ b/doc/src/sgml/ref/pg_rewind.sgml
@@ -124,6 +124,20 @@ PostgreSQL documentation
<literal>on</literal>, but is enabled by default.
</para>
+ <para>
+ The data checksum states of the source and target server must be
+ compatible. <application>pg_rewind</application> refuses to run if an
+ online checksum state transition was interrupted on either server, or
+ if data checksums are enabled on the target server, or were enabled at
+ the point of divergence, while the source server runs without them: the
+ rewound server would fail checksum verification on the blocks copied
+ from the source. The opposite combination is allowed with a warning;
+ the rewound server keeps data checksums disabled unless the source
+ server enabled them online after the point of divergence. See
+ <xref linkend="app-pgchecksums"/> for how to change the state
+ consistently across a replication setup.
+ </para>
+
<warning>
<title>Warning: Failures While Rewinding</title>
<para>
diff --git a/src/bin/pg_rewind/parsexlog.c b/src/bin/pg_rewind/parsexlog.c
index 023e23b063c..6e87b00f8c2 100644
--- a/src/bin/pg_rewind/parsexlog.c
+++ b/src/bin/pg_rewind/parsexlog.c
@@ -167,7 +167,8 @@ readOneRecord(const char *datadir, XLogRecPtr ptr, int tliIndex,
void
findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
- XLogRecPtr *lastchkptredo, const char *restoreCommand)
+ XLogRecPtr *lastchkptredo, uint32 *lastchkptdatachecksums,
+ const char *restoreCommand)
{
/* Walk backwards, starting from the given record */
XLogRecord *record;
@@ -255,6 +256,7 @@ findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
*lastchkptrec = searchptr;
*lastchkpttli = checkPoint.ThisTimeLineID;
*lastchkptredo = checkPoint.redo;
+ *lastchkptdatachecksums = checkPoint.dataChecksumState;
break;
}
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 466f4223501..9b99e628d4e 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -145,6 +145,7 @@ main(int argc, char **argv)
XLogRecPtr chkptrec;
TimeLineID chkpttli;
XLogRecPtr chkptredo;
+ uint32 chkptdatachecksums;
TimeLineID source_tli;
TimeLineID target_tli;
XLogRecPtr target_wal_endrec;
@@ -472,10 +473,51 @@ main(int argc, char **argv)
keepwal_init();
findLastCheckpoint(datadir_target, divergerec, lastcommontliIndex,
- &chkptrec, &chkpttli, &chkptredo, restore_command);
+ &chkptrec, &chkpttli, &chkptredo, &chkptdatachecksums,
+ restore_command);
pg_log_info("rewinding from last common checkpoint at %X/%08X on timeline %u",
LSN_FORMAT_ARGS(chkptrec), chkpttli);
+ /*
+ * Replay on the rewound server resumes from the last common checkpoint
+ * and adopts the data checksum state recorded there, not the state in the
+ * control file installed from the source. If checksums were enabled at
+ * the divergence point but the source runs without them, the rewound
+ * server would verify checksums while the blocks copied from the source
+ * have none. The control files cannot reveal this: an offline disable on
+ * the source after the divergence leaves both of them saying "off".
+ *
+ * Only the fully enabled state needs checking. An in-progress state at
+ * the divergence point means a transition was still running there, and
+ * all of its page rewrites are logged after that point, so replay brings
+ * the rewound server to whatever state the copied WAL ends in.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the source server or recreate the target server from a base backup.");
+ exit(1);
+ }
+
+ /*
+ * The same comparison is needed against the target: every block the
+ * rewind does not copy keeps the target's content. The control files
+ * cannot reveal this case either, in the other direction: a standby's
+ * checkpoints are written by its upstream primary, so a standby whose
+ * checksums were disabled offline still has "on" checkpoints in its WAL,
+ * while both control files may agree.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_target.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the target server");
+ pg_log_error_detail("Blocks kept from the target would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the target server with pg_checksums or recreate the target server from a base backup.");
+ exit(1);
+ }
+
/* Initialize the hash table to track the status of each file */
filehash_init();
@@ -799,6 +841,55 @@ sanityChecks(void)
pg_fatal("target server needs to use either data checksums or \"wal_log_hints = on\"");
}
+ /*
+ * The rewound target keeps its control file fields from the source, but
+ * every block the rewind does not copy keeps the target's content. If
+ * checksums are enabled on the target and disabled on the source, the
+ * result would claim enabled checksums while the blocks copied from the
+ * source have none, and the target would fail checksum verification as
+ * soon as it reads them. Refuse that combination, and refuse an
+ * interrupted online transition on either side, same as pg_checksums.
+ *
+ * The opposite mismatch is allowed: recovery under the backup label
+ * written by pg_rewind adopts the checksum state as of the divergence
+ * point, so a target whose own checkpoints carry no checksums stays
+ * without them, or converges through the replayed WAL if the source
+ * enabled them online. Only warn about it, so that an offline change on
+ * the source is not overlooked.
+ *
+ * That reasoning does not hold when the target is a standby, whose
+ * checkpoints were written by its upstream primary; see the checks after
+ * findLastCheckpoint().
+ */
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted on the target server");
+ pg_log_error_hint("Start the server and shut it down cleanly to reset the state, then retry; on a standby, let replication complete the transition first.");
+ exit(1);
+ }
+ if (ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition is incomplete on the source server");
+ pg_log_error_hint("Let the transition complete, or reset the state with a clean restart, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_OFF)
+ {
+ pg_log_error("data checksums are enabled on the target server but disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, and the target would fail checksum verification after the rewind.");
+ pg_log_error_hint("Either disable data checksums on the target server or enable them on the source server, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_OFF &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_warning("data checksums are disabled on the target server but enabled on the source server");
+ pg_log_warning_detail("The rewound server will keep data checksums disabled unless the source server enabled them online after the point of divergence.");
+ }
+
/*
* Target cluster better not be running. This doesn't guard against
* someone starting the cluster concurrently. Also, this is probably more
diff --git a/src/bin/pg_rewind/pg_rewind.h b/src/bin/pg_rewind/pg_rewind.h
index 9a981f7f246..9187c3bf72f 100644
--- a/src/bin/pg_rewind/pg_rewind.h
+++ b/src/bin/pg_rewind/pg_rewind.h
@@ -39,6 +39,7 @@ extern void findLastCheckpoint(const char *datadir, XLogRecPtr forkptr,
int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
XLogRecPtr *lastchkptredo,
+ uint32 *lastchkptdatachecksums,
const char *restoreCommand);
extern XLogRecPtr readOneRecord(const char *datadir, XLogRecPtr ptr,
int tliIndex, const char *restoreCommand);
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index e5d38fafb7d..7d07c757052 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -45,6 +45,8 @@ tests += {
't/019_standby_shutdown_catchup.pl',
't/020_cascade_divergence.pl',
't/021_rewind_divergent_transitions.pl',
+ 't/022_rewind_state.pl',
+ 't/023_rewind_standby_target.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/022_rewind_state.pl b/src/test/modules/test_checksums/t/022_rewind_state.pl
new file mode 100644
index 00000000000..b7adc968c4d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/022_rewind_state.pl
@@ -0,0 +1,133 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind checks the data checksum states of source and target.
+# Checksums enabled on the target, or at the point of divergence, with
+# a source running without them are refused: the rewound server would
+# verify checksums on blocks copied from a source that has none. The
+# opposite mismatch only warns; replay keeps the target's own state.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# Scenarios 1 and 2: checksums on from initdb. wal_log_hints keeps the
+# target eligible for pg_rewind after checksums are disabled on it.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b);
+
+# Failover to B, and divergence on A.
+$node_b->promote;
+$node_b->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_a->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_a->stop('fast');
+
+# Offline disable on the source only: enabled target, disabled source.
+$node_b->stop;
+$node_b->checksum_disable_offline;
+$node_b->start;
+test_checksum_state($node_b, 'off');
+
+my @rewind_cmd = (
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'));
+
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums are enabled on the target server but disabled on the source server/,
+ 'refuses an enabled target with a disabled source');
+
+# The refusal happens before any modification: the target still starts
+# on its own timeline.
+$node_a->start;
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'target untouched by the refused rewind');
+$node_a->stop('fast');
+
+# Scenario 2: disabling the target too makes the control files match,
+# but checksums were still enabled at the point of divergence, which is
+# the state replay on the rewound server would resume with.
+$node_a->checksum_disable_offline;
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums were enabled at the point of divergence but are disabled on the source server/,
+ 'refuses when checksums were enabled at the point of divergence');
+
+# Scenario 3: fresh pair without checksums; offline enable on the
+# source only. Allowed with a warning, and the rewound server keeps
+# checksums disabled.
+my $node_c = PostgreSQL::Test::Cluster->new('node_c');
+$node_c->init(allows_streaming => 1, no_data_checksums => 1);
+$node_c->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_c->start;
+$node_c->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_c->backup('backup');
+my $node_d = PostgreSQL::Test::Cluster->new('node_d');
+$node_d->init_from_backup($node_c, 'backup', has_streaming => 1);
+$node_d->start;
+$node_c->wait_for_catchup($node_d);
+
+$node_d->promote;
+$node_d->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_c->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_c->stop('fast');
+
+$node_d->stop;
+$node_d->checksum_enable_offline;
+$node_d->start;
+test_checksum_state($node_d, 'on');
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_c->data_dir,
+ '--source-server' => $node_d->connstr('postgres'),
+ ]);
+like(
+ $stderr,
+ qr/data checksums are disabled on the target server but enabled on the source server/,
+ 'warns for a disabled target with an enabled source');
+like($stderr, qr/Done!/, 'rewind completed despite the warning');
+
+# The rewound server follows D and keeps its own state.
+$node_c->append_conf('postgresql.conf', 'port = ' . $node_c->port);
+$node_c->enable_streaming($node_d);
+$node_c->set_standby_mode;
+$node_c->start;
+$node_d->wait_for_catchup($node_c);
+is($node_c->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+test_checksum_state($node_c, 'off');
+
+$node_c->stop;
+$node_d->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/023_rewind_standby_target.pl b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
new file mode 100644
index 00000000000..3f5a3c6be8b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
@@ -0,0 +1,141 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_rewind compares the data checksum state at the point of
+# divergence against the target as well as the source.
+#
+# A standby's checkpoints are written by its upstream primary, so a standby
+# whose checksums were disabled offline still has "on" checkpoints in its
+# WAL. Rewinding it onto a promoted sibling passes every control file
+# check (target off + source on only warns), yet recovery after the rewind
+# adopts the divergence checkpoint's state and the rewound server would
+# verify checksums over the pages it wrote while it was locally "off".
+# pg_rewind must refuse; after an offline enable on the target, which
+# rewrites all of its pages, the rewind goes through.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t0 AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+
+my $lagging = PostgreSQL::Test::Cluster->new('lagging');
+$lagging->init_from_backup($primary, 'backup', has_streaming => 1);
+$lagging->start;
+
+my $failover = PostgreSQL::Test::Cluster->new('failover');
+$failover->init_from_backup($primary, 'backup', has_streaming => 1);
+$failover->start;
+
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+
+test_checksum_state($primary, 'on');
+test_checksum_state($lagging, 'on');
+test_checksum_state($failover, 'on');
+
+# Offline-disable checksums on the lagging standby only. This is a
+# divergence 012_offline_standby.pl declares survivable: the standby keeps
+# its own state, warns, and stays readable.
+$lagging->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $lagging->data_dir);
+$lagging->start;
+test_checksum_state($lagging, 'off');
+
+# Give the lagging standby plenty of pages to write while it is locally
+# "off", and force a restartpoint so they reach disk without checksums.
+# The other standby replays the same WAL with checksums on.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+$lagging->safe_psql('postgres', "CHECKPOINT;");
+
+# Take the "failover" standby out of the picture. It is promoted from
+# this position later, so everything replayed from here on is the part of
+# the WAL that diverges and that pg_rewind will roll back. t1 is already
+# replayed everywhere, so its blocks are *not* rolled back.
+$failover->stop;
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$primary->wait_for_catchup($lagging);
+
+# Failover: promote the other standby, which forks the timeline behind the
+# replay position of the lagging standby.
+$primary->stop('fast');
+$failover->start;
+$failover->promote;
+$failover->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$failover->safe_psql('postgres', "CHECKPOINT;");
+
+# Re-attaching the lagging standby with pg_rewind must be refused: the
+# divergence checkpoint says "on" while the target is "off", so the rewound
+# server would verify checksums over the blocks it keeps.
+$lagging->stop;
+
+command_fails_like(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ],
+ qr/data checksums were enabled at the point of divergence but are disabled on the target server/,
+ 'pg_rewind refuses a target whose divergence checkpoint has checksums');
+
+# Enabling checksums offline on the target rewrites all of its pages, after
+# which the rewind is safe.
+system_or_bail('pg_checksums', '--enable', '--pgdata', $lagging->data_dir);
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ]);
+like($stderr, qr/Done!/, 'pg_rewind completes after the offline enable');
+
+$lagging->append_conf('postgresql.conf', 'port = ' . $lagging->port);
+$lagging->enable_streaming($failover);
+$lagging->set_standby_mode;
+$lagging->start;
+
+my ($rc, $out, $err) =
+ $lagging->psql('postgres',
+ "SELECT setting FROM pg_settings " . "WHERE name = 'data_checksums';");
+is($rc, 0, 'rewound standby accepts connections') or diag("stderr: $err");
+is($out, 'on', 'rewound standby resumes with checksums on');
+
+($rc, $out, $err) = $lagging->psql('postgres', "SELECT count(*) FROM t1;");
+is($rc, 0, 'blocks kept from the target stay readable')
+ or diag("stderr: $err");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($lagging->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the rewound standby');
+
+$lagging->stop('immediate');
+$failover->stop;
+done_testing();
--
2.55.0
[application/octet-stream] v9-0001-Do-not-adopt-data-checksum-state-from-another-nod.patch (127.9K, ../../CAN4CZFNdfb-yFRW7Sh3FkZ3Qc91Lz-JDQN0b-070-6G35g5Ssg@mail.gmail.com/3-v9-0001-Do-not-adopt-data-checksum-state-from-another-nod.patch)
download | inline diff:
From f4987dd3843d129b41b0589e600db1c3d1d57cd1 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:02 +0000
Subject: [PATCH v9 1/4] Do not adopt data checksum state from another node
during replay
Offline pg_checksums changes are local to one node, but replay adopted
the data checksum state carried by checkpoint records unconditionally.
After an offline enable on the primary, a standby whose pages were
never rewritten started verifying checksums it does not have; after an
offline change on a standby, the next replayed checkpoint silently
reverted it.
To fix, make the control file track this node's state alone, and have
replay cross-check the replayed state against it instead of adopting
it, warning once per divergent value and reporting when the states
agree again. pg_control gains a watermark, normally the end LSN of
the newest XLOG2_CHECKSUMS record the node has written or applied, so
that replay can skip transition records whose effect the control file
already contains, and a flag marking a state last written by
pg_checksums, which recovery must never overwrite with a replayed one.
Since the control file may only claim "on" once every page on disk
carries a checksum, persisting the state is tied to flushes:
XLOG2_CHECKSUMS replay persists every state but "on" immediately and
leaves "on" to the next restartpoint, checkpoints and restartpoints
only persist a state their flush ran under from beginning to end, and
transitions publish their state under the new
DataChecksumTransitionLock so that the states carried by WAL records
match their WAL order. The comments in xlog.c spell out the
individual rules.
Recovery from a base backup is the exception to not adopting: its
control file was copied at an arbitrary moment, so the state carried
by the starting checkpoint is the one the WAL from there on was
written under. pg_rewind keeps the target's own state and clamps the
watermark to the divergence point, since a watermark set on the
target's abandoned fork could numerically cover transition records the
source wrote after the divergence.
Bump PG_CONTROL_VERSION.
Document the offline procedure for replication setups, the lockstep
one: stop all nodes, run pg_checksums on each of them, and only then
restart. An offline change writes no WAL and has no ordering against
WAL a node has not replayed yet, so a node must not stop before
replaying all WAL of its upstream, or the change is overridden on
restart.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Author: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
---
doc/src/sgml/ref/pg_checksums.sgml | 51 +-
doc/src/sgml/wal.sgml | 21 +
src/backend/access/transam/xlog.c | 547 ++++++++++++++++--
.../utils/activity/wait_event_names.txt | 1 +
src/bin/pg_checksums/pg_checksums.c | 10 +
src/bin/pg_controldata/pg_controldata.c | 4 +
src/bin/pg_resetwal/pg_resetwal.c | 7 +
src/bin/pg_rewind/pg_rewind.c | 35 +-
src/bin/pg_upgrade/controldata.c | 2 +-
src/include/catalog/pg_control.h | 28 +-
src/include/storage/lwlocklist.h | 1 +
src/test/modules/test_checksums/Makefile | 2 +-
src/test/modules/test_checksums/meson.build | 10 +
.../test_checksums/t/012_offline_standby.pl | 289 +++++++++
.../modules/test_checksums/t/013_rewind.pl | 201 +++++++
.../modules/test_checksums/t/014_lockstep.pl | 182 ++++++
.../t/015_standby_crash_after_disable.pl | 125 ++++
.../t/016_promote_enable_crash.pl | 134 +++++
.../test_checksums/t/017_restartpoint_race.pl | 138 +++++
.../t/018_enable_crash_windows.pl | 496 ++++++++++++++++
.../t/019_standby_shutdown_catchup.pl | 126 ++++
.../t/020_cascade_divergence.pl | 115 ++++
.../t/021_rewind_divergent_transitions.pl | 193 ++++++
23 files changed, 2639 insertions(+), 79 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/012_offline_standby.pl
create mode 100644 src/test/modules/test_checksums/t/013_rewind.pl
create mode 100644 src/test/modules/test_checksums/t/014_lockstep.pl
create mode 100644 src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
create mode 100644 src/test/modules/test_checksums/t/016_promote_enable_crash.pl
create mode 100644 src/test/modules/test_checksums/t/017_restartpoint_race.pl
create mode 100644 src/test/modules/test_checksums/t/018_enable_crash_windows.pl
create mode 100644 src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
create mode 100644 src/test/modules/test_checksums/t/020_cascade_divergence.pl
create mode 100644 src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index ae66fad3f0f..6f332ade462 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -242,15 +242,48 @@ PostgreSQL documentation
data directory must not be started or else data loss may occur.
</para>
<para>
- When using a replication setup with tools which perform direct copies
- of relation file blocks (for example <xref linkend="app-pgrewind"/>),
- enabling or disabling checksums can lead to page corruptions in the
- shape of incorrect checksums if the operation is not done consistently
- across all nodes. When enabling or disabling checksums in a replication
- setup, it is thus recommended to stop all the clusters before switching
- them all consistently. Destroying all standbys, performing the operation
- on the primary and finally recreating the standbys from scratch is also
- safe.
+ Enabling or disabling checksums with
+ <application>pg_checksums</application> changes only the local data
+ directory; the new state does not propagate over replication. In a
+ replication setup the same change must be applied to every node: stop all
+ nodes, run <application>pg_checksums</application> on each of them, and
+ only then restart them. Before stopping a standby, make sure it has
+ replayed all WAL of its upstream node, for example by stopping the
+ primary first and comparing
+ <function>pg_last_wal_replay_lsn()</function> with
+ <function>pg_last_wal_receive_lsn()</function> on the standby. Tools
+ that copy relation file blocks directly between nodes, such as
+ <xref linkend="app-pgrewind"/>, likewise require both nodes to be in
+ the same data checksum state.
+ </para>
+ <para>
+ The replay requirement exists because an offline change is recorded
+ only in the control file and has no defined ordering against WAL the
+ node has not replayed yet; see
+ <xref linkend="checksums-offline-enable-disable"/>. A node stopped
+ before replaying an online checksum state change applies that change
+ when it is restarted, overriding the offline change, and the states of
+ the nodes silently diverge until a later checkpoint record triggers
+ the warning described below. Because of this it is best not to mix
+ the two mechanisms: change the state of a replication setup either
+ with the offline procedure above or with an online transition, and
+ make sure the previous change has reached every node before starting
+ the next one.
+ </para>
+ <para>
+ If the change is applied inconsistently, each node keeps its own state,
+ and a standby logs a warning when the state recorded in the replayed WAL
+ differs from its own. The same warning can appear transiently while a
+ standby catches up over WAL written before a consistent change; it stops
+ once a checkpoint record carrying the new state has been replayed.
+ </para>
+ <para>
+ A standby whose data directory was never checksummed must not have
+ checksums enabled by catching up this way. Converge the cluster by
+ running <application>pg_checksums</application> on it while stopped, by
+ enabling checksums online, or by recreating it from a base backup. Note
+ that an online enable only starts from a primary whose checksums are off,
+ so if they are already enabled there, disable them online first.
</para>
<para>
If <application>pg_checksums</application> is aborted or killed while
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index ec62d17fbcc..d88a832c2a8 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -317,6 +317,27 @@
verify checksums, on an offline cluster.
</para>
+ <para>
+ An offline change provides durability differently from an
+ <link linkend="checksums-online-enable-disable">online change</link>.
+ An online transition is WAL-logged: it is ordered against all other
+ WAL records, it is replayed after a crash, and it propagates to
+ standbys. An offline change is recorded only in the cluster's
+ control file: it writes no WAL, it is invisible to replication, and
+ it has no defined ordering against WAL the node has not replayed
+ yet. When a node later replays WAL that contains an online checksum
+ state change, that change takes effect on the node even if it was
+ written before the offline change was made.
+ </para>
+
+ <para>
+ An offline change only affects the data directory it is run on; the
+ new state does not propagate over replication. In a replication setup
+ the same change must be applied to all nodes while all of them are
+ stopped, as described in <xref linkend="app-pgchecksums"/>. A standby
+ whose state diverges logs a warning but keeps its local setting.
+ </para>
+
</sect2>
<sect2 id="checksums-online-enable-disable" xreflabel="Online Enabling of Checksums">
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index de4c96e135f..7487ca509f3 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -556,9 +556,17 @@ typedef struct XLogCtlData
*/
XLogRecPtr lastFpwDisableRecPtr;
- /* last data_checksum_version we've seen */
+ /* current data checksum state of this node */
uint32 data_checksum_version;
+ /*
+ * In-memory copies of ControlFile->data_checksum_lsn and
+ * ControlFile->data_checksum_is_local, see there. Updated together with
+ * data_checksum_version under info_lck.
+ */
+ XLogRecPtr data_checksum_lsn;
+ bool data_checksum_is_local;
+
slock_t info_lck; /* locks shared variables shown above */
/*
@@ -690,6 +698,14 @@ static ChecksumStateType LocalDataChecksumState = 0;
*/
int data_checksums = 0;
+/*
+ * Whether replay of the next checkpoint-family record must adopt the data
+ * checksum state it carries. Set when recovery starts from a base backup,
+ * where the state at the redo point takes precedence over the control file
+ * copied with the backup later.
+ */
+static bool adoptChecksumStateFromNextCheckpoint = false;
+
/* For WALInsertLockAcquire/Release functions */
static int MyLockNo = 0;
static bool holdingAllLocks = false;
@@ -730,6 +746,8 @@ static void ValidateXLOGDirectoryStructure(void);
static void CleanupBackupHistory(void);
static void UpdateMinRecoveryPoint(XLogRecPtr lsn, bool force);
static bool PerformRecoveryXLogAction(void);
+static void CheckReplayedDataChecksumState(uint32 replayed_version);
+static void AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn);
static void InitControlFile(uint64 sysidentifier, uint32 data_checksum_version);
static void WriteControlFile(void);
static void ReadControlFile(void);
@@ -757,7 +775,7 @@ static void WALInsertLockAcquireExclusive(void);
static void WALInsertLockRelease(void);
static void WALInsertLockUpdateInsertingAt(XLogRecPtr insertingAt);
-static void XLogChecksums(uint32 new_type);
+static XLogRecPtr XLogChecksums(uint32 new_type);
/*
* Insert an XLOG record represented by an already-constructed chain of data
@@ -4774,6 +4792,7 @@ void
SetDataChecksumsOnInProgress(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
/*
* The state transition is performed in a critical section with
@@ -4782,14 +4801,12 @@ SetDataChecksumsOnInProgress(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4827,6 +4844,8 @@ void
SetDataChecksumsOn(void)
{
uint64 barrier;
+ bool persist;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4847,23 +4866,11 @@ SetDataChecksumsOn(void)
SpinLockRelease(&XLogCtl->info_lck);
INJECTION_POINT("datachecksums-enable-checksums-delay", NULL);
+ INJECTION_POINT_LOAD("datachecksums-on-before-publish");
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_VERSION);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- SpinLockRelease(&XLogCtl->info_lck);
-
- /*
- * Update the controlfile before waiting since if we have an immediate
- * shutdown while waiting we want to come back up with checksums enabled.
- */
- LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- UpdateControlFile();
- LWLockRelease(ControlFileLock);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_VERSION);
barrier = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_CHECKSUM_ON);
@@ -4873,6 +4880,39 @@ SetDataChecksumsOn(void)
INJECTION_POINT("datachecksums-on-before-checkpoint", NULL);
RequestCheckpoint(CHECKPOINT_FORCE | CHECKPOINT_WAIT | CHECKPOINT_FAST);
+
+ INJECTION_POINT("datachecksums-on-after-checkpoint", NULL);
+
+ /*
+ * Persist "on" only now that the checkpoint has flushed the pages the
+ * transition rewrote. Crash recovery initializes verification from the
+ * control file but resumes from a checkpoint that can predate the
+ * transition, so an "on" persisted earlier would have replay verify pages
+ * whose rewrite never reached disk: the pages the worker found in shared
+ * buffers are not written back by its ring buffer.
+ *
+ * The checkpoint above normally persists the state itself; this covers the
+ * case where it started before the record written above and left the field
+ * alone. Crashing before this point is safe, as replay then re-establishes
+ * "on" from the full page images of the rewrite. Skip the write if the
+ * state moved on meanwhile, since whatever moved it persists its own.
+ * Compare the watermark rather than the state: a state comparison could
+ * not tell our transition from a later round trip back to "on".
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ persist = (XLogCtl->data_checksum_lsn == recptr);
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (persist)
+ {
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+ }
+
WaitForProcSignalBarrier(barrier);
}
@@ -4893,6 +4933,7 @@ void
SetDataChecksumsOff(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4918,14 +4959,12 @@ SetDataChecksumsOff(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4956,14 +4995,12 @@ SetDataChecksumsOff(void)
/* Ensure that we don't incur a checkpoint during disabling checksums */
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -5001,6 +5038,131 @@ SetLocalDataChecksumState(uint32 data_checksum_version)
data_checksums = data_checksum_version;
}
+/*
+ * Cross-check the data checksum state carried by a replayed checkpoint record
+ * against the state of this node.
+ *
+ * The state in the record belongs to whichever node wrote the WAL, and must
+ * not be adopted: an offline state change made with pg_checksums on one node
+ * of a replication set generates no WAL, and adopting would leak it into the
+ * other nodes through replay. Backup label recovery is the one exception,
+ * see AdoptReplayedDataChecksumState(). XLOG_CHECKPOINT_ONLINE needs no call
+ * here, as an online checkpoint's state already traveled in the preceding
+ * XLOG_CHECKPOINT_REDO record.
+ *
+ * Only archive recovery can see a lasting mismatch, as only there can the WAL
+ * and the control file come from different nodes or different times. In
+ * crash recovery a mismatch means replay resumed from a restartpoint
+ * predating an already-applied XLOG2_CHECKSUMS record, and replaying forward
+ * re-establishes the same state.
+ */
+static void
+CheckReplayedDataChecksumState(uint32 replayed_version)
+{
+ /*
+ * Warn once per remote value, so a lasting mismatch does not flood the
+ * log. Matching states re-arm the warning. Backend-local state is
+ * enough: replay only runs in the startup process, and restarting it
+ * re-arms as well.
+ */
+ static uint32 last_warned_version = PG_UINT32_MAX;
+ uint32 local_version;
+
+ if (!ArchiveRecoveryRequested)
+ return;
+
+ /*
+ * Re-replayed WAL below the consistency point was already cross-checked
+ * before minRecoveryPoint was last persisted, and the persisted state can
+ * legitimately be newer than what checkpoint records there carry:
+ * XLOG2_CHECKSUMS replay persists most states ahead of the restartpoint
+ * horizon. In particular the checkpoint record recovery restarts from is
+ * such a re-replay.
+ */
+ if (!reachedConsistency)
+ return;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ local_version = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (replayed_version == local_version)
+ {
+ /*
+ * Report convergence if this process warned before. Nothing else
+ * tells the operator that running pg_checksums on the other nodes, or
+ * a rebuild, took effect.
+ */
+ if (last_warned_version != PG_UINT32_MAX)
+ ereport(LOG,
+ errmsg("data checksum state \"%s\" of this node now agrees with the replayed WAL",
+ get_checksum_state_string(local_version)));
+
+ last_warned_version = PG_UINT32_MAX;
+ return;
+ }
+
+ /* the nodes legitimately differ while an online transition runs */
+ if (replayed_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ replayed_version == PG_DATA_CHECKSUM_INPROGRESS_OFF ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ return;
+
+ if (replayed_version == last_warned_version)
+ return;
+ last_warned_version = replayed_version;
+
+ ereport(WARNING,
+ errmsg("data checksum state \"%s\" of this node does not match the state \"%s\" in the replayed WAL",
+ get_checksum_state_string(local_version),
+ get_checksum_state_string(replayed_version)),
+ errdetail("The data checksum state was most likely changed with pg_checksums on another node."),
+ errhint("Apply the same change with pg_checksums on the primary and all standby servers, or rebuild this server from a base backup."));
+}
+
+/*
+ * Adopt the data checksum state found at the redo point of backup label
+ * recovery. Persist it immediately so that a crash before the first
+ * restartpoint does not resurrect the state copied with the backup; a crash
+ * at this point restarts from the same redo point, so the control file does
+ * not run ahead of the replay position. If the value is unchanged the
+ * control file already carries it, so both the barrier and the persist are
+ * skipped.
+ *
+ * lsn is the location of the checkpoint-family record the state was taken
+ * from and becomes the new watermark: the adopted state covers everything
+ * below the redo point, which replay never revisits.
+ */
+static void
+AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn)
+{
+ bool changed = false;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (XLogCtl->data_checksum_version != new_version)
+ {
+ XLogCtl->data_checksum_version = new_version;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(new_version);
+ changed = true;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (!changed)
+ return;
+
+ EmitAndWaitDataChecksumsBarrier(new_version);
+
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = new_version;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+}
+
/* guc hook */
const char *
show_data_checksums(void)
@@ -5454,6 +5616,8 @@ XLOGShmemInit(void *arg)
/* Use the checksum info from control file */
XLogCtl->data_checksum_version = ControlFile->data_checksum_version;
+ XLogCtl->data_checksum_lsn = ControlFile->data_checksum_lsn;
+ XLogCtl->data_checksum_is_local = ControlFile->data_checksum_is_local;
SetLocalDataChecksumState(XLogCtl->data_checksum_version);
SpinLockInit(&XLogCtl->Insert.insertpos_lck);
@@ -6031,6 +6195,52 @@ StartupXLOG(void)
SetCommitTsLimit(checkPoint.oldestCommitTsXid,
checkPoint.newestCommitTsXid);
+ /*
+ * When recovery starts from a base backup, the control file was copied at
+ * an arbitrary moment and its data checksum state may differ from the
+ * state at the redo point, which is what the WAL from there on was
+ * written under. Adopt the state of the starting checkpoint: a shutdown
+ * checkpoint is not replayed, so take it from the record read above; the
+ * redo point of an online checkpoint is its CHECKPOINT_REDO record, so
+ * let the replay of that record adopt it. Check backupStartPoint in
+ * addition to the label: on a crash restart during backup recovery the
+ * label file is already renamed away, but the start point persists until
+ * the backup end record.
+ *
+ * Not for a base backup taken from a standby, though. Its starting
+ * checkpoint is the standby's last restartpoint, a record written by the
+ * upstream primary, whose state is not the one the copied files were
+ * written under. The copied control file is already correct: a standby
+ * persists its state only at restartpoint horizons and never claims
+ * more than what reached disk. Such backups are recognized by
+ * backupEndPoint together with backupEndRequired; backupEndPoint is only
+ * set for "BACKUP FROM: standby" labels and persists across a crash
+ * restart. pg_rewind writes a standby label as well, but no
+ * backupEndPoint, and its recovery keeps adopting: the control file it
+ * installs carries the target's own checksum state, which can lag the
+ * redo point of the last common checkpoint the same way a restartpoint
+ * horizon can.
+ *
+ * Never adopt over a state the control file's watermark or local flag
+ * marks as newer than the starting checkpoint. A pg_checksums change is
+ * local to the node and generates no WAL, so nothing in the replayed WAL
+ * could ever restore it once overwritten; and a watermark above the redo
+ * point means the control file already contains the effect of every
+ * transition record up to there.
+ */
+ if ((haveBackupLabel || XLogRecPtrIsValid(ControlFile->backupStartPoint)) &&
+ !(XLogRecPtrIsValid(ControlFile->backupEndPoint) &&
+ ControlFile->backupEndRequired) &&
+ !ControlFile->data_checksum_is_local &&
+ checkPoint.redo > ControlFile->data_checksum_lsn)
+ {
+ if (wasShutdown)
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ checkPoint.redo);
+ else
+ adoptChecksumStateFromNextCheckpoint = true;
+ }
+
/*
* Clear out any old relcache cache files. This is *necessary* if we do
* any WAL replay, since that would probably result in the cache files
@@ -6633,11 +6843,7 @@ StartupXLOG(void)
if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
ereport(WARNING,
@@ -6654,11 +6860,7 @@ StartupXLOG(void)
else if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
}
@@ -6814,6 +7016,23 @@ static bool
PerformRecoveryXLogAction(void)
{
bool promoted = false;
+ bool flushForChecksums;
+ uint32 checksum_state;
+
+ /*
+ * The end-of-recovery record persists the data checksum state without
+ * flushing the buffer pool, but the control file may only claim "on" once
+ * every page on disk carries a checksum. If replay entered that state
+ * without a restartpoint following it, the pages rewritten by the
+ * transition are still only in the buffer pool, so take the full
+ * checkpoint below instead of the lightweight record.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ flushForChecksums = (checksum_state == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile->data_checksum_version != checksum_state);
/*
* Perform a checkpoint to update all our recovery activity to disk.
@@ -6829,7 +7048,7 @@ PerformRecoveryXLogAction(void)
* fully out of recovery mode and already accepting queries.
*/
if (ArchiveRecoveryRequested && IsUnderPostmaster &&
- PromoteIsTriggered())
+ PromoteIsTriggered() && !flushForChecksums)
{
promoted = true;
@@ -7436,6 +7655,7 @@ CreateCheckPoint(int flags)
uint32 freespace;
XLogRecPtr PriorRedoPtr;
XLogRecPtr last_important_lsn;
+ XLogRecPtr checksumLsn;
VirtualTransactionId *vxids;
int nvxids;
int oldXLogAllowed = 0;
@@ -7548,11 +7768,14 @@ CreateCheckPoint(int flags)
checkPoint.wal_level = wal_level;
/*
- * Get the current data_checksum_version value from xlogctl, valid at the
- * time of the checkpoint.
+ * Get the current data_checksum_version value from xlogctl. This is
+ * final only for a shutdown checkpoint, where no concurrent transition
+ * is possible; an online checkpoint resamples it together with the redo
+ * record below.
*/
SpinLockAcquire(&XLogCtl->info_lck);
checkPoint.dataChecksumState = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
if (shutdown)
@@ -7609,10 +7832,21 @@ CreateCheckPoint(int flags)
{
xl_checkpoint_redo redo_rec;
+ /*
+ * Sample the data checksum state and insert the redo record under
+ * DataChecksumTransitionLock, so that a concurrent transition cannot
+ * insert its XLOG2_CHECKSUMS record between the sampling and the
+ * insertion below. Without this, the redo record could follow the
+ * transition record in WAL while carrying the pre-transition state,
+ * and recovery resuming here would never learn about the transition.
+ * See XLogChecksums().
+ */
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
WALInsertLockAcquire();
redo_rec.wal_level = wal_level;
SpinLockAcquire(&XLogCtl->info_lck);
redo_rec.data_checksum_version = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
WALInsertLockRelease();
@@ -7620,6 +7854,14 @@ CreateCheckPoint(int flags)
XLogBeginInsert();
XLogRegisterData(&redo_rec, sizeof(xl_checkpoint_redo));
(void) XLogInsert(RM_XLOG_ID, XLOG_CHECKPOINT_REDO);
+ LWLockRelease(DataChecksumTransitionLock);
+
+ /*
+ * The checkpoint record must carry the same state as the redo record
+ * just inserted: the sample taken before redo determination can be
+ * stale by now, and the pair would otherwise disagree.
+ */
+ checkPoint.dataChecksumState = redo_rec.data_checksum_version;
/*
* XLogInsertRecord will have updated XLogCtl->Insert.RedoRecPtr in
@@ -7823,6 +8065,40 @@ CreateCheckPoint(int flags)
ControlFile->minRecoveryPoint = InvalidXLogRecPtr;
ControlFile->minRecoveryPointTLI = 0;
+ /*
+ * Persist the data checksum state this node runs under. Only the
+ * top-level field tracks this node; ControlFile->checkPointCopy above is
+ * a historical record used to resume replay.
+ *
+ * checkPoint.dataChecksumState was sampled under
+ * DataChecksumTransitionLock together with the redo record, so it is the
+ * state in effect at the redo point. If it was "on",
+ * the XLOG2_CHECKSUMS record announcing that precedes the redo point and
+ * every page the transition rewrote was dirtied before it, so
+ * CheckPointGuts() has just written all of them out. Recording the state
+ * here is what keeps a finished transition from being resolved as
+ * interrupted when this checkpoint is the one crash recovery resumes
+ * from: replay never sees the record announcing it.
+ *
+ * Persist it only if the state did not change while the flush was in
+ * progress. If it changed in between, the pages written out straddle two
+ * states, and the newer one could claim checksums that pages already on
+ * disk do not carry; leave the field to the next checkpoint then, the
+ * transition itself has already persisted every state that is safe
+ * without a flush.
+ * Compare the watermark rather than the state: record positions are
+ * unique, so a full round trip back to the sampled state cannot alias,
+ * while its flushed pages straddle the intermediate states all the same.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksumLsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checkPoint.dataChecksumState;
+ ControlFile->data_checksum_lsn = checksumLsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
/*
* Persist unloggedLSN value. It's reset on crash recovery, so this goes
* unused on non-shutdown checkpoints, but seems useful to store it always
@@ -7967,9 +8243,11 @@ CreateEndOfRecoveryRecord(void)
ControlFile->minRecoveryPoint = recptr;
ControlFile->minRecoveryPointTLI = xlrec.ThisTimeLineID;
- /* start with the latest checksum version (as of the end of recovery) */
+ /* persist the data checksum state this node ended recovery with */
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
@@ -8174,6 +8452,9 @@ CreateRestartPoint(int flags)
XLogRecPtr endptr;
XLogSegNo _logSegNo;
TimestampTz xtime;
+ uint32 checksum_state;
+ XLogRecPtr checksum_lsn;
+ bool checksum_is_local;
/* Concurrent checkpoint/restartpoint cannot happen */
Assert(!IsUnderPostmaster || MyBackendType == B_CHECKPOINTER);
@@ -8220,8 +8501,45 @@ CreateRestartPoint(int flags)
UpdateMinRecoveryPoint(InvalidXLogRecPtr, true);
if (flags & CHECKPOINT_IS_SHUTDOWN)
{
+ bool catchUpChecksums;
+
+ /*
+ * There is no new restartpoint to persist the data checksum state
+ * with, but a cleanly stopped node should not leave the control
+ * file behind the state replay reached: pg_checksums and
+ * pg_rewind read it, and an in-progress state there makes them
+ * refuse to run. Catching it up needs the same guarantee a
+ * restartpoint gives: that every page on disk carries a checksum,
+ * so flush the buffer pool first. Only a transition to
+ * "on" that no restartpoint followed can get here; the other
+ * states are already persisted by XLOG2_CHECKSUMS replay. Replay
+ * has ended by now, so the state cannot change under us.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ checksum_is_local = XLogCtl->data_checksum_is_local;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ catchUpChecksums =
+ (checksum_lsn != ControlFile->data_checksum_lsn &&
+ XLogRecPtrIsValid(lastCheckPointRecPtr));
+
+ if (catchUpChecksums)
+ {
+ MemSet(&CheckpointStats, 0, sizeof(CheckpointStats));
+ CheckpointStats.ckpt_start_t = GetCurrentTimestamp();
+ CheckPointGuts(lastCheckPoint.redo, flags);
+ }
+
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
+ if (catchUpChecksums)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = checksum_is_local;
+ }
UpdateControlFile();
LWLockRelease(ControlFileLock);
}
@@ -8262,6 +8580,17 @@ CreateRestartPoint(int flags)
/* Update the process title */
update_checkpoint_display(flags, true, false);
+ /*
+ * Note the data checksum state the flush below starts under. Replay runs
+ * concurrently and can change the state while the flush is in progress,
+ * in which case the flush covers pages written under both states; see
+ * where the state is persisted further down.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
CheckPointGuts(lastCheckPoint.redo, flags);
/*
@@ -8321,8 +8650,26 @@ CreateRestartPoint(int flags)
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
}
- /* we shall start with the latest checksum version */
- ControlFile->data_checksum_version = lastCheckPoint.dataChecksumState;
+ /*
+ * Persist the data checksum state of this node. Not the state of the
+ * replayed checkpoint: that one belongs to the node that wrote it and
+ * may differ after an offline change on either side.
+ * ControlFile->checkPointCopy above keeps the replayed value on
+ * purpose, being a historical record used to resume replay rather
+ * than a tracker of node state.
+ *
+ * Persist only if the flush above ran under one state throughout; see
+ * CreateCheckPoint() for why, including why this compares the
+ * watermark and not the state.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksum_lsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
}
@@ -8763,9 +9110,21 @@ XLogReportParameters(void)
}
/*
- * Log the new state of checksums
+ * Log and publish the new state of checksums
+ *
+ * Inserting the record and publishing the new state must be atomic with
+ * respect to a checkpoint sampling the state for its XLOG_CHECKPOINT_REDO
+ * record: without that, a checkpoint could read the old state after the
+ * record is already in WAL and insert a redo record that both precedes the
+ * transition in WAL order and carries the pre-transition state. Recovery
+ * resuming from such a redo point would never replay the transition record
+ * and resolve the finished transition as interrupted.
+ * DataChecksumTransitionLock serializes the two; see CreateCheckPoint().
+ *
+ * Returns the end LSN of the inserted record, which the caller persists
+ * together with the new state as the data checksum watermark.
*/
-static void
+static XLogRecPtr
XLogChecksums(uint32 new_type)
{
xl_checksum_state xlrec;
@@ -8773,12 +9132,28 @@ XLogChecksums(uint32 new_type)
xlrec.new_checksum_state = new_type;
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
+
XLogBeginInsert();
XLogRegisterData((char *) &xlrec, sizeof(xl_checksum_state));
recptr = XLogInsert(RM_XLOG2_ID, XLOG2_CHECKSUMS);
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, recptr);
+
+ /* only loaded by SetDataChecksumsOn(), a no-op for the other callers */
+ INJECTION_POINT_CACHED("datachecksums-on-before-publish", NULL);
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ XLogCtl->data_checksum_version = new_type;
+ XLogCtl->data_checksum_lsn = recptr;
+ XLogCtl->data_checksum_is_local = false;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockRelease(DataChecksumTransitionLock);
+
XLogFlush(recptr);
+
+ return recptr;
}
/*
@@ -8966,11 +9341,19 @@ xlog_redo(XLogReaderState *record)
/* ControlFile->checkPointCopy always tracks the latest ckpt XID */
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->checkPointCopy.nextXid = checkPoint.nextXid;
- ControlFile->data_checksum_version = checkPoint.dataChecksumState;
UpdateControlFile();
LWLockRelease(ControlFileLock);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(checkPoint.dataChecksumState);
+
/*
* We should've already switched to the new TLI before replaying this
* record.
@@ -9206,19 +9589,17 @@ xlog_redo(XLogReaderState *record)
else if (info == XLOG_CHECKPOINT_REDO)
{
xl_checkpoint_redo redo_rec;
- bool new_state = false;
memcpy(&redo_rec, XLogRecGetData(record), sizeof(xl_checkpoint_redo));
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = redo_rec.data_checksum_version;
- SetLocalDataChecksumState(redo_rec.data_checksum_version);
- if (redo_rec.data_checksum_version != ControlFile->data_checksum_version)
- new_state = true;
- SpinLockRelease(&XLogCtl->info_lck);
-
- if (new_state)
- EmitAndWaitDataChecksumsBarrier(redo_rec.data_checksum_version);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(redo_rec.data_checksum_version,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(redo_rec.data_checksum_version);
}
else if (info == XLOG_LOGICAL_DECODING_STATUS_CHANGE)
{
@@ -9280,25 +9661,63 @@ xlog2_redo(XLogReaderState *record)
{
xl_checksum_state state;
XLogRecPtr lsn = record->EndRecPtr;
+ XLogRecPtr watermark;
memcpy(&state, XLogRecGetData(record), sizeof(xl_checksum_state));
+ SpinLockAcquire(&XLogCtl->info_lck);
+ watermark = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ /*
+ * Skip records this node has already applied. The control file
+ * carries the watermark, so this holds across restarts: recovery
+ * resuming below a record whose effect the control file already
+ * contains must not re-apply it, or it would revert a state change
+ * made with pg_checksums in between, which moves the state without
+ * writing any record of its own.
+ */
+ if (lsn <= watermark)
+ return;
+
/* advertise the location before the new state becomes visible */
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, lsn);
SpinLockAcquire(&XLogCtl->info_lck);
XLogCtl->data_checksum_version = state.new_checksum_state;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(state.new_checksum_state);
SpinLockRelease(&XLogCtl->info_lck);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = state.new_checksum_state;
+
+ /*
+ * Persist the new state, except when it is "on". Only "on" verifies
+ * checksums during reads, and between the last restartpoint and this
+ * record there may be pages on disk flushed under the old state; a
+ * crash-restart initializes verification from the control file and
+ * replay reads those pages back, so the control file may only say
+ * "on" once everything written under the transition has been flushed,
+ * as restartpoints and the end of recovery do.
+ * The opposite direction cannot wait for the restartpoint: once this
+ * record is replayed, evicted pages are written without checksums,
+ * and a control file still saying "on" would fail verification on
+ * exactly those pages after a crash.
+ */
+ if (state.new_checksum_state != PG_DATA_CHECKSUM_VERSION)
+ {
+ ControlFile->data_checksum_version = state.new_checksum_state;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ }
/*
* Update minRecoveryPoint to ensure that if recovery is aborted, we
* recover back up to this point before allowing hot standby again.
- * The new state is durable in pg_control while its location is only
- * tracked in shared memory; a standby becoming consistent below this
- * record would let base backups resume checksum verification with the
+ * The change location is only tracked in shared memory and is lost
+ * over a restart; a standby becoming consistent below this record
+ * would let base backups resume checksum verification with the
* location unknown. The local copies cannot be updated as long as
* crash recovery is happening and we expect all the WAL to be
* replayed.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 256b3a3c02e..77fa543f9d8 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -371,6 +371,7 @@ WaitLSN "Waiting to read or update shared Wait-for-LSN state."
LogicalDecodingControl "Waiting to read or update logical decoding status information."
DataChecksumsWorker "Waiting for data checksums worker."
AioWorkerControl "Waiting to update AIO worker information."
+DataChecksumTransition "Waiting for a data checksum state transition to be written to WAL."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 3b3ae23f1a6..20a99112838 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -648,6 +648,16 @@ main(int argc, char *argv[])
ControlFile->data_checksum_version =
(mode == PG_MODE_ENABLE) ? PG_DATA_CHECKSUM_VERSION : PG_DATA_CHECKSUM_OFF;
+ /*
+ * Mark the state as changed locally, without a WAL record. Recovery
+ * then knows the state is newer than anything the WAL carries and
+ * does not let a replayed checkpoint overwrite it. The watermark is
+ * left alone: any XLOG2_CHECKSUMS record this node had applied stays
+ * covered, and only records above it, written after this change, take
+ * effect again.
+ */
+ ControlFile->data_checksum_is_local = true;
+
if (do_sync)
{
pg_log_info("syncing data directory");
diff --git a/src/bin/pg_controldata/pg_controldata.c b/src/bin/pg_controldata/pg_controldata.c
index 6fc87ed114d..c363ce3dbb7 100644
--- a/src/bin/pg_controldata/pg_controldata.c
+++ b/src/bin/pg_controldata/pg_controldata.c
@@ -349,6 +349,10 @@ main(int argc, char *argv[])
(ControlFile->float8ByVal ? _("by value") : _("by reference")));
printf(_("Data page checksum version: %u\n"),
ControlFile->data_checksum_version);
+ printf(_("Data checksum watermark: %X/%08X\n"),
+ LSN_FORMAT_ARGS(ControlFile->data_checksum_lsn));
+ printf(_("Data checksum state is node-local: %s\n"),
+ (ControlFile->data_checksum_is_local ? _("yes") : _("no")));
printf(_("Default char data signedness: %s\n"),
(ControlFile->default_char_signedness ? _("signed") : _("unsigned")));
printf(_("Mock authentication nonce: %s\n"),
diff --git a/src/bin/pg_resetwal/pg_resetwal.c b/src/bin/pg_resetwal/pg_resetwal.c
index 1542a56ca4b..cdfb9898066 100644
--- a/src/bin/pg_resetwal/pg_resetwal.c
+++ b/src/bin/pg_resetwal/pg_resetwal.c
@@ -923,6 +923,13 @@ RewriteControlFile(void)
ControlFile.backupEndPoint = InvalidXLogRecPtr;
ControlFile.backupEndRequired = false;
+ /*
+ * The old WAL is gone and the new position may lie below the old
+ * watermark, which would make replay ignore future checksum transition
+ * records. The state itself is kept.
+ */
+ ControlFile.data_checksum_lsn = InvalidXLogRecPtr;
+
/*
* Force the defaults for max_* settings. The values don't really matter
* as long as wal_level='minimal'; the postmaster will reset these fields
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 2e86fd158d0..466f4223501 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -37,7 +37,8 @@ static void usage(const char *progname);
static void perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo);
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec);
static void createBackupLabel(XLogRecPtr startpoint, TimeLineID starttli,
XLogRecPtr checkpointloc);
@@ -531,7 +532,8 @@ main(int argc, char **argv)
* This is the point of no return. Once we start copying things, there is
* no turning back!
*/
- perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo);
+ perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo,
+ divergerec);
if (showprogress)
pg_log_info("syncing target data directory");
@@ -566,7 +568,8 @@ static void
perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo)
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec)
{
XLogRecPtr endrec;
TimeLineID endtli;
@@ -738,6 +741,32 @@ perform_rewind(filemap_t *filemap, rewind_source *source,
ControlFile_new.minRecoveryPoint = endrec;
ControlFile_new.minRecoveryPointTLI = endtli;
ControlFile_new.state = DB_IN_ARCHIVE_RECOVERY;
+
+ /*
+ * Keep the target's own data checksum state. Most of the data directory
+ * is still the target's: only blocks it changed since the divergence
+ * were copied from the source, so the source's state says nothing about
+ * the pages that stay. Replay from the last common checkpoint applies
+ * any WAL-logged transition the target has not seen (the watermark tells
+ * them apart), which converges the rewound server to the source's state
+ * whenever the WAL carries it.
+ */
+ ControlFile_new.data_checksum_version = ControlFile_target.data_checksum_version;
+ ControlFile_new.data_checksum_lsn = ControlFile_target.data_checksum_lsn;
+ ControlFile_new.data_checksum_is_local = ControlFile_target.data_checksum_is_local;
+
+ /*
+ * The watermark is only meaningful within the history the node replays.
+ * Records at or below the divergence point are common to both histories
+ * and stay covered, but a watermark above it was set by a transition
+ * record on the target's own abandoned fork: numerically it can cover
+ * transition records the source wrote after the divergence, and replay
+ * would skip them as already applied. Clamp it to the divergence point,
+ * so that every transition record on the source's history takes effect.
+ */
+ if (ControlFile_new.data_checksum_lsn > divergerec)
+ ControlFile_new.data_checksum_lsn = divergerec;
+
if (!dry_run)
update_controlfile(datadir_target, &ControlFile_new, do_sync);
}
diff --git a/src/bin/pg_upgrade/controldata.c b/src/bin/pg_upgrade/controldata.c
index b3bd4ccde83..de539479bdc 100644
--- a/src/bin/pg_upgrade/controldata.c
+++ b/src/bin/pg_upgrade/controldata.c
@@ -431,7 +431,7 @@ get_control_data(ClusterInfo *cluster)
cluster->controldata.date_is_int = strstr(p, "64-bit integers") != NULL;
got_date_is_int = true;
}
- else if ((p = strstr(bufin, "checksum")) != NULL)
+ else if ((p = strstr(bufin, "Data page checksum version:")) != NULL)
{
p = strchr(p, ':');
diff --git a/src/include/catalog/pg_control.h b/src/include/catalog/pg_control.h
index 7b5404460ec..6a83316827d 100644
--- a/src/include/catalog/pg_control.h
+++ b/src/include/catalog/pg_control.h
@@ -22,7 +22,7 @@
/* Version identifier for this pg_control format */
-#define PG_CONTROL_VERSION 1903
+#define PG_CONTROL_VERSION 1904
/* Nonce key length, see below */
#define MOCK_AUTH_NONCE_LEN 32
@@ -237,6 +237,32 @@ typedef struct ControlFileData
/* Current data checksums state */
uint32 data_checksum_version;
+ /*
+ * WAL position through which data checksum transitions are covered.
+ * Replay ignores XLOG2_CHECKSUMS records ending at or below this point:
+ * their effect is already contained in data_checksum_version, or an
+ * offline pg_checksums change made after they were first applied
+ * supersedes them. Ordinarily this is the end of the newest such record
+ * this node has written or applied, but a tool may store any position
+ * that covers the same set of records. InvalidXLogRecPtr if the node
+ * has never written or applied such a record.
+ *
+ * The comparison has no timeline context, so the value is only valid
+ * within the WAL history this node replays. A tool that moves the node
+ * to another history must clamp the watermark to the point where the
+ * histories fork, as pg_rewind does, or reset it, as pg_resetwal does.
+ */
+ XLogRecPtr data_checksum_lsn;
+
+ /*
+ * True when data_checksum_version was last set by pg_checksums rather
+ * than by a WAL-logged transition. Such a state is local to this node
+ * and newer than anything the WAL carries, so recovery must not replace
+ * it with a state taken from a checkpoint record. Cleared by the next
+ * WAL-logged transition.
+ */
+ bool data_checksum_is_local;
+
/*
* True if the default signedness of char is "signed" on a platform where
* the cluster is initialized.
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index d7eb648bd27..8d858be9927 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -89,6 +89,7 @@ PG_LWLOCK(54, WaitLSN)
PG_LWLOCK(55, LogicalDecodingControl)
PG_LWLOCK(56, DataChecksumsWorker)
PG_LWLOCK(57, AioWorkerControl)
+PG_LWLOCK(58, DataChecksumTransition)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/test/modules/test_checksums/Makefile b/src/test/modules/test_checksums/Makefile
index 71455cd5577..80f54bf6d8c 100644
--- a/src/test/modules/test_checksums/Makefile
+++ b/src/test/modules/test_checksums/Makefile
@@ -9,7 +9,7 @@
#
#-------------------------------------------------------------------------
-EXTRA_INSTALL = src/test/modules/injection_points
+EXTRA_INSTALL = contrib/pg_buffercache src/test/modules/injection_points
export enable_injection_points
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index fb7129d796f..e5d38fafb7d 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -35,6 +35,16 @@ tests += {
't/009_fpi.pl',
't/010_backup_straddle.pl',
't/011_standby_straddle.pl',
+ 't/012_offline_standby.pl',
+ 't/013_rewind.pl',
+ 't/014_lockstep.pl',
+ 't/015_standby_crash_after_disable.pl',
+ 't/016_promote_enable_crash.pl',
+ 't/017_restartpoint_race.pl',
+ 't/018_enable_crash_windows.pl',
+ 't/019_standby_shutdown_catchup.pl',
+ 't/020_cascade_divergence.pl',
+ 't/021_rewind_divergent_transitions.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
new file mode 100644
index 00000000000..241a1282ae9
--- /dev/null
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -0,0 +1,289 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Offline checksum changes with pg_checksums are local to one node. A
+# standby must neither adopt the state of the primary from replayed
+# checkpoint records, nor lose its own offline change to them.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', 'autovacuum = off');
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Scenario 1: enable offline on the primary only. The standby must
+# stay off, warn about the mismatch, and remain readable.
+$standby->stop;
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+test_checksum_state($primary, 'on');
+test_checksum_state($standby, 'off');
+
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after offline enable on the primary');
+
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+my @warnings = $log =~ /(does not match the state)/g;
+is(scalar(@warnings), 1, 'mismatch warned once per remote value');
+
+# Matching states re-arm the warning: undo the divergence on the primary,
+# then diverge again to the same value, all without restarting the standby.
+$primary->stop;
+$primary->checksum_disable_offline;
+$logstart = -s $standby->logfile;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no warning while the states match again');
+like($log, qr/now agrees with the replayed WAL/,
+ 'convergence is reported once the states match again');
+
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+test_checksum_state($standby, 'off');
+
+# The local state survives both clean and immediate restarts.
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+# Still in scenario 1's divergence: take a base backup *from the
+# standby*. A backup taken on a standby uses the last restartpoint as
+# its starting checkpoint (do_pg_backup_start()), so the record at the
+# redo point was written by the upstream primary and carries the
+# primary's state. Recovery from such a backup must not adopt that
+# state: the files were copied from the standby, and the new node
+# would come up verifying checksums its files do not have.
+
+# Make sure the standby has written pages under its own "off" state, so
+# its files really do lack checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,50000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', "SELECT count(*) FROM t2;");
+
+$standby->backup('from_standby');
+my $newnode = PostgreSQL::Test::Cluster->new('newnode');
+$newnode->init_from_backup($standby, 'from_standby');
+
+# Stream from the primary, not from the standby the files were copied
+# from, so the new node replays the "on" primary's records.
+$newnode->enable_streaming($primary);
+$newnode->start;
+
+# The new node's files all came from a cluster running with checksums
+# off. Anything but "off" here means it adopted the primary's state
+# through the checkpoint record at the redo point.
+my ($rc, $stdout, $stderr) = $newnode->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'the node copied from the standby accepts connections')
+ or diag("stderr: $stderr");
+is($stdout, 'off', 'backup of an "off" standby comes up with checksums off');
+
+# A checkpoint record streamed from the "on" primary must not flip the
+# state either.
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($newnode);
+test_checksum_state($newnode, 'off');
+
+# And it must be able to read the pages the standby wrote without
+# checksums.
+(undef, $stdout, $stderr) =
+ $newnode->psql('postgres', "SELECT count(*) FROM t2;");
+is($stdout, '50000', 'pages copied from the standby are readable')
+ or diag("stderr: $stderr");
+
+my $newnode_log = PostgreSQL::Test::Utils::slurp_file($newnode->logfile);
+unlike(
+ $newnode_log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the node copied from the standby');
+
+$newnode->stop('immediate');
+
+# Converge the cluster: enable offline on the standby too.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+$primary->wait_for_catchup($standby);
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after converging');
+
+# Scenario 2: disable offline on the standby only. The replayed
+# checkpoint records of the still-enabled primary must not override it.
+$standby->stop;
+$standby->checksum_disable_offline;
+$standby->start;
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'off');
+
+# Restartpoints must persist the local state, not the replayed copy.
+$standby->safe_psql('postgres', "CHECKPOINT;");
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable with checksums disabled locally');
+
+# Scenario 3: crash-restart right after an online transition, before the
+# next restartpoint. Replay then resumes from an older restartpoint whose
+# checkpoint records still carry the pre-transition state. Those must
+# still match the state seeded from the control file, and the transition
+# itself must be re-established by re-replaying the XLOG2_CHECKSUMS record,
+# without a spurious mismatch warning along the way.
+
+# Converge first: bring the standby back to "on" offline.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+disable_data_checksums($primary, wait => 'off');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+# Crash-restart the standby immediately, before any restartpoint has had a
+# chance to persist the new state to its control file.
+$logstart = -s $standby->logfile;
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001',
+ 'standby readable after crash-restart across an online transition');
+
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no spurious mismatch warning after crash-restart across an online transition'
+);
+
+# Scenario 4: a standby stopped while replaying an interrupted online
+# transition keeps the interrupted state in its own control file. A
+# primary is never caught this way, as its checksums launcher resolves
+# inprogress-on back to off from its exit cleanup. A standby has no
+# launcher; it carries forward whatever the last replayed record left
+# it in.
+
+# Block an online enable on the primary at inprogress-on with a
+# blocking temp table, same trick as in 004_offline.pl.
+my $bsession = $primary->background_psql('postgres');
+$bsession->query_safe('CREATE TEMPORARY TABLE tt (a integer);');
+enable_data_checksums($primary, wait => 'inprogress-on');
+
+# The standby picks up the in-progress state from the XLOG2_CHECKSUMS record.
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# Stop the standby cleanly; its restartpoint persists inprogress-on to
+# its own control file, since nothing on a standby resolves it away.
+$standby->stop;
+$standby->start;
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# The primary still sits at inprogress-on, so a base backup taken now
+# copies a control file and a redo point that both carry the
+# in-progress state; replay of the XLOG2_CHECKSUMS records completes
+# the transition on the new standby.
+$primary->backup('inprogress_backup');
+my $standby2 = PostgreSQL::Test::Cluster->new('standby2');
+$standby2->init_from_backup($primary, 'inprogress_backup',
+ has_streaming => 1);
+$standby2->start;
+
+# Backup label recovery adopts the redo point's state immediately,
+# before any further WAL is replayed.
+wait_for_checksum_state($standby2, 'inprogress-on');
+
+$bsession->quit;
+wait_for_checksum_state($primary, 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable once the transition completes');
+
+# Scenario 4 continued: the backup taken mid-transition must also
+# complete the transition and stay readable.
+$primary->wait_for_catchup($standby2);
+wait_for_checksum_state($standby2, 'on');
+
+is($standby2->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'mid-transition backup standby readable after the transition');
+
+$standby2->stop;
+
+# Nothing along this standby's replay can legitimately disagree: it
+# started out at the same inprogress-on state as the primary.
+my $standby2_log = PostgreSQL::Test::Utils::slurp_file($standby2->logfile);
+unlike(
+ $standby2_log,
+ qr/does not match the state/,
+ 'no mismatch warning on a backup taken mid-transition');
+
+# No extra CHECKPOINT needed: the transition's completion checkpoint
+# wrote every page with a checksum, and the shutdown restartpoint
+# flushes the rest.
+command_ok([ 'pg_checksums', '--check', '-D', $standby2->data_dir ],
+ 'checksums valid on the mid-transition backup standby');
+
+$standby->stop;
+$primary->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/013_rewind.pl b/src/test/modules/test_checksums/t/013_rewind.pl
new file mode 100644
index 00000000000..a791e24317d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/013_rewind.pl
@@ -0,0 +1,201 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test pg_rewind across an online data checksum enable.
+#
+# A clean switchover leaves the shutdown checkpoint of the old primary
+# as the last common checkpoint between the two nodes. When data
+# checksums are enabled online on the new primary before the old one is
+# rewound, pg_rewind installs the control file of the new primary, which
+# already claims checksums are fully enabled, while replay begins at the
+# shutdown checkpoint whose record still carries the old state. Replay
+# of the WAL stretch from before the enable must run with checksums off,
+# as recorded in the checkpoint, else it would verify pages which never
+# had checksums written and fail recovery.
+#
+# The new primary requests a checkpoint right after promotion, and
+# replaying its CHECKPOINT_REDO record would repair the state before
+# any interesting WAL is reached. Hold the checkpointer on the standby
+# in a restartpoint over the promotion, like in the recovery test
+# 041_checkpoint_at_promote, so that the post-promotion writes end up
+# in WAL before the first checkpoint of the new timeline.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Old primary. full_page_writes is off so that the updates done on the
+# promoted node do not carry full page images, forcing replay on the
+# rewound node to read the pages from disk. wal_log_hints is required
+# by pg_rewind on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+full_page_writes = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+log_checkpoints = on
+]);
+$node_a->start;
+
+if (!$node_a->check_extension('injection_points'))
+{
+ plan skip_all => 'Extension injection_points not installed';
+}
+
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+$node_a->safe_psql('postgres', "CREATE EXTENSION injection_points;");
+
+# Set the hint bits on t before taking the backup, so that reads on the
+# promoted node do not emit full page images for its pages later.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->safe_psql('postgres', "SELECT count(*) FROM t;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_a, 'off');
+test_checksum_state($node_b, 'off');
+
+# Hold the next restartpoint on the standby.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point', 'wait');");
+
+# Give the restartpoint a checkpoint record to work on, then start it
+# in a background session; it will block on the injection point with
+# the checkpointer busy until released.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+my $bg_psql = $node_b->background_psql('postgres', on_error_stop => 0);
+$bg_psql->query_until(
+ qr/starting_restartpoint/, q(
+ \echo starting_restartpoint
+ CHECKPOINT;
+));
+$node_b->wait_for_event('checkpointer', 'create-restart-point');
+
+# Clean switchover: the shutdown checkpoint of A streams to B and
+# becomes the last common checkpoint.
+$node_a->stop('fast');
+
+my ($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+$stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+my $shutdown_ckpt = $1;
+$node_b->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_ckpt'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+
+my $logstart = -s $node_b->logfile;
+$node_b->promote;
+
+# Accidental restart of the old primary, diverging its timeline.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->stop('fast');
+
+# Updates on the new primary before its first checkpoint; replay of
+# these on the rewound node has to read the pages from disk.
+$node_b->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 25 = 0;");
+ok( !$node_b->log_contains("checkpoint complete", $logstart),
+ "no checkpoint on the new timeline before the updates");
+
+# Release the checkpointer; the queued post-promotion checkpoint runs
+# after the updates.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+$bg_psql->quit;
+
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+# pg_rewind refuses to run with full_page_writes disabled on the
+# source; the updates it was disabled for are already in WAL.
+$node_b->safe_psql('postgres', "ALTER SYSTEM SET full_page_writes = on;");
+$node_b->safe_psql('postgres', "SELECT pg_reload_conf();");
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the new primary');
+
+# Replay on the rewound node must start at the shutdown checkpoint of
+# the switchover, with the control file of the new primary.
+my $backup_label = slurp_file($node_a->data_dir . '/backup_label');
+$backup_label =~ /^CHECKPOINT LOCATION: ([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from backup_label";
+is($1, $shutdown_ckpt, 'replay starts at the switchover checkpoint');
+
+($stdout, $stderr) = run_command(
+ [
+ 'pg_waldump',
+ '-p' => $node_a->data_dir . '/pg_wal',
+ '-t' => 1,
+ '-s' => $shutdown_ckpt,
+ '-n' => 1,
+ ]);
+like($stdout, qr/CHECKPOINT_SHUTDOWN/,
+ 'last common checkpoint is a shutdown checkpoint');
+
+# pg_rewind keeps the target's own checksum state in the control file it
+# writes; the online enable reaches the rewound node through WAL replay
+# below, not through the copied control file.
+($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+like(
+ $stdout,
+ qr/^Data page checksum version:\s*0$/m,
+ 'rewound node keeps its own checksum state in the control file');
+
+# Start the rewound node as a standby of the new primary. Replay runs
+# through the pre-enable WAL stretch and the online enable.
+#
+# The rewind replaced the configuration files with those of the new
+# primary, so put the port back.
+my $connstr = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_b->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/014_lockstep.pl b/src/test/modules/test_checksums/t/014_lockstep.pl
new file mode 100644
index 00000000000..713f37fb22c
--- /dev/null
+++ b/src/test/modules/test_checksums/t/014_lockstep.pl
@@ -0,0 +1,182 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# The lockstep procedure for offline checksum changes in a replication
+# setup: stop all nodes, run pg_checksums on all of them, restart.
+# Replay of checkpoint records written before the change must not
+# revert the state of the standby.
+#
+# The same pair then runs the procedure in the disable direction, and
+# finally across a divergence: after a failover both nodes get their
+# checksums enabled offline, while the last common checkpoint still
+# carries "off". pg_rewind must keep the offline "on" on the rewound
+# node, since no record in the replayed WAL could ever restore it.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# wal_log_hints keeps the pair eligible for pg_rewind without checksums.
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Part 1: the lockstep procedure in the enable direction. Stop the
+# standby first: the WAL written after this point is replayed only
+# after the offline switch, and every checkpoint record in it still
+# carries the old state.
+$standby->stop;
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 0;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 1;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+# The lockstep procedure.
+$primary->checksum_enable_offline;
+$standby->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to off.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "off" in the replayed WAL/,
+ 0);
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep enable');
+
+# Crash the standby and replay the same stretch again.
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'on');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after crash restart');
+
+# Once a post-switch checkpoint has been replayed the states match and
+# no warning may be logged.
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no mismatch warning once the states match');
+
+# Every page the standby wrote in this window must carry a checksum.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$standby->stop;
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'checksums valid on the standby');
+
+# Part 2: the lockstep procedure in the other direction. The standby
+# is already stopped; give it a pre-switch WAL stretch to replay, with
+# every checkpoint record in it still carrying "on" -- an explicit
+# checkpoint plus the shutdown checkpoint written by stop().
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 2;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+$primary->checksum_disable_offline;
+$standby->checksum_disable_offline;
+
+my $disable_logstart = -s $standby->logfile;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to on.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $disable_logstart);
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep disable');
+
+# Part 3: failover, divergence and pg_rewind across offline enables.
+# The roles swap from here on: the standby becomes the new primary and
+# the old primary is rewound to follow it, but the variables keep the
+# names they had above.
+#
+# Failover: promote the standby, then diverge the old primary.
+$standby->promote;
+$standby->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$primary->stop;
+
+# The lockstep procedure across the divergence: both nodes get their
+# checksums enabled offline.
+$primary->checksum_enable_offline;
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+
+# The states match, so the rewind proceeds without complaint.
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $primary->data_dir,
+ '--source-server' => $standby->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled offline on both nodes');
+
+# The rewind replaced the configuration files with those of the source,
+# so put the port back. The copied file also carries the source's own
+# stale primary_conninfo; enable_streaming() below appends ours, which
+# wins as the later entry.
+$primary->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$primary->port]}
+]);
+$primary->enable_streaming($standby);
+
+my $rewind_logstart = -s $primary->logfile;
+$primary->start;
+$standby->wait_for_catchup($primary);
+
+is($primary->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+
+# The offline enable survives: replay from the common checkpoint must
+# not resurrect the pre-divergence "off".
+test_checksum_state($primary, 'on');
+
+$log =
+ PostgreSQL::Test::Utils::slurp_file($primary->logfile, $rewind_logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no divergence reported between the rewound server and its source');
+
+$primary->stop;
+$standby->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
new file mode 100644
index 00000000000..315b27b93d4
--- /dev/null
+++ b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
@@ -0,0 +1,125 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that a standby crashing after a replayed online disable, but before
+# its next restartpoint, restarts cleanly.
+#
+# Pages dirtied before the XLOG2_CHECKSUMS record and evicted after it are
+# written out under the new "off" state, without checksums. The control
+# file must follow the record immediately in this direction: a crash-restart
+# initializes checksum verification from the control file, and one still
+# saying "on" would fail verification on exactly those pages while replaying
+# records older than the state change.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$primary->start;
+
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,1000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+
+# Small buffer pool so that replaying a bulk load evicts the pages we care
+# about, and no restartpoints so the control file stays where it is.
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 1MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_delay = 10000
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'on');
+
+# No full page images, so that replay has to read the pages of "t" back from
+# disk instead of overwriting them from the WAL.
+$primary->append_conf('postgresql.conf', 'full_page_writes = off');
+$primary->reload;
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" on the standby, *before* the state change. They are
+# not flushed: the standby has no restartpoint from here on.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+# Online disable. The standby replays it and switches to "off", but its
+# control file is not updated.
+$primary->safe_psql('postgres', 'SELECT pg_disable_data_checksums();');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+my ($ctl_before) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($ctl_state) = $ctl_before =~ /Data page checksum version:\s+(\d+)/;
+note(
+ "standby control file data_checksum_version after the disable: "
+ . "$ctl_state (0 = off, 1 = on)");
+
+# Force the standby to evict the dirty pages of "t" now that it is "off":
+# they are written back without checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE filler AS SELECT generate_series(1,300000) AS a;");
+$primary->wait_for_catchup($standby);
+
+# Crash the standby before it gets a chance to run a restartpoint.
+$standby->stop('immediate');
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'standby restarts after crashing between a checksum state '
+ . 'change and the next restartpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'the pages written while "off" are readable')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("standby log tail:\n" . join("\n", @lines[ -12 .. -1 ]))
+ if @lines >= 12;
+ fail('the pages written while "off" are readable');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/016_promote_enable_crash.pl b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
new file mode 100644
index 00000000000..97f84da5685
--- /dev/null
+++ b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a standby promoted right after replaying an online enable, before any
+# restartpoint has flushed the rewritten pages.
+#
+# The end-of-recovery record persists the checksum state without flushing the
+# buffer pool, while the control file still points at a restartpoint older than
+# the transition. Persisting "on" there would make a crash before the
+# post-promotion checkpoint verify checksums over the pages that were flushed
+# while the state was still "off"; promotion must take the full checkpoint
+# instead.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Dirty the pages of "t" without full page images, then push them out to disk
+# on the standby while checksums are still off.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Online enable on the primary; the standby replays it into its own buffers,
+# where the rewritten pages stay dirty (no restartpoint, no bgwriter).
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+note("standby control file before the promotion: $before");
+
+# Promote. The end-of-recovery record persists the live state without any
+# buffer flush; the checkpoint requested afterwards is not immediate.
+$standby->promote;
+$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery();')
+ or die 'timed out waiting for the promotion';
+$standby->stop('immediate');
+
+($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($ckpt) = $ctl =~ /Latest checkpoint location:\s+(\S+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note(
+ "standby control file after the promotion: $after, checkpoint $ckpt, redo $redo"
+);
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'promoted node restarts after crashing before its first checkpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/017_restartpoint_race.pl b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
new file mode 100644
index 00000000000..65aa834f366
--- /dev/null
+++ b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
@@ -0,0 +1,138 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a restartpoint whose flush races the replay of an online enable.
+#
+# Replay keeps running while CheckPointGuts() writes out the buffer pool, so
+# the state at the end of the flush can be newer than the one the flush ran
+# under. Persisting it would claim checksums for pages the flush wrote out
+# before the transition, against a redo pointer that predates it.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# C1: the checkpoint the pending restartpoint will target.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" without full page images and push them to disk on
+# the standby while it is still "off".
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Start a restartpoint for C1 and hold it right after CheckPointGuts().
+$standby->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point','wait');");
+my $bg = $standby->background_psql('postgres');
+$bg->query_until(qr//, "\\echo restartpoint\nCHECKPOINT;\n");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint injection point';
+
+# The startup process keeps replaying while the checkpointer is held: the
+# whole online enable lands, and the rewritten pages stay dirty.
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+# Let the restartpoint finish; it persists the state it samples now.
+$standby->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint to finish';
+$standby->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+
+$standby->stop('immediate');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note("standby control file after the restartpoint: $after, redo $redo");
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started, 'standby restarts after a restartpoint that raced the enable');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/018_enable_crash_windows.pl b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
new file mode 100644
index 00000000000..17f96ff6272
--- /dev/null
+++ b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
@@ -0,0 +1,496 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Crash and checkpoint windows inside an online data checksum enable on
+# a single node. Each scenario starts from "off" on the same cluster,
+# runs the enable into a held injection point, and checks what a crash
+# or a concurrent checkpoint in that window may and may not do to the
+# control file and the pages on disk.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+use IPC::Run;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Bring the node back to "off" with no leftover table between
+# scenarios. The crash restarts already resolve an interrupted enable
+# to "off", so only disable when something is left to disable. Those
+# restarts also clear any attached injection points. The checkpoint
+# bounds the WAL the next scenario's crash recovery has to replay.
+sub reset_scenario
+{
+ my ($node) = @_;
+
+ BAIL_OUT('node is not running after the previous scenario')
+ unless $node->is_alive;
+
+ my $state = $node->safe_psql('postgres', 'SHOW data_checksums;');
+ disable_data_checksums($node, wait => 'off') if $state ne 'off';
+ $node->safe_psql('postgres', 'DROP TABLE IF EXISTS t;');
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+ test_checksum_state($node, 'off');
+}
+
+my $node = PostgreSQL::Test::Cluster->new('crash_windows');
+$node->init(no_data_checksums => 1);
+
+# Scenario 1 needs the pages of "t" to reach disk without a full page
+# image, hence full_page_writes and wal_log_hints off. Scenario 2
+# needs them to stay dirty in shared buffers until a checkpoint, hence
+# no bgwriter and a large shared_buffers. wal_level is pinned because
+# init() writes "minimal" for a non-streaming node. The rest merely
+# tolerate the settings.
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+wal_level = replica
+max_wal_size = 10GB
+full_page_writes = off
+wal_log_hints = off
+shared_buffers = 512MB
+bgwriter_lru_maxpages = 0
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$node->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+
+test_checksum_state($node, 'off');
+
+# Scenario 1: a primary crashes inside SetDataChecksumsOn(), just before the
+# forced checkpoint that flushes the rewritten pages, with crash recovery
+# resuming from a checkpoint older than the transition. The control file
+# must still say "inprogress-on" there: replay does not adopt the state of
+# the checkpoint record it resumes from, so an "on" written before the flush
+# would stay in effect while replay reads pages whose rewrite never reached
+# disk. Here the pages went out through the rewriting worker's ring buffer,
+# so only the control file state discriminates; see the next scenario for
+# the case where they did not.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Establish the checkpoint that crash recovery will resume from.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty the pages of "t" without full page images, then push them out
+ # to disk while checksums are still off.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+ my $evicted = $node->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+ note("evict_relation on t: $evicted, relpath $relpath");
+
+ # Stop the enable right after the control file has been updated to
+ # "on" but before the checkpoint that flushes the rewritten pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+ $node->safe_psql('postgres', 'SELECT pg_enable_data_checksums();');
+
+ $node->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'datachecksums-on-before-checkpoint';"
+ ) or die 'timed out waiting for the injection point';
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ $node->stop('immediate');
+
+ # Show the on-disk checksum field of the first page of "t".
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 2: the same crash as above, but with the rewritten pages resident
+# in shared buffers instead of gone out through the ring. The rewriting
+# worker reads through a BAS_VACUUM ring, which writes pages back as the
+# ring recycles, but a page already resident in shared buffers is not read
+# through the ring: ReadBufferExtended() hands back the existing buffer, and
+# it stays dirty until a checkpoint. The control file may therefore not
+# say "on" before that checkpoint has run, or crash recovery would resume
+# from a checkpoint older than the transition with verification already
+# enabled, and replay records that read those still-unchecksummed pages.
+# full_page_writes is off so that the records do not simply overwrite the
+# pages with a full page image.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,100000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Establish the checkpoint that crash recovery will resume from, with the
+ # pages of "t" written out while checksums are still off.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty those pages again without emitting full page images, and leave
+ # them in shared buffers. Replay of these records has to read the
+ # pages from disk.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+
+ my $dirty_before = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_before, '>', 0,
+ 'pages of t are resident and dirty before enabling checksums');
+
+ # Hold the enabling right before the checkpoint that flushes the rewritten
+ # pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ # The rewritten pages must still be sitting dirty in shared buffers, or
+ # the window this test is about does not exist.
+ my $dirty_after = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_after, '>', 0,
+ 'rewritten pages of t are still dirty before the checkpoint');
+
+ $node->stop('immediate');
+
+ # The on-disk copy is the one written before the transition, without a
+ # checksum.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ is($pd_checksum, 0, 'block 0 of t on disk carries no checksum');
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 3: a checkpoint that runs between the XLOG2_CHECKSUMS("on")
+# record and the control file write at the end of SetDataChecksumsOn() must
+# not make a crash throw the completed transition away. SetDataChecksumsOn()
+# writes the record, flips shared memory to "on", emits the barrier and only
+# then requests the checkpoint that flushes the rewritten pages; the control
+# file is written after that checkpoint returns. A crash in that window is
+# harmless only as long as recovery still starts before the record. Any
+# checkpoint completing in the window moves the redo point past the record,
+# so recovery would never see it, would come up with the control file's
+# "inprogress-on" and StartupXLOG() would demote that to "off", even though
+# every page on disk carries a checksum by then. CreateCheckPoint()
+# therefore persists the state the checkpoint ran under, the same way
+# CreateRestartPoint() does. Here the window is made deterministic by
+# holding the launcher at the datachecksums-on-before-checkpoint injection
+# point and checkpointing from another session.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Hold the launcher after the record, the shared memory flip and the
+ # barrier, but before the checkpoint SetDataChecksumsOn() requests
+ # itself.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ # Every backend already sees "on" and writes checksums.
+ test_checksum_state($node, 'on');
+
+ # ... while the control file still says "inprogress-on".
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($before, '3', 'control file says "inprogress-on" inside the window');
+
+ # A concurrent checkpoint. It flushes the rewritten pages and moves
+ # the redo point past the XLOG2_CHECKSUMS("on") record, so it has to
+ # record the "on" state in the control file as well.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($after, '1',
+ 'the concurrent checkpoint records "on" in the control file');
+
+ $node->stop('immediate');
+
+ # The checkpoint flushed the rewritten pages, so they carry a checksum on
+ # disk: the transition really did complete.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ isnt($pd_checksum, 0, 'block 0 of t on disk carries a checksum');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition is complete on disk, so the cluster has to come back
+ # "on".
+ test_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+}
+reset_scenario($node);
+
+# Scenario 4: a crash between the checkpoint an online enable requests and
+# the control file write that follows it must not lose the transition. The
+# last steps of SetDataChecksumsOn() are
+#
+# WAL record -> shmem -> barrier -> checkpoint -> persist
+#
+# The checkpoint is what makes "on" safe to persist, but it also moves the
+# redo point above the XLOG2_CHECKSUMS record that carries the new state.
+# Crash recovery started from that checkpoint therefore never replays the
+# record, so without the state the checkpoint itself persists, a crash in
+# the remaining window would bring the cluster back at "inprogress-on",
+# which StartupXLOG() resolves to "off", discarding a transition whose
+# pages are all on disk with a checksum. The previous scenario exercises
+# the same window through a checkpoint requested by another session; this
+# one closes it from the other side: the transition's own checkpoint is the
+# one that moves the redo point, so persisting the state from
+# CreateCheckPoint() is what has to save it, not the write below the
+# injection point.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Hold the launcher after the checkpoint that licenses "on" has
+ # completed and before the state reaches the control file.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-after-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-after-checkpoint');
+
+ # The transition is complete as far as the running cluster is concerned.
+ test_checksum_state($node, 'on');
+
+ # The checkpoint has already recorded it, so the pending write below the
+ # injection point has nothing left to do.
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($version) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($version, '1',
+ 'the requested checkpoint recorded "on" in the control file');
+
+ # Crash before the control file write that follows the checkpoint.
+ $node->stop('immediate');
+
+ $node->start;
+
+ # Every page on disk carries a checksum and the checkpoint that
+ # flushed them completed, so the cluster has to come back verifying
+ # them.
+ test_checksum_state($node, 'on');
+
+ is($node->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '10000', 'relation readable after the crash');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ # The data directory must be one the offline tools accept.
+ $node->stop;
+ $node->command_ok([ 'pg_checksums', '--check', '-D', $node->data_dir ],
+ 'pg_checksums accepts the data directory');
+
+ # Leave the node running for the reset.
+ $node->start;
+}
+reset_scenario($node);
+
+# Scenario 5: a checkpoint racing SetDataChecksumsOn() between the insertion
+# of the XLOG2_CHECKSUMS("on") record and the shared memory update must not
+# insert an XLOG_CHECKPOINT_REDO record that follows the transition in WAL
+# order while still carrying "inprogress-on". Recovery resuming from such a
+# redo point never replays the preceding transition record, comes up in
+# "inprogress-on" and resolves the finished transition as interrupted.
+# XLogChecksums() closes the window by inserting the record and publishing
+# the new state under DataChecksumTransitionLock, which CreateCheckPoint()
+# takes around sampling the state and inserting the redo record. Here the
+# launcher is held between the two steps at the
+# datachecksums-on-before-publish injection point, and a concurrent
+# CHECKPOINT has to block on the lock instead of completing inside the
+# window.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # The datachecksums-on-before-publish point fires inside a critical
+ # section, where the wait machinery must not allocate. Waiting once
+ # at the datachecksums-enable-checksums-delay point, which the
+ # launcher runs outside the critical section, initializes it; see
+ # 050_redo_segment_missing.pl for the same recipe around
+ # create-checkpoint-run.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-enable-checksums-delay','wait');"
+ );
+
+ # Hold the launcher after the XLOG2_CHECKSUMS("on") record is in WAL but
+ # before the new state is published in shared memory.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-publish','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-enable-checksums-delay');
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-enable-checksums-delay');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-enable-checksums-delay');");
+
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-publish');
+
+ # The record is in WAL, the published state is still the old one.
+ test_checksum_state($node, 'inprogress-on');
+
+ # A concurrent checkpoint. It must block on DataChecksumTransitionLock
+ # before inserting its redo record rather than complete inside the window.
+ my $checkpointer = IPC::Run::start(
+ [ 'psql', '-XAtq', '-d', $node->connstr('postgres'), '-c', 'CHECKPOINT;' ],
+ '<' => '/dev/null',
+ '>' => '/dev/null',
+ '2>' => '/dev/null',
+ IPC::Run::timer($PostgreSQL::Test::Utils::timeout_default));
+
+ ok( $node->poll_query_until(
+ 'postgres',
+ "SELECT wait_event = 'DataChecksumTransition' "
+ . "FROM pg_stat_activity WHERE backend_type = 'checkpointer';"),
+ 'concurrent checkpoint blocks on DataChecksumTransitionLock');
+
+ # Release the launcher; the checkpoint then samples the published "on" and
+ # its redo record follows the transition record.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-publish');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-publish');");
+
+ $checkpointer->finish;
+
+ # Crash while the launcher's own checkpoint may still be in flight.
+ # Recovery resumes from the concurrent checkpoint's redo point, which
+ # now lies above the transition record and carries "on".
+ $node->stop('immediate');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition completed, so the cluster has to come back "on".
+ wait_for_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ $node->stop;
+}
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
new file mode 100644
index 00000000000..aaa35f8f7dd
--- /dev/null
+++ b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
@@ -0,0 +1,126 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A standby stopped after replaying an online enable, but before the
+# checkpoint record that follows it on the primary.
+#
+# The shutdown restartpoint has no new checkpoint record to work from
+# and is skipped, so the control file would keep the in-progress state
+# the transition passed through, even though replay left the node at
+# "on" and rewrote every page. pg_checksums reads that field and
+# refuses to run on an in-progress state, so the shutdown has to flush
+# and catch it up instead.
+#
+# The caught-up control file then carries the watermark of the "on"
+# record, which the second half of this test relies on: an offline
+# disable made while the standby is down must survive the re-replay of
+# that record on the next startup, or the change would be silently
+# reverted.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+wal_keep_size = 1GB
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Anchor the standby's last restartpoint here, so that nothing replayed
+# from now on gives the shutdown restartpoint a newer checkpoint record
+# to use.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Hold the enable right after the state change record, before the
+# checkpoint it requests once the transition is complete.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+enable_data_checksums($primary);
+$primary->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+wait_for_checksum_state($primary, 'on');
+
+# The state change record is already flushed and streams on its own;
+# no checkpoint record follows while the launcher is held.
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+$standby->stop('fast');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($state, '1',
+ 'shutdown catches the control file up with the replayed state');
+
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'pg_checksums verifies the stopped standby');
+
+# Disable checksums offline while the standby is down. The restart
+# below resumes replay under the last restartpoint, re-reading the "on"
+# record; the watermark persisted with the catch-up above marks it as
+# already applied, so the offline change must survive.
+$standby->checksum_disable_offline;
+
+# Let the enable finish on the primary.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-checkpoint');");
+$primary->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-checkpoint');");
+
+# The launcher exits only after the checkpoint it requests completes,
+# so a checkpoint record now follows the transition in WAL.
+$primary->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_catalog.pg_stat_activity "
+ . "WHERE backend_type = 'datachecksums launcher';");
+
+my $log_offset = -s $standby->logfile;
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+
+# The nodes legitimately diverged, which replayed checkpoints report.
+$standby->wait_for_log(
+ qr/data checksum state "off" of this node does not match the state "on" in the replayed WAL/,
+ $log_offset);
+
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/020_cascade_divergence.pl b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
new file mode 100644
index 00000000000..b4b474fa583
--- /dev/null
+++ b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
@@ -0,0 +1,115 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# An offline data checksum state change made on one node of a cascading setup
+# is reported all the way down the chain.
+#
+# CheckReplayedDataChecksumState() is reached from xlog_redo() when a
+# checkpoint record (XLOG_CHECKPOINT_SHUTDOWN or XLOG_CHECKPOINT_REDO) is
+# replayed after consistency has been reached. Those records are written by
+# the root primary only and are relayed verbatim by every intermediate
+# standby, so a cascaded standby that was changed offline learns about the
+# divergence too. An intermediate standby's own restartpoints are not
+# WAL-logged and cannot surface it.
+#
+# Note that this makes detection checkpoint-driven: an idle or read-only
+# primary surfaces the divergence no sooner than its next checkpoint, so up to
+# checkpoint_timeout may pass before the operator sees anything. Closing that
+# window would require comparing the state when a walreceiver connects.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# checkpoint_timeout is deliberately long so that the only checkpoint record in
+# this test is the one requested explicitly below.
+my $conf = qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+);
+
+my $primary = PostgreSQL::Test::Cluster->new('cascade_primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', $conf);
+$primary->start;
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,1000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('cascade_standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf('postgresql.conf', $conf);
+$standby->start;
+
+$standby->backup('backup');
+my $cascade = PostgreSQL::Test::Cluster->new('cascade_cascade');
+$cascade->init_from_backup($standby, 'backup', has_streaming => 1);
+$cascade->append_conf('postgresql.conf', $conf);
+$cascade->start;
+
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+test_checksum_state($cascade, 'off');
+
+# Change the cascaded standby offline, and only it.
+$cascade->stop;
+$cascade->command_ok([ 'pg_checksums', '--enable', '-D', $cascade->data_dir ],
+ 'pg_checksums enables checksums on the cascaded standby only');
+
+my $logstart = -s $cascade->logfile;
+$cascade->start;
+
+test_checksum_state($cascade, 'on');
+
+# Ordinary WAL traffic carries no state, so the cascaded standby keeps
+# streaming from a chain whose state is "off" without noticing.
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (1);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+is($cascade->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '1001', 'cascaded standby keeps streaming from a divergent chain');
+
+# Restartpoints on the intermediate standby are not WAL-logged, so this does
+# not surface anything on the cascaded standby either.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (2);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+my $log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+unlike(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'an intermediate restartpoint does not surface the divergence');
+
+# A checkpoint on the root primary does: the record is relayed down the whole
+# chain, so both the direct and the cascaded standby compare it against their
+# own state. wait_for_catchup() below waits for replay, not just receipt, so
+# the record has been through xlog_redo() on both nodes once it returns.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+$log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+like(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'the cascaded standby reports the divergence on the primary checkpoint');
+
+$cascade->stop;
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
new file mode 100644
index 00000000000..a6c8a1091ee
--- /dev/null
+++ b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
@@ -0,0 +1,193 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind across online data checksum transitions on both sides of a
+# divergence, with the two nodes trading roles between the scenarios.
+#
+# Scenario 1: after a switchover the new primary enables checksums
+# online, while the old primary restarts on its old timeline, advances
+# its WAL beyond the enable records and runs an online enable/disable
+# cycle of its own. The old primary ends "off" with a checksum
+# watermark numerically above every checksum record the new primary has
+# written. pg_rewind clamps the watermark it keeps to the divergence
+# point; without the clamp, replay on the rewound node would skip the
+# source's enable as already applied and stay "off" under an "on"
+# primary.
+#
+# Scenario 2: checksums are disabled again, and after another
+# switchover both nodes enable them online independently, so the
+# divergence checkpoint carries "off" while both control files say
+# "on". The rewind is allowed, the target keeps its own "on" state,
+# and replay re-walks the source's enable onto the already enabled
+# node.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+sub controldata_watermark
+{
+ my ($node) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $node->data_dir ]);
+ $stdout =~ /^Data checksum watermark:\s*([0-9A-F]+)\/([0-9A-F]+)$/m
+ or die "watermark missing from pg_controldata output";
+ return (hex($1) << 32) + hex($2);
+}
+
+# Wait until the standby has replayed the shutdown checkpoint of the
+# stopped primary, so that a subsequent promotion diverges after it and
+# the shutdown checkpoint becomes the last common checkpoint.
+sub wait_for_shutdown_checkpoint_replay
+{
+ my ($primary, $standby) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $primary->data_dir ]);
+ $stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+ my $shutdown_checkpoint = $1;
+
+ $standby->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_checkpoint'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+}
+
+# Old primary, checksums off. wal_log_hints is required by pg_rewind
+# on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+# Clean switchover to B; enable checksums online on it.
+$node_a->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_a, $node_b);
+$node_b->promote;
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+$node_b->stop('fast');
+my $watermark_b = controldata_watermark($node_b);
+$node_b->start;
+
+# Accidental restart of the old primary on the old timeline. Advance
+# its WAL beyond the enable watermark of B, then run an online enable
+# and disable cycle: the node ends "off" with a watermark above every
+# checksum record B has written.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t_pad AS SELECT generate_series(1,200000) AS a;");
+enable_data_checksums($node_a, wait => 'on');
+disable_data_checksums($node_a, wait => 'off');
+test_checksum_state($node_a, 'off');
+$node_a->stop('fast');
+
+my $watermark_a = controldata_watermark($node_a);
+die "test broken: target watermark not above the source's enable"
+ unless $watermark_a > $watermark_b;
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the promoted node');
+
+# Start the rewound node as a standby of B. Replay from the last
+# common checkpoint runs through B's online enable, which the target
+# never saw, so the rewound node must converge to "on".
+my $connstr_b = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr_b application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+# Scenario 2, reusing the pair with the roles reversed. Disable
+# checksums online so the next divergence point carries "off", and let
+# A replay the change.
+disable_data_checksums($node_b, wait => 'off');
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'off');
+
+# Clean switchover back to A; enable checksums online on it.
+$node_b->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_b, $node_a);
+$node_a->promote;
+enable_data_checksums($node_a, wait => 'on');
+test_checksum_state($node_a, 'on');
+my $source_enable_watermark = controldata_watermark($node_a);
+
+# The old primary restarts on its old timeline and enables checksums
+# online independently: both control files say "on", the divergence
+# checkpoint says "off".
+$node_b->start;
+$node_b->safe_psql('postgres', "INSERT INTO t_div VALUES (2);");
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+$node_b->stop('fast');
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_b->data_dir,
+ '--source-server' => $node_a->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled on both nodes');
+
+my $connstr_a = $node_a->connstr;
+$node_b->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_b->port]}
+primary_conninfo = '$connstr_a application_name=@{[$node_b->name]}'
+]);
+$node_b->set_standby_mode;
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_b, 'on');
+
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_b->stop('fast');
+is(controldata_watermark($node_b), $source_enable_watermark,
+ 'rewound node replayed the source checksum transition');
+command_ok([ 'pg_checksums', '--check', '-D', $node_b->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the twice-rewound node');
+
+done_testing();
--
2.55.0
[application/octet-stream] v9-0004-pg_combinebackup-Refuse-mixed-data-checksum-state.patch (9.5K, ../../CAN4CZFNdfb-yFRW7Sh3FkZ3Qc91Lz-JDQN0b-070-6G35g5Ssg@mail.gmail.com/4-v9-0004-pg_combinebackup-Refuse-mixed-data-checksum-state.patch)
download | inline diff:
From 6ee92e7a6db57a14956f6bc5b474462c95a0f2d7 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Tue, 25 Aug 2026 15:06:58 +0000
Subject: [PATCH v9 4/4] pg_combinebackup: Refuse mixed data checksum states in
a backup chain
check_control_files() detected a chain whose backups were taken under
different data checksum states, warned, and proceeded. The output
directory keeps the last backup's control file, so a full backup taken
with checksums off combined with an incremental taken after an offline
enable produces a cluster whose control file says "on" while most of
its blocks carry no checksums; it starts, and then every connection
dies on the first unchecksummed catalog page. An offline enable
between two backups of a chain is all it takes, since it rewrites
every page without logging anything, so the incremental backup does
not re-ship the pages.
Turn the warning into an error, matching what pg_rewind does for the
equivalent combinations. The check stays asymmetric on purpose: when
the last backup was taken with checksums off, stale checksums from an
earlier backup are never verified, and that chain remains usable.
---
doc/src/sgml/ref/pg_combinebackup.sgml | 19 +--
src/bin/pg_combinebackup/pg_combinebackup.c | 14 +-
src/test/modules/test_checksums/meson.build | 1 +
.../t/024_combinebackup_mixed.pl | 134 ++++++++++++++++++
4 files changed, 155 insertions(+), 13 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
diff --git a/doc/src/sgml/ref/pg_combinebackup.sgml b/doc/src/sgml/ref/pg_combinebackup.sgml
index 9a6d201e0b8..f5d7e177b36 100644
--- a/doc/src/sgml/ref/pg_combinebackup.sgml
+++ b/doc/src/sgml/ref/pg_combinebackup.sgml
@@ -306,18 +306,19 @@ PostgreSQL documentation
<para>
<literal>pg_combinebackup</literal> does not recompute page checksums when
- writing the output directory. Therefore, if any of the backups used for
- reconstruction were taken with checksums disabled, but the final backup was
- taken with checksums enabled, the resulting directory may contain pages
- with invalid checksums.
+ writing the output directory. It therefore refuses a chain in which some
+ of the backups used for reconstruction were taken with checksums disabled
+ but the final backup was taken with checksums enabled: the resulting
+ directory would contain pages with invalid checksums, and the cluster
+ would fail checksum verification as soon as it reads them. Take a new
+ full backup after enabling data checksums with
+ <xref linkend="app-pgchecksums"/>.
</para>
<para>
- To avoid this problem, taking a new full backup after changing the checksum
- state of the cluster using <xref linkend="app-pgchecksums"/> is
- recommended. Otherwise, you can disable and then optionally reenable
- checksums on the directory produced by <literal>pg_combinebackup</literal>
- in order to correct the problem.
+ The reverse case is accepted: when the final backup was taken with
+ checksums disabled, stale checksums copied from an earlier backup are
+ never verified.
</para>
</refsect1>
diff --git a/src/bin/pg_combinebackup/pg_combinebackup.c b/src/bin/pg_combinebackup/pg_combinebackup.c
index 86e2ee37c40..254a27b125b 100644
--- a/src/bin/pg_combinebackup/pg_combinebackup.c
+++ b/src/bin/pg_combinebackup/pg_combinebackup.c
@@ -670,13 +670,19 @@ check_control_files(int n_backups, char **backup_dirs)
pg_log_debug("system identifier is %" PRIu64, system_identifier);
/*
- * Warn the user if not all backups are in the same state with regards to
- * checksums.
+ * Reject a chain whose backups were not all in the same state with
+ * regards to checksums. The loop above only flags that when the last
+ * backup has checksums enabled: the blocks taken from the older backups
+ * have no checksums, and pg_combinebackup does not recompute them, so the
+ * combined cluster would fail verification as soon as it reads them. The
+ * other direction is harmless, since stale checksums are never verified
+ * while checksums are disabled.
*/
if (data_checksum_mismatch)
{
- pg_log_warning("only some backups have checksums enabled");
- pg_log_warning_hint("Disable, and optionally reenable, checksums on the output directory to avoid failures.");
+ pg_log_error("only some backups have checksums enabled");
+ pg_log_error_hint("Take a new full backup after changing the data checksum state with pg_checksums.");
+ exit(1);
}
return system_identifier;
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index 7d07c757052..7eccd5156b9 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -47,6 +47,7 @@ tests += {
't/021_rewind_divergent_transitions.pl',
't/022_rewind_state.pl',
't/023_rewind_standby_target.pl',
+ 't/024_combinebackup_mixed.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
new file mode 100644
index 00000000000..ce9dee9413b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_combinebackup refuses a chain whose backups were taken under
+# different data checksum states when the final state has them enabled.
+#
+# The output directory keeps the last backup's control file, so a full
+# backup taken with checksums off plus an incremental taken after an
+# offline enable would produce a directory whose control file says "on"
+# while most of its blocks have no checksums. An offline enable between
+# the two backups is enough to get there: it rewrites every page but logs
+# nothing, so the incremental backup does not re-ship the pages.
+#
+# The reverse order stays allowed: with the final backup taken with
+# checksums off, stale checksums from an earlier backup are never verified.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $node = PostgreSQL::Test::Cluster->new('node');
+$node->init(
+ no_data_checksums => 1,
+ has_archiving => 1,
+ allows_streaming => 1);
+$node->append_conf('postgresql.conf', 'summarize_wal = on');
+$node->append_conf('postgresql.conf', 'autovacuum = off');
+$node->start;
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+test_checksum_state($node, 'off');
+
+$node->backup('full');
+
+# Offline enable: rewrites every page, logs nothing.
+$node->stop;
+system_or_bail('pg_checksums', '--enable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'on');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full/backup_manifest',
+ ],
+ 'incremental backup with checksums on');
+
+# Combining the chain must be refused: the result would say "on" while the
+# blocks inherited from the full backup have no checksums.
+command_fails_like(
+ [
+ 'pg_combinebackup',
+ $node->backup_dir . '/full',
+ $node->backup_dir . '/incr',
+ '--output' => $node->backup_dir . '/combined',
+ ],
+ qr/only some backups have checksums enabled/,
+ 'pg_combinebackup refuses a chain crossing an offline enable');
+
+# The other direction: full backup with checksums on, offline disable, then
+# an incremental. The last backup wins, so the combined cluster comes up off.
+$node->backup('full2');
+
+$node->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr2',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full2/backup_manifest',
+ ],
+ 'incremental backup with checksums off');
+
+my $restored = PostgreSQL::Test::Cluster->new('restored');
+$restored->init_from_backup(
+ $node, 'incr2',
+ combine_with_prior => ['full2'],
+ has_restoring => 1,
+ standby => 0);
+
+$restored->start;
+
+my ($rc, $stdout, $stderr) = $restored->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'combined backup accepts connections') or diag("stderr: $stderr");
+is($stdout, 'off', 'combined backup follows the last backup state');
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t3;');
+is($rc, 0, 'blocks from the incremental backup are readable')
+ or diag("stderr: $stderr");
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t1;');
+is($rc, 0, 'blocks inherited from the full backup are readable')
+ or diag("stderr: $stderr");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($restored->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures in the combined backup');
+
+$restored->stop('immediate');
+$node->stop;
+
+done_testing();
--
2.55.0
[application/octet-stream] v9-0002-pg_checksums-Refuse-interrupted-transitions-note-.patch (7.7K, ../../CAN4CZFNdfb-yFRW7Sh3FkZ3Qc91Lz-JDQN0b-070-6G35g5Ssg@mail.gmail.com/5-v9-0002-pg_checksums-Refuse-interrupted-transitions-note-.patch)
download | inline diff:
From 8ea56cddee88222cbf2d906fdf847bdb22a1e959 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:19 +0000
Subject: [PATCH v9 2/4] pg_checksums: Refuse interrupted transitions, note
change is local
An inprogress-on or inprogress-off control file means an online
transition was cut short; running the offline tool on top of it mixes
two procedures. Refuse it in every mode, with a hint pointing at the
start-stop cycle (or, on a standby, at letting replication finish)
that resets the state. Also print that the change applies to one
data directory only, with standby-aware wording, since in a
replication setup the same change must be made on every node.
On a primary, inprogress-on cannot survive a graceful stop: the
checksums launcher resolves it from its exit cleanup, and a crashed
primary is already rejected by the existing "cluster must be shut
down" check. inprogress-off can, however: it is set by the backend
running pg_disable_data_checksums(), and a fast shutdown arriving
between its two barriers leaves it behind in a cleanly shut down
control file, where the next start-stop cycle resolves it at end of
recovery. A standby can be stopped with either state, since it has
no launcher and only carries forward whatever state the last replayed
WAL record left it in, with a restartpoint persisting that as-is.
The standby is also the deterministic way to reach the new guard,
which is why the test coverage uses one; the primary window would
need an injection point between the two barriers.
The pg_checksums page said the tool still processes all relation files
regardless of an interrupted online transition; describe the refusal
instead.
---
doc/src/sgml/ref/pg_checksums.sgml | 11 ++++---
src/bin/pg_checksums/pg_checksums.c | 21 +++++++++++++
.../test_checksums/t/012_offline_standby.pl | 31 ++++++++++++++-----
3 files changed, 51 insertions(+), 12 deletions(-)
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index 6f332ade462..b337950be3d 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -49,11 +49,12 @@ PostgreSQL documentation
</para>
<para>
- When enabling checksums with <application>pg_checksums</application>, if
- checksums were in the process of being enabled using
- <xref linkend="checksums-online-enable-disable"/> when the cluster was shut
- down, <application>pg_checksums</application> will still process all
- relation files regardless of the progress of online checksum processing.
+ If checksums were in the process of being enabled or disabled using
+ <xref linkend="checksums-online-enable-disable"/> when the cluster was
+ shut down, the control file still records that interrupted state, and
+ <application>pg_checksums</application> refuses to run in any mode.
+ Start the cluster and shut it down cleanly to reset the state, then
+ retry; on a standby, let replication complete the transition first.
</para>
<para>
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 20a99112838..a49b6f368b7 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -586,6 +586,21 @@ main(int argc, char *argv[])
ControlFile->state != DB_SHUTDOWNED_IN_RECOVERY)
pg_fatal("cluster must be shut down");
+ /*
+ * An inprogress state means an online transition was cut short. A
+ * standby stopped mid-transition carries either state; a cleanly shut
+ * down primary can still carry inprogress-off, which a fast shutdown
+ * during pg_disable_data_checksums() leaves behind, while inprogress-on
+ * is always resolved by the launcher's exit cleanup.
+ */
+ if (ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted");
+ pg_log_error_hint("Start and cleanly shut down the cluster once to reset the data checksum state, then retry. On a standby, let replication complete the transition first.");
+ exit(1);
+ }
+
if (ControlFile->data_checksum_version != PG_DATA_CHECKSUM_VERSION &&
mode == PG_MODE_CHECK)
pg_fatal("data checksums are not enabled in cluster");
@@ -673,6 +688,12 @@ main(int argc, char *argv[])
printf(_("Checksums enabled in cluster\n"));
else
printf(_("Checksums disabled in cluster\n"));
+
+ printf(_("This change applies to this data directory only.\n"));
+ if (ControlFile->state == DB_SHUTDOWNED_IN_RECOVERY)
+ printf(_("This node appears to be a standby; apply the same change to the primary and all other standbys.\n"));
+ else
+ printf(_("In a replication setup, apply the same change to every node while all are stopped, before restarting any of them.\n"));
}
return 0;
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
index 241a1282ae9..5efe11f993e 100644
--- a/src/test/modules/test_checksums/t/012_offline_standby.pl
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -149,7 +149,10 @@ $newnode->stop('immediate');
# Converge the cluster: enable offline on the standby too.
$standby->stop;
-$standby->checksum_enable_offline;
+command_checks_all(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ 0, [qr/appears to be a standby/],
+ [], 'standby-role notice on offline enable');
$standby->start;
test_checksum_state($standby, 'on');
$primary->wait_for_catchup($standby);
@@ -217,11 +220,11 @@ unlike(
);
# Scenario 4: a standby stopped while replaying an interrupted online
-# transition keeps the interrupted state in its own control file. A
-# primary is never caught this way, as its checksums launcher resolves
-# inprogress-on back to off from its exit cleanup. A standby has no
-# launcher; it carries forward whatever the last replayed record left
-# it in.
+# transition keeps the interrupted state in its own control file, and
+# pg_checksums must refuse to touch it. A primary is never caught this
+# way, as its checksums launcher resolves inprogress-on back to off from
+# its exit cleanup. A standby has no launcher; it carries forward
+# whatever the last replayed record left it in.
# Block an online enable on the primary at inprogress-on with a
# blocking temp table, same trick as in 004_offline.pl.
@@ -235,6 +238,20 @@ wait_for_checksum_state($standby, 'inprogress-on');
# Stop the standby cleanly; its restartpoint persists inprogress-on to
# its own control file, since nothing on a standby resolves it away.
$standby->stop;
+
+command_fails_like(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --enable refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --check refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--disable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --disable refuses a standby stopped mid-transition');
+
$standby->start;
wait_for_checksum_state($standby, 'inprogress-on');
@@ -257,7 +274,7 @@ wait_for_checksum_state($primary, 'on');
$primary->wait_for_catchup($standby);
wait_for_checksum_state($standby, 'on');
-is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
'10001', 'standby readable once the transition completes');
# Scenario 4 continued: the backup taken mid-transition must also
--
2.55.0
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-02 07:45 Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
parent: Zsolt Parragi <zsolt.parragi@percona.com>
0 siblings, 1 reply; 43+ messages in thread
From: Bertrand Drouvot @ 2026-09-02 07:45 UTC (permalink / raw)
To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; pgsql-hackers@lists.postgresql.org
Hi,
On Tue, Sep 01, 2026 at 02:50:04PM +0100, Zsolt Parragi wrote:
> Thanks!
>
> I applied these changes to v9 with some additional comment editing. I
> also squashed 0005 into 0001 because it describes what's implemented
> there, and I also tried to significantly reduce the commit message of
> 0001. Otherwise everything else is unchanged.
Thanks!
I initially thought there could be two more issues: one involving a base backup
spanning an online enable and another involving a crash during the first recovery
after pg_rewind. Further testing showed that neither was an issue.
So I'm happy with the current v9-0001 behavior. I now just have a couple of
wording comments:
=== 1
+ only then restart them. Before stopping a standby, make sure it has
+ replayed all WAL of its upstream node, for example by stopping the
+ primary first and comparing
+ <function>pg_last_wal_replay_lsn()</function> with
+ <function>pg_last_wal_receive_lsn()</function> on the standby.
Equality only proves that all received WAL has been replayed, not that all
upstream WAL was received. Maybe we should compare against the stopped
primary's shutdown checkpoint location, as 021 does?
=== 2
+ /*
+ * Mark the state as changed locally, without a WAL record. Recovery
+ * then knows the state is newer than anything the WAL carries and
+ * does not let a replayed checkpoint overwrite it. The watermark is
+ * left alone: any XLOG2_CHECKSUMS record this node had applied stays
+ * covered, and only records above it, written after this change, take
+ * effect again.
+ */
An offline change has no ordering against WAL not yet replayed, so records above
the watermark may have been written before the offline change. Maybe this should
be worded in terms of records covered by the watermark, without implying
chronological ordering?
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-02 12:13 Zsolt Parragi <zsolt.parragi@percona.com>
parent: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 1 reply; 43+ messages in thread
From: Zsolt Parragi @ 2026-09-02 12:13 UTC (permalink / raw)
To: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; pgsql-hackers@lists.postgresql.org
Thanks, fixed in v10, along with another comment in pg_control.h,
otherwise it is identical to v9.
Attachments:
[application/octet-stream] v10-0003-pg_rewind-Check-the-data-checksum-states-of-sour.patch (21.6K, ../../CAN4CZFOOHZmVnL-B2D+V7ODmVFgkYucgQD+-qegDTKZgJ3dtQg@mail.gmail.com/2-v10-0003-pg_rewind-Check-the-data-checksum-states-of-sour.patch)
download | inline diff:
From 044ff44049a4d61fe3f6a7503d89de95a5aed0e9 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 18:40:22 +0000
Subject: [PATCH v10 3/4] pg_rewind: Check the data checksum states of source
and target
pg_rewind installs the source's control file on the target, but every
block it does not copy keeps the target's content. With checksums
enabled on the target and disabled on the source, the rewound server
claims enabled checksums while the blocks copied from the source have
none, and fails checksum verification as soon as it reads them; in the
worst case every connection attempt dies on an unverifiable catalog
page.
Comparing the control files is not enough. Replay on the rewound
server resumes from the last common checkpoint and adopts the data
checksum state recorded there, so an offline disable on the source
after the divergence leaves both control files saying "off" while the
rewound server still resumes with verification enabled. Check the
state carried by the divergence checkpoint as well, which pg_rewind
already reads.
Refuse both cases, and an interrupted online transition on either
side, same as pg_checksums does. The opposite mismatch, checksums on
the source only, is allowed with a warning: recovery under the backup
label adopts the divergence state, so the rewound server keeps
checksums disabled (or converges through the replayed WAL if the
source enabled them online), and no page can fail verification.
---
doc/src/sgml/ref/pg_rewind.sgml | 14 ++
src/bin/pg_rewind/parsexlog.c | 4 +-
src/bin/pg_rewind/pg_rewind.c | 93 +++++++++++-
src/bin/pg_rewind/pg_rewind.h | 1 +
src/test/modules/test_checksums/meson.build | 2 +
.../test_checksums/t/022_rewind_state.pl | 133 +++++++++++++++++
.../t/023_rewind_standby_target.pl | 141 ++++++++++++++++++
7 files changed, 386 insertions(+), 2 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/022_rewind_state.pl
create mode 100644 src/test/modules/test_checksums/t/023_rewind_standby_target.pl
diff --git a/doc/src/sgml/ref/pg_rewind.sgml b/doc/src/sgml/ref/pg_rewind.sgml
index b95e3868e9d..ac3d0c9328f 100644
--- a/doc/src/sgml/ref/pg_rewind.sgml
+++ b/doc/src/sgml/ref/pg_rewind.sgml
@@ -124,6 +124,20 @@ PostgreSQL documentation
<literal>on</literal>, but is enabled by default.
</para>
+ <para>
+ The data checksum states of the source and target server must be
+ compatible. <application>pg_rewind</application> refuses to run if an
+ online checksum state transition was interrupted on either server, or
+ if data checksums are enabled on the target server, or were enabled at
+ the point of divergence, while the source server runs without them: the
+ rewound server would fail checksum verification on the blocks copied
+ from the source. The opposite combination is allowed with a warning;
+ the rewound server keeps data checksums disabled unless the source
+ server enabled them online after the point of divergence. See
+ <xref linkend="app-pgchecksums"/> for how to change the state
+ consistently across a replication setup.
+ </para>
+
<warning>
<title>Warning: Failures While Rewinding</title>
<para>
diff --git a/src/bin/pg_rewind/parsexlog.c b/src/bin/pg_rewind/parsexlog.c
index 023e23b063c..6e87b00f8c2 100644
--- a/src/bin/pg_rewind/parsexlog.c
+++ b/src/bin/pg_rewind/parsexlog.c
@@ -167,7 +167,8 @@ readOneRecord(const char *datadir, XLogRecPtr ptr, int tliIndex,
void
findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
- XLogRecPtr *lastchkptredo, const char *restoreCommand)
+ XLogRecPtr *lastchkptredo, uint32 *lastchkptdatachecksums,
+ const char *restoreCommand)
{
/* Walk backwards, starting from the given record */
XLogRecord *record;
@@ -255,6 +256,7 @@ findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
*lastchkptrec = searchptr;
*lastchkpttli = checkPoint.ThisTimeLineID;
*lastchkptredo = checkPoint.redo;
+ *lastchkptdatachecksums = checkPoint.dataChecksumState;
break;
}
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 466f4223501..9b99e628d4e 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -145,6 +145,7 @@ main(int argc, char **argv)
XLogRecPtr chkptrec;
TimeLineID chkpttli;
XLogRecPtr chkptredo;
+ uint32 chkptdatachecksums;
TimeLineID source_tli;
TimeLineID target_tli;
XLogRecPtr target_wal_endrec;
@@ -472,10 +473,51 @@ main(int argc, char **argv)
keepwal_init();
findLastCheckpoint(datadir_target, divergerec, lastcommontliIndex,
- &chkptrec, &chkpttli, &chkptredo, restore_command);
+ &chkptrec, &chkpttli, &chkptredo, &chkptdatachecksums,
+ restore_command);
pg_log_info("rewinding from last common checkpoint at %X/%08X on timeline %u",
LSN_FORMAT_ARGS(chkptrec), chkpttli);
+ /*
+ * Replay on the rewound server resumes from the last common checkpoint
+ * and adopts the data checksum state recorded there, not the state in the
+ * control file installed from the source. If checksums were enabled at
+ * the divergence point but the source runs without them, the rewound
+ * server would verify checksums while the blocks copied from the source
+ * have none. The control files cannot reveal this: an offline disable on
+ * the source after the divergence leaves both of them saying "off".
+ *
+ * Only the fully enabled state needs checking. An in-progress state at
+ * the divergence point means a transition was still running there, and
+ * all of its page rewrites are logged after that point, so replay brings
+ * the rewound server to whatever state the copied WAL ends in.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the source server or recreate the target server from a base backup.");
+ exit(1);
+ }
+
+ /*
+ * The same comparison is needed against the target: every block the
+ * rewind does not copy keeps the target's content. The control files
+ * cannot reveal this case either, in the other direction: a standby's
+ * checkpoints are written by its upstream primary, so a standby whose
+ * checksums were disabled offline still has "on" checkpoints in its WAL,
+ * while both control files may agree.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_target.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the target server");
+ pg_log_error_detail("Blocks kept from the target would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the target server with pg_checksums or recreate the target server from a base backup.");
+ exit(1);
+ }
+
/* Initialize the hash table to track the status of each file */
filehash_init();
@@ -799,6 +841,55 @@ sanityChecks(void)
pg_fatal("target server needs to use either data checksums or \"wal_log_hints = on\"");
}
+ /*
+ * The rewound target keeps its control file fields from the source, but
+ * every block the rewind does not copy keeps the target's content. If
+ * checksums are enabled on the target and disabled on the source, the
+ * result would claim enabled checksums while the blocks copied from the
+ * source have none, and the target would fail checksum verification as
+ * soon as it reads them. Refuse that combination, and refuse an
+ * interrupted online transition on either side, same as pg_checksums.
+ *
+ * The opposite mismatch is allowed: recovery under the backup label
+ * written by pg_rewind adopts the checksum state as of the divergence
+ * point, so a target whose own checkpoints carry no checksums stays
+ * without them, or converges through the replayed WAL if the source
+ * enabled them online. Only warn about it, so that an offline change on
+ * the source is not overlooked.
+ *
+ * That reasoning does not hold when the target is a standby, whose
+ * checkpoints were written by its upstream primary; see the checks after
+ * findLastCheckpoint().
+ */
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted on the target server");
+ pg_log_error_hint("Start the server and shut it down cleanly to reset the state, then retry; on a standby, let replication complete the transition first.");
+ exit(1);
+ }
+ if (ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition is incomplete on the source server");
+ pg_log_error_hint("Let the transition complete, or reset the state with a clean restart, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_OFF)
+ {
+ pg_log_error("data checksums are enabled on the target server but disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, and the target would fail checksum verification after the rewind.");
+ pg_log_error_hint("Either disable data checksums on the target server or enable them on the source server, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_OFF &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_warning("data checksums are disabled on the target server but enabled on the source server");
+ pg_log_warning_detail("The rewound server will keep data checksums disabled unless the source server enabled them online after the point of divergence.");
+ }
+
/*
* Target cluster better not be running. This doesn't guard against
* someone starting the cluster concurrently. Also, this is probably more
diff --git a/src/bin/pg_rewind/pg_rewind.h b/src/bin/pg_rewind/pg_rewind.h
index 9a981f7f246..9187c3bf72f 100644
--- a/src/bin/pg_rewind/pg_rewind.h
+++ b/src/bin/pg_rewind/pg_rewind.h
@@ -39,6 +39,7 @@ extern void findLastCheckpoint(const char *datadir, XLogRecPtr forkptr,
int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
XLogRecPtr *lastchkptredo,
+ uint32 *lastchkptdatachecksums,
const char *restoreCommand);
extern XLogRecPtr readOneRecord(const char *datadir, XLogRecPtr ptr,
int tliIndex, const char *restoreCommand);
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index e5d38fafb7d..7d07c757052 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -45,6 +45,8 @@ tests += {
't/019_standby_shutdown_catchup.pl',
't/020_cascade_divergence.pl',
't/021_rewind_divergent_transitions.pl',
+ 't/022_rewind_state.pl',
+ 't/023_rewind_standby_target.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/022_rewind_state.pl b/src/test/modules/test_checksums/t/022_rewind_state.pl
new file mode 100644
index 00000000000..b7adc968c4d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/022_rewind_state.pl
@@ -0,0 +1,133 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind checks the data checksum states of source and target.
+# Checksums enabled on the target, or at the point of divergence, with
+# a source running without them are refused: the rewound server would
+# verify checksums on blocks copied from a source that has none. The
+# opposite mismatch only warns; replay keeps the target's own state.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# Scenarios 1 and 2: checksums on from initdb. wal_log_hints keeps the
+# target eligible for pg_rewind after checksums are disabled on it.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b);
+
+# Failover to B, and divergence on A.
+$node_b->promote;
+$node_b->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_a->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_a->stop('fast');
+
+# Offline disable on the source only: enabled target, disabled source.
+$node_b->stop;
+$node_b->checksum_disable_offline;
+$node_b->start;
+test_checksum_state($node_b, 'off');
+
+my @rewind_cmd = (
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'));
+
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums are enabled on the target server but disabled on the source server/,
+ 'refuses an enabled target with a disabled source');
+
+# The refusal happens before any modification: the target still starts
+# on its own timeline.
+$node_a->start;
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'target untouched by the refused rewind');
+$node_a->stop('fast');
+
+# Scenario 2: disabling the target too makes the control files match,
+# but checksums were still enabled at the point of divergence, which is
+# the state replay on the rewound server would resume with.
+$node_a->checksum_disable_offline;
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums were enabled at the point of divergence but are disabled on the source server/,
+ 'refuses when checksums were enabled at the point of divergence');
+
+# Scenario 3: fresh pair without checksums; offline enable on the
+# source only. Allowed with a warning, and the rewound server keeps
+# checksums disabled.
+my $node_c = PostgreSQL::Test::Cluster->new('node_c');
+$node_c->init(allows_streaming => 1, no_data_checksums => 1);
+$node_c->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_c->start;
+$node_c->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_c->backup('backup');
+my $node_d = PostgreSQL::Test::Cluster->new('node_d');
+$node_d->init_from_backup($node_c, 'backup', has_streaming => 1);
+$node_d->start;
+$node_c->wait_for_catchup($node_d);
+
+$node_d->promote;
+$node_d->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_c->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_c->stop('fast');
+
+$node_d->stop;
+$node_d->checksum_enable_offline;
+$node_d->start;
+test_checksum_state($node_d, 'on');
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_c->data_dir,
+ '--source-server' => $node_d->connstr('postgres'),
+ ]);
+like(
+ $stderr,
+ qr/data checksums are disabled on the target server but enabled on the source server/,
+ 'warns for a disabled target with an enabled source');
+like($stderr, qr/Done!/, 'rewind completed despite the warning');
+
+# The rewound server follows D and keeps its own state.
+$node_c->append_conf('postgresql.conf', 'port = ' . $node_c->port);
+$node_c->enable_streaming($node_d);
+$node_c->set_standby_mode;
+$node_c->start;
+$node_d->wait_for_catchup($node_c);
+is($node_c->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+test_checksum_state($node_c, 'off');
+
+$node_c->stop;
+$node_d->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/023_rewind_standby_target.pl b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
new file mode 100644
index 00000000000..3f5a3c6be8b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
@@ -0,0 +1,141 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_rewind compares the data checksum state at the point of
+# divergence against the target as well as the source.
+#
+# A standby's checkpoints are written by its upstream primary, so a standby
+# whose checksums were disabled offline still has "on" checkpoints in its
+# WAL. Rewinding it onto a promoted sibling passes every control file
+# check (target off + source on only warns), yet recovery after the rewind
+# adopts the divergence checkpoint's state and the rewound server would
+# verify checksums over the pages it wrote while it was locally "off".
+# pg_rewind must refuse; after an offline enable on the target, which
+# rewrites all of its pages, the rewind goes through.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t0 AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+
+my $lagging = PostgreSQL::Test::Cluster->new('lagging');
+$lagging->init_from_backup($primary, 'backup', has_streaming => 1);
+$lagging->start;
+
+my $failover = PostgreSQL::Test::Cluster->new('failover');
+$failover->init_from_backup($primary, 'backup', has_streaming => 1);
+$failover->start;
+
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+
+test_checksum_state($primary, 'on');
+test_checksum_state($lagging, 'on');
+test_checksum_state($failover, 'on');
+
+# Offline-disable checksums on the lagging standby only. This is a
+# divergence 012_offline_standby.pl declares survivable: the standby keeps
+# its own state, warns, and stays readable.
+$lagging->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $lagging->data_dir);
+$lagging->start;
+test_checksum_state($lagging, 'off');
+
+# Give the lagging standby plenty of pages to write while it is locally
+# "off", and force a restartpoint so they reach disk without checksums.
+# The other standby replays the same WAL with checksums on.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+$lagging->safe_psql('postgres', "CHECKPOINT;");
+
+# Take the "failover" standby out of the picture. It is promoted from
+# this position later, so everything replayed from here on is the part of
+# the WAL that diverges and that pg_rewind will roll back. t1 is already
+# replayed everywhere, so its blocks are *not* rolled back.
+$failover->stop;
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$primary->wait_for_catchup($lagging);
+
+# Failover: promote the other standby, which forks the timeline behind the
+# replay position of the lagging standby.
+$primary->stop('fast');
+$failover->start;
+$failover->promote;
+$failover->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$failover->safe_psql('postgres', "CHECKPOINT;");
+
+# Re-attaching the lagging standby with pg_rewind must be refused: the
+# divergence checkpoint says "on" while the target is "off", so the rewound
+# server would verify checksums over the blocks it keeps.
+$lagging->stop;
+
+command_fails_like(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ],
+ qr/data checksums were enabled at the point of divergence but are disabled on the target server/,
+ 'pg_rewind refuses a target whose divergence checkpoint has checksums');
+
+# Enabling checksums offline on the target rewrites all of its pages, after
+# which the rewind is safe.
+system_or_bail('pg_checksums', '--enable', '--pgdata', $lagging->data_dir);
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ]);
+like($stderr, qr/Done!/, 'pg_rewind completes after the offline enable');
+
+$lagging->append_conf('postgresql.conf', 'port = ' . $lagging->port);
+$lagging->enable_streaming($failover);
+$lagging->set_standby_mode;
+$lagging->start;
+
+my ($rc, $out, $err) =
+ $lagging->psql('postgres',
+ "SELECT setting FROM pg_settings " . "WHERE name = 'data_checksums';");
+is($rc, 0, 'rewound standby accepts connections') or diag("stderr: $err");
+is($out, 'on', 'rewound standby resumes with checksums on');
+
+($rc, $out, $err) = $lagging->psql('postgres', "SELECT count(*) FROM t1;");
+is($rc, 0, 'blocks kept from the target stay readable')
+ or diag("stderr: $err");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($lagging->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the rewound standby');
+
+$lagging->stop('immediate');
+$failover->stop;
+done_testing();
--
2.55.0
[application/octet-stream] v10-0004-pg_combinebackup-Refuse-mixed-data-checksum-stat.patch (9.5K, ../../CAN4CZFOOHZmVnL-B2D+V7ODmVFgkYucgQD+-qegDTKZgJ3dtQg@mail.gmail.com/3-v10-0004-pg_combinebackup-Refuse-mixed-data-checksum-stat.patch)
download | inline diff:
From 8b5bf47c816abab56a967a3dfbad0001bfd055d9 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Tue, 25 Aug 2026 15:06:58 +0000
Subject: [PATCH v10 4/4] pg_combinebackup: Refuse mixed data checksum states
in a backup chain
check_control_files() detected a chain whose backups were taken under
different data checksum states, warned, and proceeded. The output
directory keeps the last backup's control file, so a full backup taken
with checksums off combined with an incremental taken after an offline
enable produces a cluster whose control file says "on" while most of
its blocks carry no checksums; it starts, and then every connection
dies on the first unchecksummed catalog page. An offline enable
between two backups of a chain is all it takes, since it rewrites
every page without logging anything, so the incremental backup does
not re-ship the pages.
Turn the warning into an error, matching what pg_rewind does for the
equivalent combinations. The check stays asymmetric on purpose: when
the last backup was taken with checksums off, stale checksums from an
earlier backup are never verified, and that chain remains usable.
---
doc/src/sgml/ref/pg_combinebackup.sgml | 19 +--
src/bin/pg_combinebackup/pg_combinebackup.c | 14 +-
src/test/modules/test_checksums/meson.build | 1 +
.../t/024_combinebackup_mixed.pl | 134 ++++++++++++++++++
4 files changed, 155 insertions(+), 13 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
diff --git a/doc/src/sgml/ref/pg_combinebackup.sgml b/doc/src/sgml/ref/pg_combinebackup.sgml
index 9a6d201e0b8..f5d7e177b36 100644
--- a/doc/src/sgml/ref/pg_combinebackup.sgml
+++ b/doc/src/sgml/ref/pg_combinebackup.sgml
@@ -306,18 +306,19 @@ PostgreSQL documentation
<para>
<literal>pg_combinebackup</literal> does not recompute page checksums when
- writing the output directory. Therefore, if any of the backups used for
- reconstruction were taken with checksums disabled, but the final backup was
- taken with checksums enabled, the resulting directory may contain pages
- with invalid checksums.
+ writing the output directory. It therefore refuses a chain in which some
+ of the backups used for reconstruction were taken with checksums disabled
+ but the final backup was taken with checksums enabled: the resulting
+ directory would contain pages with invalid checksums, and the cluster
+ would fail checksum verification as soon as it reads them. Take a new
+ full backup after enabling data checksums with
+ <xref linkend="app-pgchecksums"/>.
</para>
<para>
- To avoid this problem, taking a new full backup after changing the checksum
- state of the cluster using <xref linkend="app-pgchecksums"/> is
- recommended. Otherwise, you can disable and then optionally reenable
- checksums on the directory produced by <literal>pg_combinebackup</literal>
- in order to correct the problem.
+ The reverse case is accepted: when the final backup was taken with
+ checksums disabled, stale checksums copied from an earlier backup are
+ never verified.
</para>
</refsect1>
diff --git a/src/bin/pg_combinebackup/pg_combinebackup.c b/src/bin/pg_combinebackup/pg_combinebackup.c
index 86e2ee37c40..254a27b125b 100644
--- a/src/bin/pg_combinebackup/pg_combinebackup.c
+++ b/src/bin/pg_combinebackup/pg_combinebackup.c
@@ -670,13 +670,19 @@ check_control_files(int n_backups, char **backup_dirs)
pg_log_debug("system identifier is %" PRIu64, system_identifier);
/*
- * Warn the user if not all backups are in the same state with regards to
- * checksums.
+ * Reject a chain whose backups were not all in the same state with
+ * regards to checksums. The loop above only flags that when the last
+ * backup has checksums enabled: the blocks taken from the older backups
+ * have no checksums, and pg_combinebackup does not recompute them, so the
+ * combined cluster would fail verification as soon as it reads them. The
+ * other direction is harmless, since stale checksums are never verified
+ * while checksums are disabled.
*/
if (data_checksum_mismatch)
{
- pg_log_warning("only some backups have checksums enabled");
- pg_log_warning_hint("Disable, and optionally reenable, checksums on the output directory to avoid failures.");
+ pg_log_error("only some backups have checksums enabled");
+ pg_log_error_hint("Take a new full backup after changing the data checksum state with pg_checksums.");
+ exit(1);
}
return system_identifier;
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index 7d07c757052..7eccd5156b9 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -47,6 +47,7 @@ tests += {
't/021_rewind_divergent_transitions.pl',
't/022_rewind_state.pl',
't/023_rewind_standby_target.pl',
+ 't/024_combinebackup_mixed.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
new file mode 100644
index 00000000000..ce9dee9413b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_combinebackup refuses a chain whose backups were taken under
+# different data checksum states when the final state has them enabled.
+#
+# The output directory keeps the last backup's control file, so a full
+# backup taken with checksums off plus an incremental taken after an
+# offline enable would produce a directory whose control file says "on"
+# while most of its blocks have no checksums. An offline enable between
+# the two backups is enough to get there: it rewrites every page but logs
+# nothing, so the incremental backup does not re-ship the pages.
+#
+# The reverse order stays allowed: with the final backup taken with
+# checksums off, stale checksums from an earlier backup are never verified.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $node = PostgreSQL::Test::Cluster->new('node');
+$node->init(
+ no_data_checksums => 1,
+ has_archiving => 1,
+ allows_streaming => 1);
+$node->append_conf('postgresql.conf', 'summarize_wal = on');
+$node->append_conf('postgresql.conf', 'autovacuum = off');
+$node->start;
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+test_checksum_state($node, 'off');
+
+$node->backup('full');
+
+# Offline enable: rewrites every page, logs nothing.
+$node->stop;
+system_or_bail('pg_checksums', '--enable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'on');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full/backup_manifest',
+ ],
+ 'incremental backup with checksums on');
+
+# Combining the chain must be refused: the result would say "on" while the
+# blocks inherited from the full backup have no checksums.
+command_fails_like(
+ [
+ 'pg_combinebackup',
+ $node->backup_dir . '/full',
+ $node->backup_dir . '/incr',
+ '--output' => $node->backup_dir . '/combined',
+ ],
+ qr/only some backups have checksums enabled/,
+ 'pg_combinebackup refuses a chain crossing an offline enable');
+
+# The other direction: full backup with checksums on, offline disable, then
+# an incremental. The last backup wins, so the combined cluster comes up off.
+$node->backup('full2');
+
+$node->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr2',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full2/backup_manifest',
+ ],
+ 'incremental backup with checksums off');
+
+my $restored = PostgreSQL::Test::Cluster->new('restored');
+$restored->init_from_backup(
+ $node, 'incr2',
+ combine_with_prior => ['full2'],
+ has_restoring => 1,
+ standby => 0);
+
+$restored->start;
+
+my ($rc, $stdout, $stderr) = $restored->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'combined backup accepts connections') or diag("stderr: $stderr");
+is($stdout, 'off', 'combined backup follows the last backup state');
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t3;');
+is($rc, 0, 'blocks from the incremental backup are readable')
+ or diag("stderr: $stderr");
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t1;');
+is($rc, 0, 'blocks inherited from the full backup are readable')
+ or diag("stderr: $stderr");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($restored->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures in the combined backup');
+
+$restored->stop('immediate');
+$node->stop;
+
+done_testing();
--
2.55.0
[application/octet-stream] v10-0001-Do-not-adopt-data-checksum-state-from-another-no.patch (128.2K, ../../CAN4CZFOOHZmVnL-B2D+V7ODmVFgkYucgQD+-qegDTKZgJ3dtQg@mail.gmail.com/4-v10-0001-Do-not-adopt-data-checksum-state-from-another-no.patch)
download | inline diff:
From 6e299b84551cca91f282056e2f1c1e3d72a87cc8 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:02 +0000
Subject: [PATCH v10 1/4] Do not adopt data checksum state from another node
during replay
Offline pg_checksums changes are local to one node, but replay adopted
the data checksum state carried by checkpoint records unconditionally.
After an offline enable on the primary, a standby whose pages were
never rewritten started verifying checksums it does not have; after an
offline change on a standby, the next replayed checkpoint silently
reverted it.
To fix, make the control file track this node's state alone, and have
replay cross-check the replayed state against it instead of adopting
it, warning once per divergent value and reporting when the states
agree again. pg_control gains a watermark, normally the end LSN of
the newest XLOG2_CHECKSUMS record the node has written or applied, so
that replay can skip transition records whose effect the control file
already contains, and a flag marking a state last written by
pg_checksums, which recovery must never overwrite with a replayed one.
Since the control file may only claim "on" once every page on disk
carries a checksum, persisting the state is tied to flushes:
XLOG2_CHECKSUMS replay persists every state but "on" immediately and
leaves "on" to the next restartpoint, checkpoints and restartpoints
only persist a state their flush ran under from beginning to end, and
transitions publish their state under the new
DataChecksumTransitionLock so that the states carried by WAL records
match their WAL order. The comments in xlog.c spell out the
individual rules.
Recovery from a base backup is the exception to not adopting: its
control file was copied at an arbitrary moment, so the state carried
by the starting checkpoint is the one the WAL from there on was
written under. pg_rewind keeps the target's own state and clamps the
watermark to the divergence point, since a watermark set on the
target's abandoned fork could numerically cover transition records the
source wrote after the divergence.
Bump PG_CONTROL_VERSION.
Document the offline procedure for replication setups, the lockstep
one: stop all nodes, run pg_checksums on each of them, and only then
restart. An offline change writes no WAL and has no ordering against
WAL a node has not replayed yet, so a node must not stop before
replaying all WAL of its upstream, or the change is overridden on
restart.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Author: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
---
doc/src/sgml/ref/pg_checksums.sgml | 54 +-
doc/src/sgml/wal.sgml | 21 +
src/backend/access/transam/xlog.c | 547 ++++++++++++++++--
.../utils/activity/wait_event_names.txt | 1 +
src/bin/pg_checksums/pg_checksums.c | 10 +
src/bin/pg_controldata/pg_controldata.c | 4 +
src/bin/pg_resetwal/pg_resetwal.c | 7 +
src/bin/pg_rewind/pg_rewind.c | 35 +-
src/bin/pg_upgrade/controldata.c | 2 +-
src/include/catalog/pg_control.h | 29 +-
src/include/storage/lwlocklist.h | 1 +
src/test/modules/test_checksums/Makefile | 2 +-
src/test/modules/test_checksums/meson.build | 10 +
.../test_checksums/t/012_offline_standby.pl | 289 +++++++++
.../modules/test_checksums/t/013_rewind.pl | 201 +++++++
.../modules/test_checksums/t/014_lockstep.pl | 182 ++++++
.../t/015_standby_crash_after_disable.pl | 125 ++++
.../t/016_promote_enable_crash.pl | 134 +++++
.../test_checksums/t/017_restartpoint_race.pl | 138 +++++
.../t/018_enable_crash_windows.pl | 496 ++++++++++++++++
.../t/019_standby_shutdown_catchup.pl | 126 ++++
.../t/020_cascade_divergence.pl | 115 ++++
.../t/021_rewind_divergent_transitions.pl | 193 ++++++
23 files changed, 2643 insertions(+), 79 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/012_offline_standby.pl
create mode 100644 src/test/modules/test_checksums/t/013_rewind.pl
create mode 100644 src/test/modules/test_checksums/t/014_lockstep.pl
create mode 100644 src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
create mode 100644 src/test/modules/test_checksums/t/016_promote_enable_crash.pl
create mode 100644 src/test/modules/test_checksums/t/017_restartpoint_race.pl
create mode 100644 src/test/modules/test_checksums/t/018_enable_crash_windows.pl
create mode 100644 src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
create mode 100644 src/test/modules/test_checksums/t/020_cascade_divergence.pl
create mode 100644 src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index ae66fad3f0f..969497586f4 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -242,15 +242,51 @@ PostgreSQL documentation
data directory must not be started or else data loss may occur.
</para>
<para>
- When using a replication setup with tools which perform direct copies
- of relation file blocks (for example <xref linkend="app-pgrewind"/>),
- enabling or disabling checksums can lead to page corruptions in the
- shape of incorrect checksums if the operation is not done consistently
- across all nodes. When enabling or disabling checksums in a replication
- setup, it is thus recommended to stop all the clusters before switching
- them all consistently. Destroying all standbys, performing the operation
- on the primary and finally recreating the standbys from scratch is also
- safe.
+ Enabling or disabling checksums with
+ <application>pg_checksums</application> changes only the local data
+ directory; the new state does not propagate over replication. In a
+ replication setup the same change must be applied to every node: stop all
+ nodes, run <application>pg_checksums</application> on each of them, and
+ only then restart them. Before stopping a standby, make sure it has
+ replayed all WAL of the primary: stop the primary first, read its
+ <quote>Latest checkpoint location</quote> with
+ <xref linkend="app-pgcontroldata"/>, and check that
+ <function>pg_last_wal_replay_lsn()</function> on the standby has
+ advanced past it. Comparing the replay position with
+ <function>pg_last_wal_receive_lsn()</function> is not enough, as it
+ only shows that the WAL the standby received has been replayed. Tools
+ that copy relation file blocks directly between nodes, such as
+ <xref linkend="app-pgrewind"/>, likewise require both nodes to be in
+ the same data checksum state.
+ </para>
+ <para>
+ The replay requirement exists because an offline change is recorded
+ only in the control file and has no defined ordering against WAL the
+ node has not replayed yet; see
+ <xref linkend="checksums-offline-enable-disable"/>. A node stopped
+ before replaying an online checksum state change applies that change
+ when it is restarted, overriding the offline change, and the states of
+ the nodes silently diverge until a later checkpoint record triggers
+ the warning described below. Because of this it is best not to mix
+ the two mechanisms: change the state of a replication setup either
+ with the offline procedure above or with an online transition, and
+ make sure the previous change has reached every node before starting
+ the next one.
+ </para>
+ <para>
+ If the change is applied inconsistently, each node keeps its own state,
+ and a standby logs a warning when the state recorded in the replayed WAL
+ differs from its own. The same warning can appear transiently while a
+ standby catches up over WAL written before a consistent change; it stops
+ once a checkpoint record carrying the new state has been replayed.
+ </para>
+ <para>
+ A standby whose data directory was never checksummed must not have
+ checksums enabled by catching up this way. Converge the cluster by
+ running <application>pg_checksums</application> on it while stopped, by
+ enabling checksums online, or by recreating it from a base backup. Note
+ that an online enable only starts from a primary whose checksums are off,
+ so if they are already enabled there, disable them online first.
</para>
<para>
If <application>pg_checksums</application> is aborted or killed while
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index ec62d17fbcc..d88a832c2a8 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -317,6 +317,27 @@
verify checksums, on an offline cluster.
</para>
+ <para>
+ An offline change provides durability differently from an
+ <link linkend="checksums-online-enable-disable">online change</link>.
+ An online transition is WAL-logged: it is ordered against all other
+ WAL records, it is replayed after a crash, and it propagates to
+ standbys. An offline change is recorded only in the cluster's
+ control file: it writes no WAL, it is invisible to replication, and
+ it has no defined ordering against WAL the node has not replayed
+ yet. When a node later replays WAL that contains an online checksum
+ state change, that change takes effect on the node even if it was
+ written before the offline change was made.
+ </para>
+
+ <para>
+ An offline change only affects the data directory it is run on; the
+ new state does not propagate over replication. In a replication setup
+ the same change must be applied to all nodes while all of them are
+ stopped, as described in <xref linkend="app-pgchecksums"/>. A standby
+ whose state diverges logs a warning but keeps its local setting.
+ </para>
+
</sect2>
<sect2 id="checksums-online-enable-disable" xreflabel="Online Enabling of Checksums">
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index de4c96e135f..7487ca509f3 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -556,9 +556,17 @@ typedef struct XLogCtlData
*/
XLogRecPtr lastFpwDisableRecPtr;
- /* last data_checksum_version we've seen */
+ /* current data checksum state of this node */
uint32 data_checksum_version;
+ /*
+ * In-memory copies of ControlFile->data_checksum_lsn and
+ * ControlFile->data_checksum_is_local, see there. Updated together with
+ * data_checksum_version under info_lck.
+ */
+ XLogRecPtr data_checksum_lsn;
+ bool data_checksum_is_local;
+
slock_t info_lck; /* locks shared variables shown above */
/*
@@ -690,6 +698,14 @@ static ChecksumStateType LocalDataChecksumState = 0;
*/
int data_checksums = 0;
+/*
+ * Whether replay of the next checkpoint-family record must adopt the data
+ * checksum state it carries. Set when recovery starts from a base backup,
+ * where the state at the redo point takes precedence over the control file
+ * copied with the backup later.
+ */
+static bool adoptChecksumStateFromNextCheckpoint = false;
+
/* For WALInsertLockAcquire/Release functions */
static int MyLockNo = 0;
static bool holdingAllLocks = false;
@@ -730,6 +746,8 @@ static void ValidateXLOGDirectoryStructure(void);
static void CleanupBackupHistory(void);
static void UpdateMinRecoveryPoint(XLogRecPtr lsn, bool force);
static bool PerformRecoveryXLogAction(void);
+static void CheckReplayedDataChecksumState(uint32 replayed_version);
+static void AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn);
static void InitControlFile(uint64 sysidentifier, uint32 data_checksum_version);
static void WriteControlFile(void);
static void ReadControlFile(void);
@@ -757,7 +775,7 @@ static void WALInsertLockAcquireExclusive(void);
static void WALInsertLockRelease(void);
static void WALInsertLockUpdateInsertingAt(XLogRecPtr insertingAt);
-static void XLogChecksums(uint32 new_type);
+static XLogRecPtr XLogChecksums(uint32 new_type);
/*
* Insert an XLOG record represented by an already-constructed chain of data
@@ -4774,6 +4792,7 @@ void
SetDataChecksumsOnInProgress(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
/*
* The state transition is performed in a critical section with
@@ -4782,14 +4801,12 @@ SetDataChecksumsOnInProgress(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4827,6 +4844,8 @@ void
SetDataChecksumsOn(void)
{
uint64 barrier;
+ bool persist;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4847,23 +4866,11 @@ SetDataChecksumsOn(void)
SpinLockRelease(&XLogCtl->info_lck);
INJECTION_POINT("datachecksums-enable-checksums-delay", NULL);
+ INJECTION_POINT_LOAD("datachecksums-on-before-publish");
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_VERSION);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- SpinLockRelease(&XLogCtl->info_lck);
-
- /*
- * Update the controlfile before waiting since if we have an immediate
- * shutdown while waiting we want to come back up with checksums enabled.
- */
- LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- UpdateControlFile();
- LWLockRelease(ControlFileLock);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_VERSION);
barrier = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_CHECKSUM_ON);
@@ -4873,6 +4880,39 @@ SetDataChecksumsOn(void)
INJECTION_POINT("datachecksums-on-before-checkpoint", NULL);
RequestCheckpoint(CHECKPOINT_FORCE | CHECKPOINT_WAIT | CHECKPOINT_FAST);
+
+ INJECTION_POINT("datachecksums-on-after-checkpoint", NULL);
+
+ /*
+ * Persist "on" only now that the checkpoint has flushed the pages the
+ * transition rewrote. Crash recovery initializes verification from the
+ * control file but resumes from a checkpoint that can predate the
+ * transition, so an "on" persisted earlier would have replay verify pages
+ * whose rewrite never reached disk: the pages the worker found in shared
+ * buffers are not written back by its ring buffer.
+ *
+ * The checkpoint above normally persists the state itself; this covers the
+ * case where it started before the record written above and left the field
+ * alone. Crashing before this point is safe, as replay then re-establishes
+ * "on" from the full page images of the rewrite. Skip the write if the
+ * state moved on meanwhile, since whatever moved it persists its own.
+ * Compare the watermark rather than the state: a state comparison could
+ * not tell our transition from a later round trip back to "on".
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ persist = (XLogCtl->data_checksum_lsn == recptr);
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (persist)
+ {
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+ }
+
WaitForProcSignalBarrier(barrier);
}
@@ -4893,6 +4933,7 @@ void
SetDataChecksumsOff(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4918,14 +4959,12 @@ SetDataChecksumsOff(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4956,14 +4995,12 @@ SetDataChecksumsOff(void)
/* Ensure that we don't incur a checkpoint during disabling checksums */
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -5001,6 +5038,131 @@ SetLocalDataChecksumState(uint32 data_checksum_version)
data_checksums = data_checksum_version;
}
+/*
+ * Cross-check the data checksum state carried by a replayed checkpoint record
+ * against the state of this node.
+ *
+ * The state in the record belongs to whichever node wrote the WAL, and must
+ * not be adopted: an offline state change made with pg_checksums on one node
+ * of a replication set generates no WAL, and adopting would leak it into the
+ * other nodes through replay. Backup label recovery is the one exception,
+ * see AdoptReplayedDataChecksumState(). XLOG_CHECKPOINT_ONLINE needs no call
+ * here, as an online checkpoint's state already traveled in the preceding
+ * XLOG_CHECKPOINT_REDO record.
+ *
+ * Only archive recovery can see a lasting mismatch, as only there can the WAL
+ * and the control file come from different nodes or different times. In
+ * crash recovery a mismatch means replay resumed from a restartpoint
+ * predating an already-applied XLOG2_CHECKSUMS record, and replaying forward
+ * re-establishes the same state.
+ */
+static void
+CheckReplayedDataChecksumState(uint32 replayed_version)
+{
+ /*
+ * Warn once per remote value, so a lasting mismatch does not flood the
+ * log. Matching states re-arm the warning. Backend-local state is
+ * enough: replay only runs in the startup process, and restarting it
+ * re-arms as well.
+ */
+ static uint32 last_warned_version = PG_UINT32_MAX;
+ uint32 local_version;
+
+ if (!ArchiveRecoveryRequested)
+ return;
+
+ /*
+ * Re-replayed WAL below the consistency point was already cross-checked
+ * before minRecoveryPoint was last persisted, and the persisted state can
+ * legitimately be newer than what checkpoint records there carry:
+ * XLOG2_CHECKSUMS replay persists most states ahead of the restartpoint
+ * horizon. In particular the checkpoint record recovery restarts from is
+ * such a re-replay.
+ */
+ if (!reachedConsistency)
+ return;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ local_version = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (replayed_version == local_version)
+ {
+ /*
+ * Report convergence if this process warned before. Nothing else
+ * tells the operator that running pg_checksums on the other nodes, or
+ * a rebuild, took effect.
+ */
+ if (last_warned_version != PG_UINT32_MAX)
+ ereport(LOG,
+ errmsg("data checksum state \"%s\" of this node now agrees with the replayed WAL",
+ get_checksum_state_string(local_version)));
+
+ last_warned_version = PG_UINT32_MAX;
+ return;
+ }
+
+ /* the nodes legitimately differ while an online transition runs */
+ if (replayed_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ replayed_version == PG_DATA_CHECKSUM_INPROGRESS_OFF ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ return;
+
+ if (replayed_version == last_warned_version)
+ return;
+ last_warned_version = replayed_version;
+
+ ereport(WARNING,
+ errmsg("data checksum state \"%s\" of this node does not match the state \"%s\" in the replayed WAL",
+ get_checksum_state_string(local_version),
+ get_checksum_state_string(replayed_version)),
+ errdetail("The data checksum state was most likely changed with pg_checksums on another node."),
+ errhint("Apply the same change with pg_checksums on the primary and all standby servers, or rebuild this server from a base backup."));
+}
+
+/*
+ * Adopt the data checksum state found at the redo point of backup label
+ * recovery. Persist it immediately so that a crash before the first
+ * restartpoint does not resurrect the state copied with the backup; a crash
+ * at this point restarts from the same redo point, so the control file does
+ * not run ahead of the replay position. If the value is unchanged the
+ * control file already carries it, so both the barrier and the persist are
+ * skipped.
+ *
+ * lsn is the location of the checkpoint-family record the state was taken
+ * from and becomes the new watermark: the adopted state covers everything
+ * below the redo point, which replay never revisits.
+ */
+static void
+AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn)
+{
+ bool changed = false;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (XLogCtl->data_checksum_version != new_version)
+ {
+ XLogCtl->data_checksum_version = new_version;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(new_version);
+ changed = true;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (!changed)
+ return;
+
+ EmitAndWaitDataChecksumsBarrier(new_version);
+
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = new_version;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+}
+
/* guc hook */
const char *
show_data_checksums(void)
@@ -5454,6 +5616,8 @@ XLOGShmemInit(void *arg)
/* Use the checksum info from control file */
XLogCtl->data_checksum_version = ControlFile->data_checksum_version;
+ XLogCtl->data_checksum_lsn = ControlFile->data_checksum_lsn;
+ XLogCtl->data_checksum_is_local = ControlFile->data_checksum_is_local;
SetLocalDataChecksumState(XLogCtl->data_checksum_version);
SpinLockInit(&XLogCtl->Insert.insertpos_lck);
@@ -6031,6 +6195,52 @@ StartupXLOG(void)
SetCommitTsLimit(checkPoint.oldestCommitTsXid,
checkPoint.newestCommitTsXid);
+ /*
+ * When recovery starts from a base backup, the control file was copied at
+ * an arbitrary moment and its data checksum state may differ from the
+ * state at the redo point, which is what the WAL from there on was
+ * written under. Adopt the state of the starting checkpoint: a shutdown
+ * checkpoint is not replayed, so take it from the record read above; the
+ * redo point of an online checkpoint is its CHECKPOINT_REDO record, so
+ * let the replay of that record adopt it. Check backupStartPoint in
+ * addition to the label: on a crash restart during backup recovery the
+ * label file is already renamed away, but the start point persists until
+ * the backup end record.
+ *
+ * Not for a base backup taken from a standby, though. Its starting
+ * checkpoint is the standby's last restartpoint, a record written by the
+ * upstream primary, whose state is not the one the copied files were
+ * written under. The copied control file is already correct: a standby
+ * persists its state only at restartpoint horizons and never claims
+ * more than what reached disk. Such backups are recognized by
+ * backupEndPoint together with backupEndRequired; backupEndPoint is only
+ * set for "BACKUP FROM: standby" labels and persists across a crash
+ * restart. pg_rewind writes a standby label as well, but no
+ * backupEndPoint, and its recovery keeps adopting: the control file it
+ * installs carries the target's own checksum state, which can lag the
+ * redo point of the last common checkpoint the same way a restartpoint
+ * horizon can.
+ *
+ * Never adopt over a state the control file's watermark or local flag
+ * marks as newer than the starting checkpoint. A pg_checksums change is
+ * local to the node and generates no WAL, so nothing in the replayed WAL
+ * could ever restore it once overwritten; and a watermark above the redo
+ * point means the control file already contains the effect of every
+ * transition record up to there.
+ */
+ if ((haveBackupLabel || XLogRecPtrIsValid(ControlFile->backupStartPoint)) &&
+ !(XLogRecPtrIsValid(ControlFile->backupEndPoint) &&
+ ControlFile->backupEndRequired) &&
+ !ControlFile->data_checksum_is_local &&
+ checkPoint.redo > ControlFile->data_checksum_lsn)
+ {
+ if (wasShutdown)
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ checkPoint.redo);
+ else
+ adoptChecksumStateFromNextCheckpoint = true;
+ }
+
/*
* Clear out any old relcache cache files. This is *necessary* if we do
* any WAL replay, since that would probably result in the cache files
@@ -6633,11 +6843,7 @@ StartupXLOG(void)
if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
ereport(WARNING,
@@ -6654,11 +6860,7 @@ StartupXLOG(void)
else if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
}
@@ -6814,6 +7016,23 @@ static bool
PerformRecoveryXLogAction(void)
{
bool promoted = false;
+ bool flushForChecksums;
+ uint32 checksum_state;
+
+ /*
+ * The end-of-recovery record persists the data checksum state without
+ * flushing the buffer pool, but the control file may only claim "on" once
+ * every page on disk carries a checksum. If replay entered that state
+ * without a restartpoint following it, the pages rewritten by the
+ * transition are still only in the buffer pool, so take the full
+ * checkpoint below instead of the lightweight record.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ flushForChecksums = (checksum_state == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile->data_checksum_version != checksum_state);
/*
* Perform a checkpoint to update all our recovery activity to disk.
@@ -6829,7 +7048,7 @@ PerformRecoveryXLogAction(void)
* fully out of recovery mode and already accepting queries.
*/
if (ArchiveRecoveryRequested && IsUnderPostmaster &&
- PromoteIsTriggered())
+ PromoteIsTriggered() && !flushForChecksums)
{
promoted = true;
@@ -7436,6 +7655,7 @@ CreateCheckPoint(int flags)
uint32 freespace;
XLogRecPtr PriorRedoPtr;
XLogRecPtr last_important_lsn;
+ XLogRecPtr checksumLsn;
VirtualTransactionId *vxids;
int nvxids;
int oldXLogAllowed = 0;
@@ -7548,11 +7768,14 @@ CreateCheckPoint(int flags)
checkPoint.wal_level = wal_level;
/*
- * Get the current data_checksum_version value from xlogctl, valid at the
- * time of the checkpoint.
+ * Get the current data_checksum_version value from xlogctl. This is
+ * final only for a shutdown checkpoint, where no concurrent transition
+ * is possible; an online checkpoint resamples it together with the redo
+ * record below.
*/
SpinLockAcquire(&XLogCtl->info_lck);
checkPoint.dataChecksumState = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
if (shutdown)
@@ -7609,10 +7832,21 @@ CreateCheckPoint(int flags)
{
xl_checkpoint_redo redo_rec;
+ /*
+ * Sample the data checksum state and insert the redo record under
+ * DataChecksumTransitionLock, so that a concurrent transition cannot
+ * insert its XLOG2_CHECKSUMS record between the sampling and the
+ * insertion below. Without this, the redo record could follow the
+ * transition record in WAL while carrying the pre-transition state,
+ * and recovery resuming here would never learn about the transition.
+ * See XLogChecksums().
+ */
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
WALInsertLockAcquire();
redo_rec.wal_level = wal_level;
SpinLockAcquire(&XLogCtl->info_lck);
redo_rec.data_checksum_version = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
WALInsertLockRelease();
@@ -7620,6 +7854,14 @@ CreateCheckPoint(int flags)
XLogBeginInsert();
XLogRegisterData(&redo_rec, sizeof(xl_checkpoint_redo));
(void) XLogInsert(RM_XLOG_ID, XLOG_CHECKPOINT_REDO);
+ LWLockRelease(DataChecksumTransitionLock);
+
+ /*
+ * The checkpoint record must carry the same state as the redo record
+ * just inserted: the sample taken before redo determination can be
+ * stale by now, and the pair would otherwise disagree.
+ */
+ checkPoint.dataChecksumState = redo_rec.data_checksum_version;
/*
* XLogInsertRecord will have updated XLogCtl->Insert.RedoRecPtr in
@@ -7823,6 +8065,40 @@ CreateCheckPoint(int flags)
ControlFile->minRecoveryPoint = InvalidXLogRecPtr;
ControlFile->minRecoveryPointTLI = 0;
+ /*
+ * Persist the data checksum state this node runs under. Only the
+ * top-level field tracks this node; ControlFile->checkPointCopy above is
+ * a historical record used to resume replay.
+ *
+ * checkPoint.dataChecksumState was sampled under
+ * DataChecksumTransitionLock together with the redo record, so it is the
+ * state in effect at the redo point. If it was "on",
+ * the XLOG2_CHECKSUMS record announcing that precedes the redo point and
+ * every page the transition rewrote was dirtied before it, so
+ * CheckPointGuts() has just written all of them out. Recording the state
+ * here is what keeps a finished transition from being resolved as
+ * interrupted when this checkpoint is the one crash recovery resumes
+ * from: replay never sees the record announcing it.
+ *
+ * Persist it only if the state did not change while the flush was in
+ * progress. If it changed in between, the pages written out straddle two
+ * states, and the newer one could claim checksums that pages already on
+ * disk do not carry; leave the field to the next checkpoint then, the
+ * transition itself has already persisted every state that is safe
+ * without a flush.
+ * Compare the watermark rather than the state: record positions are
+ * unique, so a full round trip back to the sampled state cannot alias,
+ * while its flushed pages straddle the intermediate states all the same.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksumLsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checkPoint.dataChecksumState;
+ ControlFile->data_checksum_lsn = checksumLsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
/*
* Persist unloggedLSN value. It's reset on crash recovery, so this goes
* unused on non-shutdown checkpoints, but seems useful to store it always
@@ -7967,9 +8243,11 @@ CreateEndOfRecoveryRecord(void)
ControlFile->minRecoveryPoint = recptr;
ControlFile->minRecoveryPointTLI = xlrec.ThisTimeLineID;
- /* start with the latest checksum version (as of the end of recovery) */
+ /* persist the data checksum state this node ended recovery with */
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
@@ -8174,6 +8452,9 @@ CreateRestartPoint(int flags)
XLogRecPtr endptr;
XLogSegNo _logSegNo;
TimestampTz xtime;
+ uint32 checksum_state;
+ XLogRecPtr checksum_lsn;
+ bool checksum_is_local;
/* Concurrent checkpoint/restartpoint cannot happen */
Assert(!IsUnderPostmaster || MyBackendType == B_CHECKPOINTER);
@@ -8220,8 +8501,45 @@ CreateRestartPoint(int flags)
UpdateMinRecoveryPoint(InvalidXLogRecPtr, true);
if (flags & CHECKPOINT_IS_SHUTDOWN)
{
+ bool catchUpChecksums;
+
+ /*
+ * There is no new restartpoint to persist the data checksum state
+ * with, but a cleanly stopped node should not leave the control
+ * file behind the state replay reached: pg_checksums and
+ * pg_rewind read it, and an in-progress state there makes them
+ * refuse to run. Catching it up needs the same guarantee a
+ * restartpoint gives: that every page on disk carries a checksum,
+ * so flush the buffer pool first. Only a transition to
+ * "on" that no restartpoint followed can get here; the other
+ * states are already persisted by XLOG2_CHECKSUMS replay. Replay
+ * has ended by now, so the state cannot change under us.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ checksum_is_local = XLogCtl->data_checksum_is_local;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ catchUpChecksums =
+ (checksum_lsn != ControlFile->data_checksum_lsn &&
+ XLogRecPtrIsValid(lastCheckPointRecPtr));
+
+ if (catchUpChecksums)
+ {
+ MemSet(&CheckpointStats, 0, sizeof(CheckpointStats));
+ CheckpointStats.ckpt_start_t = GetCurrentTimestamp();
+ CheckPointGuts(lastCheckPoint.redo, flags);
+ }
+
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
+ if (catchUpChecksums)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = checksum_is_local;
+ }
UpdateControlFile();
LWLockRelease(ControlFileLock);
}
@@ -8262,6 +8580,17 @@ CreateRestartPoint(int flags)
/* Update the process title */
update_checkpoint_display(flags, true, false);
+ /*
+ * Note the data checksum state the flush below starts under. Replay runs
+ * concurrently and can change the state while the flush is in progress,
+ * in which case the flush covers pages written under both states; see
+ * where the state is persisted further down.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
CheckPointGuts(lastCheckPoint.redo, flags);
/*
@@ -8321,8 +8650,26 @@ CreateRestartPoint(int flags)
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
}
- /* we shall start with the latest checksum version */
- ControlFile->data_checksum_version = lastCheckPoint.dataChecksumState;
+ /*
+ * Persist the data checksum state of this node. Not the state of the
+ * replayed checkpoint: that one belongs to the node that wrote it and
+ * may differ after an offline change on either side.
+ * ControlFile->checkPointCopy above keeps the replayed value on
+ * purpose, being a historical record used to resume replay rather
+ * than a tracker of node state.
+ *
+ * Persist only if the flush above ran under one state throughout; see
+ * CreateCheckPoint() for why, including why this compares the
+ * watermark and not the state.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksum_lsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
}
@@ -8763,9 +9110,21 @@ XLogReportParameters(void)
}
/*
- * Log the new state of checksums
+ * Log and publish the new state of checksums
+ *
+ * Inserting the record and publishing the new state must be atomic with
+ * respect to a checkpoint sampling the state for its XLOG_CHECKPOINT_REDO
+ * record: without that, a checkpoint could read the old state after the
+ * record is already in WAL and insert a redo record that both precedes the
+ * transition in WAL order and carries the pre-transition state. Recovery
+ * resuming from such a redo point would never replay the transition record
+ * and resolve the finished transition as interrupted.
+ * DataChecksumTransitionLock serializes the two; see CreateCheckPoint().
+ *
+ * Returns the end LSN of the inserted record, which the caller persists
+ * together with the new state as the data checksum watermark.
*/
-static void
+static XLogRecPtr
XLogChecksums(uint32 new_type)
{
xl_checksum_state xlrec;
@@ -8773,12 +9132,28 @@ XLogChecksums(uint32 new_type)
xlrec.new_checksum_state = new_type;
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
+
XLogBeginInsert();
XLogRegisterData((char *) &xlrec, sizeof(xl_checksum_state));
recptr = XLogInsert(RM_XLOG2_ID, XLOG2_CHECKSUMS);
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, recptr);
+
+ /* only loaded by SetDataChecksumsOn(), a no-op for the other callers */
+ INJECTION_POINT_CACHED("datachecksums-on-before-publish", NULL);
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ XLogCtl->data_checksum_version = new_type;
+ XLogCtl->data_checksum_lsn = recptr;
+ XLogCtl->data_checksum_is_local = false;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockRelease(DataChecksumTransitionLock);
+
XLogFlush(recptr);
+
+ return recptr;
}
/*
@@ -8966,11 +9341,19 @@ xlog_redo(XLogReaderState *record)
/* ControlFile->checkPointCopy always tracks the latest ckpt XID */
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->checkPointCopy.nextXid = checkPoint.nextXid;
- ControlFile->data_checksum_version = checkPoint.dataChecksumState;
UpdateControlFile();
LWLockRelease(ControlFileLock);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(checkPoint.dataChecksumState);
+
/*
* We should've already switched to the new TLI before replaying this
* record.
@@ -9206,19 +9589,17 @@ xlog_redo(XLogReaderState *record)
else if (info == XLOG_CHECKPOINT_REDO)
{
xl_checkpoint_redo redo_rec;
- bool new_state = false;
memcpy(&redo_rec, XLogRecGetData(record), sizeof(xl_checkpoint_redo));
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = redo_rec.data_checksum_version;
- SetLocalDataChecksumState(redo_rec.data_checksum_version);
- if (redo_rec.data_checksum_version != ControlFile->data_checksum_version)
- new_state = true;
- SpinLockRelease(&XLogCtl->info_lck);
-
- if (new_state)
- EmitAndWaitDataChecksumsBarrier(redo_rec.data_checksum_version);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(redo_rec.data_checksum_version,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(redo_rec.data_checksum_version);
}
else if (info == XLOG_LOGICAL_DECODING_STATUS_CHANGE)
{
@@ -9280,25 +9661,63 @@ xlog2_redo(XLogReaderState *record)
{
xl_checksum_state state;
XLogRecPtr lsn = record->EndRecPtr;
+ XLogRecPtr watermark;
memcpy(&state, XLogRecGetData(record), sizeof(xl_checksum_state));
+ SpinLockAcquire(&XLogCtl->info_lck);
+ watermark = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ /*
+ * Skip records this node has already applied. The control file
+ * carries the watermark, so this holds across restarts: recovery
+ * resuming below a record whose effect the control file already
+ * contains must not re-apply it, or it would revert a state change
+ * made with pg_checksums in between, which moves the state without
+ * writing any record of its own.
+ */
+ if (lsn <= watermark)
+ return;
+
/* advertise the location before the new state becomes visible */
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, lsn);
SpinLockAcquire(&XLogCtl->info_lck);
XLogCtl->data_checksum_version = state.new_checksum_state;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(state.new_checksum_state);
SpinLockRelease(&XLogCtl->info_lck);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = state.new_checksum_state;
+
+ /*
+ * Persist the new state, except when it is "on". Only "on" verifies
+ * checksums during reads, and between the last restartpoint and this
+ * record there may be pages on disk flushed under the old state; a
+ * crash-restart initializes verification from the control file and
+ * replay reads those pages back, so the control file may only say
+ * "on" once everything written under the transition has been flushed,
+ * as restartpoints and the end of recovery do.
+ * The opposite direction cannot wait for the restartpoint: once this
+ * record is replayed, evicted pages are written without checksums,
+ * and a control file still saying "on" would fail verification on
+ * exactly those pages after a crash.
+ */
+ if (state.new_checksum_state != PG_DATA_CHECKSUM_VERSION)
+ {
+ ControlFile->data_checksum_version = state.new_checksum_state;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ }
/*
* Update minRecoveryPoint to ensure that if recovery is aborted, we
* recover back up to this point before allowing hot standby again.
- * The new state is durable in pg_control while its location is only
- * tracked in shared memory; a standby becoming consistent below this
- * record would let base backups resume checksum verification with the
+ * The change location is only tracked in shared memory and is lost
+ * over a restart; a standby becoming consistent below this record
+ * would let base backups resume checksum verification with the
* location unknown. The local copies cannot be updated as long as
* crash recovery is happening and we expect all the WAL to be
* replayed.
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 256b3a3c02e..77fa543f9d8 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -371,6 +371,7 @@ WaitLSN "Waiting to read or update shared Wait-for-LSN state."
LogicalDecodingControl "Waiting to read or update logical decoding status information."
DataChecksumsWorker "Waiting for data checksums worker."
AioWorkerControl "Waiting to update AIO worker information."
+DataChecksumTransition "Waiting for a data checksum state transition to be written to WAL."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 3b3ae23f1a6..8c4ae59222d 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -648,6 +648,16 @@ main(int argc, char *argv[])
ControlFile->data_checksum_version =
(mode == PG_MODE_ENABLE) ? PG_DATA_CHECKSUM_VERSION : PG_DATA_CHECKSUM_OFF;
+ /*
+ * Mark the state as changed locally, without a WAL record. Recovery
+ * then does not let a replayed checkpoint overwrite it, as no record
+ * could restore the change afterwards. The watermark is left alone:
+ * XLOG2_CHECKSUMS records at or below it stay covered, while records
+ * above it, which this node has not applied yet, still take effect
+ * on replay no matter when they were written.
+ */
+ ControlFile->data_checksum_is_local = true;
+
if (do_sync)
{
pg_log_info("syncing data directory");
diff --git a/src/bin/pg_controldata/pg_controldata.c b/src/bin/pg_controldata/pg_controldata.c
index 6fc87ed114d..c363ce3dbb7 100644
--- a/src/bin/pg_controldata/pg_controldata.c
+++ b/src/bin/pg_controldata/pg_controldata.c
@@ -349,6 +349,10 @@ main(int argc, char *argv[])
(ControlFile->float8ByVal ? _("by value") : _("by reference")));
printf(_("Data page checksum version: %u\n"),
ControlFile->data_checksum_version);
+ printf(_("Data checksum watermark: %X/%08X\n"),
+ LSN_FORMAT_ARGS(ControlFile->data_checksum_lsn));
+ printf(_("Data checksum state is node-local: %s\n"),
+ (ControlFile->data_checksum_is_local ? _("yes") : _("no")));
printf(_("Default char data signedness: %s\n"),
(ControlFile->default_char_signedness ? _("signed") : _("unsigned")));
printf(_("Mock authentication nonce: %s\n"),
diff --git a/src/bin/pg_resetwal/pg_resetwal.c b/src/bin/pg_resetwal/pg_resetwal.c
index 1542a56ca4b..cdfb9898066 100644
--- a/src/bin/pg_resetwal/pg_resetwal.c
+++ b/src/bin/pg_resetwal/pg_resetwal.c
@@ -923,6 +923,13 @@ RewriteControlFile(void)
ControlFile.backupEndPoint = InvalidXLogRecPtr;
ControlFile.backupEndRequired = false;
+ /*
+ * The old WAL is gone and the new position may lie below the old
+ * watermark, which would make replay ignore future checksum transition
+ * records. The state itself is kept.
+ */
+ ControlFile.data_checksum_lsn = InvalidXLogRecPtr;
+
/*
* Force the defaults for max_* settings. The values don't really matter
* as long as wal_level='minimal'; the postmaster will reset these fields
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 2e86fd158d0..466f4223501 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -37,7 +37,8 @@ static void usage(const char *progname);
static void perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo);
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec);
static void createBackupLabel(XLogRecPtr startpoint, TimeLineID starttli,
XLogRecPtr checkpointloc);
@@ -531,7 +532,8 @@ main(int argc, char **argv)
* This is the point of no return. Once we start copying things, there is
* no turning back!
*/
- perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo);
+ perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo,
+ divergerec);
if (showprogress)
pg_log_info("syncing target data directory");
@@ -566,7 +568,8 @@ static void
perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo)
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec)
{
XLogRecPtr endrec;
TimeLineID endtli;
@@ -738,6 +741,32 @@ perform_rewind(filemap_t *filemap, rewind_source *source,
ControlFile_new.minRecoveryPoint = endrec;
ControlFile_new.minRecoveryPointTLI = endtli;
ControlFile_new.state = DB_IN_ARCHIVE_RECOVERY;
+
+ /*
+ * Keep the target's own data checksum state. Most of the data directory
+ * is still the target's: only blocks it changed since the divergence
+ * were copied from the source, so the source's state says nothing about
+ * the pages that stay. Replay from the last common checkpoint applies
+ * any WAL-logged transition the target has not seen (the watermark tells
+ * them apart), which converges the rewound server to the source's state
+ * whenever the WAL carries it.
+ */
+ ControlFile_new.data_checksum_version = ControlFile_target.data_checksum_version;
+ ControlFile_new.data_checksum_lsn = ControlFile_target.data_checksum_lsn;
+ ControlFile_new.data_checksum_is_local = ControlFile_target.data_checksum_is_local;
+
+ /*
+ * The watermark is only meaningful within the history the node replays.
+ * Records at or below the divergence point are common to both histories
+ * and stay covered, but a watermark above it was set by a transition
+ * record on the target's own abandoned fork: numerically it can cover
+ * transition records the source wrote after the divergence, and replay
+ * would skip them as already applied. Clamp it to the divergence point,
+ * so that every transition record on the source's history takes effect.
+ */
+ if (ControlFile_new.data_checksum_lsn > divergerec)
+ ControlFile_new.data_checksum_lsn = divergerec;
+
if (!dry_run)
update_controlfile(datadir_target, &ControlFile_new, do_sync);
}
diff --git a/src/bin/pg_upgrade/controldata.c b/src/bin/pg_upgrade/controldata.c
index b3bd4ccde83..de539479bdc 100644
--- a/src/bin/pg_upgrade/controldata.c
+++ b/src/bin/pg_upgrade/controldata.c
@@ -431,7 +431,7 @@ get_control_data(ClusterInfo *cluster)
cluster->controldata.date_is_int = strstr(p, "64-bit integers") != NULL;
got_date_is_int = true;
}
- else if ((p = strstr(bufin, "checksum")) != NULL)
+ else if ((p = strstr(bufin, "Data page checksum version:")) != NULL)
{
p = strchr(p, ':');
diff --git a/src/include/catalog/pg_control.h b/src/include/catalog/pg_control.h
index 7b5404460ec..177dd27ea78 100644
--- a/src/include/catalog/pg_control.h
+++ b/src/include/catalog/pg_control.h
@@ -22,7 +22,7 @@
/* Version identifier for this pg_control format */
-#define PG_CONTROL_VERSION 1903
+#define PG_CONTROL_VERSION 1904
/* Nonce key length, see below */
#define MOCK_AUTH_NONCE_LEN 32
@@ -237,6 +237,33 @@ typedef struct ControlFileData
/* Current data checksums state */
uint32 data_checksum_version;
+ /*
+ * WAL position through which data checksum transitions are covered.
+ * Replay ignores XLOG2_CHECKSUMS records ending at or below this point:
+ * their effect is already contained in data_checksum_version, or an
+ * offline pg_checksums change made after they were first applied
+ * supersedes them. Ordinarily this is the end of the newest such record
+ * this node has written or applied, but a tool may store any position
+ * that covers the same set of records. InvalidXLogRecPtr if the node
+ * has never written or applied such a record.
+ *
+ * The comparison has no timeline context, so the value is only valid
+ * within the WAL history this node replays. A tool that moves the node
+ * to another history must clamp the watermark to the point where the
+ * histories fork, as pg_rewind does, or reset it, as pg_resetwal does.
+ */
+ XLogRecPtr data_checksum_lsn;
+
+ /*
+ * True when data_checksum_version was last set by pg_checksums rather
+ * than by a WAL-logged transition. Such a state is local to this node
+ * and not derived from WAL, so recovery must not replace it with a state
+ * taken from a checkpoint record; nothing in the WAL could restore the
+ * change once it is overwritten. Cleared by the next WAL-logged
+ * transition.
+ */
+ bool data_checksum_is_local;
+
/*
* True if the default signedness of char is "signed" on a platform where
* the cluster is initialized.
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index d7eb648bd27..8d858be9927 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -89,6 +89,7 @@ PG_LWLOCK(54, WaitLSN)
PG_LWLOCK(55, LogicalDecodingControl)
PG_LWLOCK(56, DataChecksumsWorker)
PG_LWLOCK(57, AioWorkerControl)
+PG_LWLOCK(58, DataChecksumTransition)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/test/modules/test_checksums/Makefile b/src/test/modules/test_checksums/Makefile
index 71455cd5577..80f54bf6d8c 100644
--- a/src/test/modules/test_checksums/Makefile
+++ b/src/test/modules/test_checksums/Makefile
@@ -9,7 +9,7 @@
#
#-------------------------------------------------------------------------
-EXTRA_INSTALL = src/test/modules/injection_points
+EXTRA_INSTALL = contrib/pg_buffercache src/test/modules/injection_points
export enable_injection_points
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index fb7129d796f..e5d38fafb7d 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -35,6 +35,16 @@ tests += {
't/009_fpi.pl',
't/010_backup_straddle.pl',
't/011_standby_straddle.pl',
+ 't/012_offline_standby.pl',
+ 't/013_rewind.pl',
+ 't/014_lockstep.pl',
+ 't/015_standby_crash_after_disable.pl',
+ 't/016_promote_enable_crash.pl',
+ 't/017_restartpoint_race.pl',
+ 't/018_enable_crash_windows.pl',
+ 't/019_standby_shutdown_catchup.pl',
+ 't/020_cascade_divergence.pl',
+ 't/021_rewind_divergent_transitions.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
new file mode 100644
index 00000000000..241a1282ae9
--- /dev/null
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -0,0 +1,289 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Offline checksum changes with pg_checksums are local to one node. A
+# standby must neither adopt the state of the primary from replayed
+# checkpoint records, nor lose its own offline change to them.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', 'autovacuum = off');
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Scenario 1: enable offline on the primary only. The standby must
+# stay off, warn about the mismatch, and remain readable.
+$standby->stop;
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+test_checksum_state($primary, 'on');
+test_checksum_state($standby, 'off');
+
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after offline enable on the primary');
+
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+my @warnings = $log =~ /(does not match the state)/g;
+is(scalar(@warnings), 1, 'mismatch warned once per remote value');
+
+# Matching states re-arm the warning: undo the divergence on the primary,
+# then diverge again to the same value, all without restarting the standby.
+$primary->stop;
+$primary->checksum_disable_offline;
+$logstart = -s $standby->logfile;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no warning while the states match again');
+like($log, qr/now agrees with the replayed WAL/,
+ 'convergence is reported once the states match again');
+
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+test_checksum_state($standby, 'off');
+
+# The local state survives both clean and immediate restarts.
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+# Still in scenario 1's divergence: take a base backup *from the
+# standby*. A backup taken on a standby uses the last restartpoint as
+# its starting checkpoint (do_pg_backup_start()), so the record at the
+# redo point was written by the upstream primary and carries the
+# primary's state. Recovery from such a backup must not adopt that
+# state: the files were copied from the standby, and the new node
+# would come up verifying checksums its files do not have.
+
+# Make sure the standby has written pages under its own "off" state, so
+# its files really do lack checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,50000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', "SELECT count(*) FROM t2;");
+
+$standby->backup('from_standby');
+my $newnode = PostgreSQL::Test::Cluster->new('newnode');
+$newnode->init_from_backup($standby, 'from_standby');
+
+# Stream from the primary, not from the standby the files were copied
+# from, so the new node replays the "on" primary's records.
+$newnode->enable_streaming($primary);
+$newnode->start;
+
+# The new node's files all came from a cluster running with checksums
+# off. Anything but "off" here means it adopted the primary's state
+# through the checkpoint record at the redo point.
+my ($rc, $stdout, $stderr) = $newnode->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'the node copied from the standby accepts connections')
+ or diag("stderr: $stderr");
+is($stdout, 'off', 'backup of an "off" standby comes up with checksums off');
+
+# A checkpoint record streamed from the "on" primary must not flip the
+# state either.
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($newnode);
+test_checksum_state($newnode, 'off');
+
+# And it must be able to read the pages the standby wrote without
+# checksums.
+(undef, $stdout, $stderr) =
+ $newnode->psql('postgres', "SELECT count(*) FROM t2;");
+is($stdout, '50000', 'pages copied from the standby are readable')
+ or diag("stderr: $stderr");
+
+my $newnode_log = PostgreSQL::Test::Utils::slurp_file($newnode->logfile);
+unlike(
+ $newnode_log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the node copied from the standby');
+
+$newnode->stop('immediate');
+
+# Converge the cluster: enable offline on the standby too.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+$primary->wait_for_catchup($standby);
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after converging');
+
+# Scenario 2: disable offline on the standby only. The replayed
+# checkpoint records of the still-enabled primary must not override it.
+$standby->stop;
+$standby->checksum_disable_offline;
+$standby->start;
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'off');
+
+# Restartpoints must persist the local state, not the replayed copy.
+$standby->safe_psql('postgres', "CHECKPOINT;");
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable with checksums disabled locally');
+
+# Scenario 3: crash-restart right after an online transition, before the
+# next restartpoint. Replay then resumes from an older restartpoint whose
+# checkpoint records still carry the pre-transition state. Those must
+# still match the state seeded from the control file, and the transition
+# itself must be re-established by re-replaying the XLOG2_CHECKSUMS record,
+# without a spurious mismatch warning along the way.
+
+# Converge first: bring the standby back to "on" offline.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+disable_data_checksums($primary, wait => 'off');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+# Crash-restart the standby immediately, before any restartpoint has had a
+# chance to persist the new state to its control file.
+$logstart = -s $standby->logfile;
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001',
+ 'standby readable after crash-restart across an online transition');
+
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no spurious mismatch warning after crash-restart across an online transition'
+);
+
+# Scenario 4: a standby stopped while replaying an interrupted online
+# transition keeps the interrupted state in its own control file. A
+# primary is never caught this way, as its checksums launcher resolves
+# inprogress-on back to off from its exit cleanup. A standby has no
+# launcher; it carries forward whatever the last replayed record left
+# it in.
+
+# Block an online enable on the primary at inprogress-on with a
+# blocking temp table, same trick as in 004_offline.pl.
+my $bsession = $primary->background_psql('postgres');
+$bsession->query_safe('CREATE TEMPORARY TABLE tt (a integer);');
+enable_data_checksums($primary, wait => 'inprogress-on');
+
+# The standby picks up the in-progress state from the XLOG2_CHECKSUMS record.
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# Stop the standby cleanly; its restartpoint persists inprogress-on to
+# its own control file, since nothing on a standby resolves it away.
+$standby->stop;
+$standby->start;
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# The primary still sits at inprogress-on, so a base backup taken now
+# copies a control file and a redo point that both carry the
+# in-progress state; replay of the XLOG2_CHECKSUMS records completes
+# the transition on the new standby.
+$primary->backup('inprogress_backup');
+my $standby2 = PostgreSQL::Test::Cluster->new('standby2');
+$standby2->init_from_backup($primary, 'inprogress_backup',
+ has_streaming => 1);
+$standby2->start;
+
+# Backup label recovery adopts the redo point's state immediately,
+# before any further WAL is replayed.
+wait_for_checksum_state($standby2, 'inprogress-on');
+
+$bsession->quit;
+wait_for_checksum_state($primary, 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable once the transition completes');
+
+# Scenario 4 continued: the backup taken mid-transition must also
+# complete the transition and stay readable.
+$primary->wait_for_catchup($standby2);
+wait_for_checksum_state($standby2, 'on');
+
+is($standby2->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'mid-transition backup standby readable after the transition');
+
+$standby2->stop;
+
+# Nothing along this standby's replay can legitimately disagree: it
+# started out at the same inprogress-on state as the primary.
+my $standby2_log = PostgreSQL::Test::Utils::slurp_file($standby2->logfile);
+unlike(
+ $standby2_log,
+ qr/does not match the state/,
+ 'no mismatch warning on a backup taken mid-transition');
+
+# No extra CHECKPOINT needed: the transition's completion checkpoint
+# wrote every page with a checksum, and the shutdown restartpoint
+# flushes the rest.
+command_ok([ 'pg_checksums', '--check', '-D', $standby2->data_dir ],
+ 'checksums valid on the mid-transition backup standby');
+
+$standby->stop;
+$primary->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/013_rewind.pl b/src/test/modules/test_checksums/t/013_rewind.pl
new file mode 100644
index 00000000000..a791e24317d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/013_rewind.pl
@@ -0,0 +1,201 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test pg_rewind across an online data checksum enable.
+#
+# A clean switchover leaves the shutdown checkpoint of the old primary
+# as the last common checkpoint between the two nodes. When data
+# checksums are enabled online on the new primary before the old one is
+# rewound, pg_rewind installs the control file of the new primary, which
+# already claims checksums are fully enabled, while replay begins at the
+# shutdown checkpoint whose record still carries the old state. Replay
+# of the WAL stretch from before the enable must run with checksums off,
+# as recorded in the checkpoint, else it would verify pages which never
+# had checksums written and fail recovery.
+#
+# The new primary requests a checkpoint right after promotion, and
+# replaying its CHECKPOINT_REDO record would repair the state before
+# any interesting WAL is reached. Hold the checkpointer on the standby
+# in a restartpoint over the promotion, like in the recovery test
+# 041_checkpoint_at_promote, so that the post-promotion writes end up
+# in WAL before the first checkpoint of the new timeline.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Old primary. full_page_writes is off so that the updates done on the
+# promoted node do not carry full page images, forcing replay on the
+# rewound node to read the pages from disk. wal_log_hints is required
+# by pg_rewind on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+full_page_writes = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+log_checkpoints = on
+]);
+$node_a->start;
+
+if (!$node_a->check_extension('injection_points'))
+{
+ plan skip_all => 'Extension injection_points not installed';
+}
+
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+$node_a->safe_psql('postgres', "CREATE EXTENSION injection_points;");
+
+# Set the hint bits on t before taking the backup, so that reads on the
+# promoted node do not emit full page images for its pages later.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->safe_psql('postgres', "SELECT count(*) FROM t;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_a, 'off');
+test_checksum_state($node_b, 'off');
+
+# Hold the next restartpoint on the standby.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point', 'wait');");
+
+# Give the restartpoint a checkpoint record to work on, then start it
+# in a background session; it will block on the injection point with
+# the checkpointer busy until released.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+my $bg_psql = $node_b->background_psql('postgres', on_error_stop => 0);
+$bg_psql->query_until(
+ qr/starting_restartpoint/, q(
+ \echo starting_restartpoint
+ CHECKPOINT;
+));
+$node_b->wait_for_event('checkpointer', 'create-restart-point');
+
+# Clean switchover: the shutdown checkpoint of A streams to B and
+# becomes the last common checkpoint.
+$node_a->stop('fast');
+
+my ($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+$stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+my $shutdown_ckpt = $1;
+$node_b->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_ckpt'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+
+my $logstart = -s $node_b->logfile;
+$node_b->promote;
+
+# Accidental restart of the old primary, diverging its timeline.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->stop('fast');
+
+# Updates on the new primary before its first checkpoint; replay of
+# these on the rewound node has to read the pages from disk.
+$node_b->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 25 = 0;");
+ok( !$node_b->log_contains("checkpoint complete", $logstart),
+ "no checkpoint on the new timeline before the updates");
+
+# Release the checkpointer; the queued post-promotion checkpoint runs
+# after the updates.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+$bg_psql->quit;
+
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+# pg_rewind refuses to run with full_page_writes disabled on the
+# source; the updates it was disabled for are already in WAL.
+$node_b->safe_psql('postgres', "ALTER SYSTEM SET full_page_writes = on;");
+$node_b->safe_psql('postgres', "SELECT pg_reload_conf();");
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the new primary');
+
+# Replay on the rewound node must start at the shutdown checkpoint of
+# the switchover, with the control file of the new primary.
+my $backup_label = slurp_file($node_a->data_dir . '/backup_label');
+$backup_label =~ /^CHECKPOINT LOCATION: ([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from backup_label";
+is($1, $shutdown_ckpt, 'replay starts at the switchover checkpoint');
+
+($stdout, $stderr) = run_command(
+ [
+ 'pg_waldump',
+ '-p' => $node_a->data_dir . '/pg_wal',
+ '-t' => 1,
+ '-s' => $shutdown_ckpt,
+ '-n' => 1,
+ ]);
+like($stdout, qr/CHECKPOINT_SHUTDOWN/,
+ 'last common checkpoint is a shutdown checkpoint');
+
+# pg_rewind keeps the target's own checksum state in the control file it
+# writes; the online enable reaches the rewound node through WAL replay
+# below, not through the copied control file.
+($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+like(
+ $stdout,
+ qr/^Data page checksum version:\s*0$/m,
+ 'rewound node keeps its own checksum state in the control file');
+
+# Start the rewound node as a standby of the new primary. Replay runs
+# through the pre-enable WAL stretch and the online enable.
+#
+# The rewind replaced the configuration files with those of the new
+# primary, so put the port back.
+my $connstr = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_b->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/014_lockstep.pl b/src/test/modules/test_checksums/t/014_lockstep.pl
new file mode 100644
index 00000000000..713f37fb22c
--- /dev/null
+++ b/src/test/modules/test_checksums/t/014_lockstep.pl
@@ -0,0 +1,182 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# The lockstep procedure for offline checksum changes in a replication
+# setup: stop all nodes, run pg_checksums on all of them, restart.
+# Replay of checkpoint records written before the change must not
+# revert the state of the standby.
+#
+# The same pair then runs the procedure in the disable direction, and
+# finally across a divergence: after a failover both nodes get their
+# checksums enabled offline, while the last common checkpoint still
+# carries "off". pg_rewind must keep the offline "on" on the rewound
+# node, since no record in the replayed WAL could ever restore it.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# wal_log_hints keeps the pair eligible for pg_rewind without checksums.
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Part 1: the lockstep procedure in the enable direction. Stop the
+# standby first: the WAL written after this point is replayed only
+# after the offline switch, and every checkpoint record in it still
+# carries the old state.
+$standby->stop;
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 0;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 1;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+# The lockstep procedure.
+$primary->checksum_enable_offline;
+$standby->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to off.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "off" in the replayed WAL/,
+ 0);
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep enable');
+
+# Crash the standby and replay the same stretch again.
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'on');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after crash restart');
+
+# Once a post-switch checkpoint has been replayed the states match and
+# no warning may be logged.
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no mismatch warning once the states match');
+
+# Every page the standby wrote in this window must carry a checksum.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$standby->stop;
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'checksums valid on the standby');
+
+# Part 2: the lockstep procedure in the other direction. The standby
+# is already stopped; give it a pre-switch WAL stretch to replay, with
+# every checkpoint record in it still carrying "on" -- an explicit
+# checkpoint plus the shutdown checkpoint written by stop().
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 2;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+$primary->checksum_disable_offline;
+$standby->checksum_disable_offline;
+
+my $disable_logstart = -s $standby->logfile;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to on.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $disable_logstart);
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep disable');
+
+# Part 3: failover, divergence and pg_rewind across offline enables.
+# The roles swap from here on: the standby becomes the new primary and
+# the old primary is rewound to follow it, but the variables keep the
+# names they had above.
+#
+# Failover: promote the standby, then diverge the old primary.
+$standby->promote;
+$standby->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$primary->stop;
+
+# The lockstep procedure across the divergence: both nodes get their
+# checksums enabled offline.
+$primary->checksum_enable_offline;
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+
+# The states match, so the rewind proceeds without complaint.
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $primary->data_dir,
+ '--source-server' => $standby->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled offline on both nodes');
+
+# The rewind replaced the configuration files with those of the source,
+# so put the port back. The copied file also carries the source's own
+# stale primary_conninfo; enable_streaming() below appends ours, which
+# wins as the later entry.
+$primary->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$primary->port]}
+]);
+$primary->enable_streaming($standby);
+
+my $rewind_logstart = -s $primary->logfile;
+$primary->start;
+$standby->wait_for_catchup($primary);
+
+is($primary->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+
+# The offline enable survives: replay from the common checkpoint must
+# not resurrect the pre-divergence "off".
+test_checksum_state($primary, 'on');
+
+$log =
+ PostgreSQL::Test::Utils::slurp_file($primary->logfile, $rewind_logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no divergence reported between the rewound server and its source');
+
+$primary->stop;
+$standby->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
new file mode 100644
index 00000000000..315b27b93d4
--- /dev/null
+++ b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
@@ -0,0 +1,125 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that a standby crashing after a replayed online disable, but before
+# its next restartpoint, restarts cleanly.
+#
+# Pages dirtied before the XLOG2_CHECKSUMS record and evicted after it are
+# written out under the new "off" state, without checksums. The control
+# file must follow the record immediately in this direction: a crash-restart
+# initializes checksum verification from the control file, and one still
+# saying "on" would fail verification on exactly those pages while replaying
+# records older than the state change.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$primary->start;
+
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,1000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+
+# Small buffer pool so that replaying a bulk load evicts the pages we care
+# about, and no restartpoints so the control file stays where it is.
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 1MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_delay = 10000
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'on');
+
+# No full page images, so that replay has to read the pages of "t" back from
+# disk instead of overwriting them from the WAL.
+$primary->append_conf('postgresql.conf', 'full_page_writes = off');
+$primary->reload;
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" on the standby, *before* the state change. They are
+# not flushed: the standby has no restartpoint from here on.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+# Online disable. The standby replays it and switches to "off", but its
+# control file is not updated.
+$primary->safe_psql('postgres', 'SELECT pg_disable_data_checksums();');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+my ($ctl_before) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($ctl_state) = $ctl_before =~ /Data page checksum version:\s+(\d+)/;
+note(
+ "standby control file data_checksum_version after the disable: "
+ . "$ctl_state (0 = off, 1 = on)");
+
+# Force the standby to evict the dirty pages of "t" now that it is "off":
+# they are written back without checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE filler AS SELECT generate_series(1,300000) AS a;");
+$primary->wait_for_catchup($standby);
+
+# Crash the standby before it gets a chance to run a restartpoint.
+$standby->stop('immediate');
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'standby restarts after crashing between a checksum state '
+ . 'change and the next restartpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'the pages written while "off" are readable')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("standby log tail:\n" . join("\n", @lines[ -12 .. -1 ]))
+ if @lines >= 12;
+ fail('the pages written while "off" are readable');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/016_promote_enable_crash.pl b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
new file mode 100644
index 00000000000..97f84da5685
--- /dev/null
+++ b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a standby promoted right after replaying an online enable, before any
+# restartpoint has flushed the rewritten pages.
+#
+# The end-of-recovery record persists the checksum state without flushing the
+# buffer pool, while the control file still points at a restartpoint older than
+# the transition. Persisting "on" there would make a crash before the
+# post-promotion checkpoint verify checksums over the pages that were flushed
+# while the state was still "off"; promotion must take the full checkpoint
+# instead.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Dirty the pages of "t" without full page images, then push them out to disk
+# on the standby while checksums are still off.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Online enable on the primary; the standby replays it into its own buffers,
+# where the rewritten pages stay dirty (no restartpoint, no bgwriter).
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+note("standby control file before the promotion: $before");
+
+# Promote. The end-of-recovery record persists the live state without any
+# buffer flush; the checkpoint requested afterwards is not immediate.
+$standby->promote;
+$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery();')
+ or die 'timed out waiting for the promotion';
+$standby->stop('immediate');
+
+($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($ckpt) = $ctl =~ /Latest checkpoint location:\s+(\S+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note(
+ "standby control file after the promotion: $after, checkpoint $ckpt, redo $redo"
+);
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'promoted node restarts after crashing before its first checkpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/017_restartpoint_race.pl b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
new file mode 100644
index 00000000000..65aa834f366
--- /dev/null
+++ b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
@@ -0,0 +1,138 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a restartpoint whose flush races the replay of an online enable.
+#
+# Replay keeps running while CheckPointGuts() writes out the buffer pool, so
+# the state at the end of the flush can be newer than the one the flush ran
+# under. Persisting it would claim checksums for pages the flush wrote out
+# before the transition, against a redo pointer that predates it.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# C1: the checkpoint the pending restartpoint will target.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" without full page images and push them to disk on
+# the standby while it is still "off".
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Start a restartpoint for C1 and hold it right after CheckPointGuts().
+$standby->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point','wait');");
+my $bg = $standby->background_psql('postgres');
+$bg->query_until(qr//, "\\echo restartpoint\nCHECKPOINT;\n");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint injection point';
+
+# The startup process keeps replaying while the checkpointer is held: the
+# whole online enable lands, and the rewritten pages stay dirty.
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+# Let the restartpoint finish; it persists the state it samples now.
+$standby->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint to finish';
+$standby->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+
+$standby->stop('immediate');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note("standby control file after the restartpoint: $after, redo $redo");
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started, 'standby restarts after a restartpoint that raced the enable');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/018_enable_crash_windows.pl b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
new file mode 100644
index 00000000000..17f96ff6272
--- /dev/null
+++ b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
@@ -0,0 +1,496 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Crash and checkpoint windows inside an online data checksum enable on
+# a single node. Each scenario starts from "off" on the same cluster,
+# runs the enable into a held injection point, and checks what a crash
+# or a concurrent checkpoint in that window may and may not do to the
+# control file and the pages on disk.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+use IPC::Run;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Bring the node back to "off" with no leftover table between
+# scenarios. The crash restarts already resolve an interrupted enable
+# to "off", so only disable when something is left to disable. Those
+# restarts also clear any attached injection points. The checkpoint
+# bounds the WAL the next scenario's crash recovery has to replay.
+sub reset_scenario
+{
+ my ($node) = @_;
+
+ BAIL_OUT('node is not running after the previous scenario')
+ unless $node->is_alive;
+
+ my $state = $node->safe_psql('postgres', 'SHOW data_checksums;');
+ disable_data_checksums($node, wait => 'off') if $state ne 'off';
+ $node->safe_psql('postgres', 'DROP TABLE IF EXISTS t;');
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+ test_checksum_state($node, 'off');
+}
+
+my $node = PostgreSQL::Test::Cluster->new('crash_windows');
+$node->init(no_data_checksums => 1);
+
+# Scenario 1 needs the pages of "t" to reach disk without a full page
+# image, hence full_page_writes and wal_log_hints off. Scenario 2
+# needs them to stay dirty in shared buffers until a checkpoint, hence
+# no bgwriter and a large shared_buffers. wal_level is pinned because
+# init() writes "minimal" for a non-streaming node. The rest merely
+# tolerate the settings.
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+wal_level = replica
+max_wal_size = 10GB
+full_page_writes = off
+wal_log_hints = off
+shared_buffers = 512MB
+bgwriter_lru_maxpages = 0
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$node->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+
+test_checksum_state($node, 'off');
+
+# Scenario 1: a primary crashes inside SetDataChecksumsOn(), just before the
+# forced checkpoint that flushes the rewritten pages, with crash recovery
+# resuming from a checkpoint older than the transition. The control file
+# must still say "inprogress-on" there: replay does not adopt the state of
+# the checkpoint record it resumes from, so an "on" written before the flush
+# would stay in effect while replay reads pages whose rewrite never reached
+# disk. Here the pages went out through the rewriting worker's ring buffer,
+# so only the control file state discriminates; see the next scenario for
+# the case where they did not.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Establish the checkpoint that crash recovery will resume from.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty the pages of "t" without full page images, then push them out
+ # to disk while checksums are still off.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+ my $evicted = $node->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+ note("evict_relation on t: $evicted, relpath $relpath");
+
+ # Stop the enable right after the control file has been updated to
+ # "on" but before the checkpoint that flushes the rewritten pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+ $node->safe_psql('postgres', 'SELECT pg_enable_data_checksums();');
+
+ $node->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'datachecksums-on-before-checkpoint';"
+ ) or die 'timed out waiting for the injection point';
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ $node->stop('immediate');
+
+ # Show the on-disk checksum field of the first page of "t".
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 2: the same crash as above, but with the rewritten pages resident
+# in shared buffers instead of gone out through the ring. The rewriting
+# worker reads through a BAS_VACUUM ring, which writes pages back as the
+# ring recycles, but a page already resident in shared buffers is not read
+# through the ring: ReadBufferExtended() hands back the existing buffer, and
+# it stays dirty until a checkpoint. The control file may therefore not
+# say "on" before that checkpoint has run, or crash recovery would resume
+# from a checkpoint older than the transition with verification already
+# enabled, and replay records that read those still-unchecksummed pages.
+# full_page_writes is off so that the records do not simply overwrite the
+# pages with a full page image.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,100000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Establish the checkpoint that crash recovery will resume from, with the
+ # pages of "t" written out while checksums are still off.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty those pages again without emitting full page images, and leave
+ # them in shared buffers. Replay of these records has to read the
+ # pages from disk.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+
+ my $dirty_before = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_before, '>', 0,
+ 'pages of t are resident and dirty before enabling checksums');
+
+ # Hold the enabling right before the checkpoint that flushes the rewritten
+ # pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ # The rewritten pages must still be sitting dirty in shared buffers, or
+ # the window this test is about does not exist.
+ my $dirty_after = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_after, '>', 0,
+ 'rewritten pages of t are still dirty before the checkpoint');
+
+ $node->stop('immediate');
+
+ # The on-disk copy is the one written before the transition, without a
+ # checksum.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ is($pd_checksum, 0, 'block 0 of t on disk carries no checksum');
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 3: a checkpoint that runs between the XLOG2_CHECKSUMS("on")
+# record and the control file write at the end of SetDataChecksumsOn() must
+# not make a crash throw the completed transition away. SetDataChecksumsOn()
+# writes the record, flips shared memory to "on", emits the barrier and only
+# then requests the checkpoint that flushes the rewritten pages; the control
+# file is written after that checkpoint returns. A crash in that window is
+# harmless only as long as recovery still starts before the record. Any
+# checkpoint completing in the window moves the redo point past the record,
+# so recovery would never see it, would come up with the control file's
+# "inprogress-on" and StartupXLOG() would demote that to "off", even though
+# every page on disk carries a checksum by then. CreateCheckPoint()
+# therefore persists the state the checkpoint ran under, the same way
+# CreateRestartPoint() does. Here the window is made deterministic by
+# holding the launcher at the datachecksums-on-before-checkpoint injection
+# point and checkpointing from another session.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Hold the launcher after the record, the shared memory flip and the
+ # barrier, but before the checkpoint SetDataChecksumsOn() requests
+ # itself.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ # Every backend already sees "on" and writes checksums.
+ test_checksum_state($node, 'on');
+
+ # ... while the control file still says "inprogress-on".
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($before, '3', 'control file says "inprogress-on" inside the window');
+
+ # A concurrent checkpoint. It flushes the rewritten pages and moves
+ # the redo point past the XLOG2_CHECKSUMS("on") record, so it has to
+ # record the "on" state in the control file as well.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($after, '1',
+ 'the concurrent checkpoint records "on" in the control file');
+
+ $node->stop('immediate');
+
+ # The checkpoint flushed the rewritten pages, so they carry a checksum on
+ # disk: the transition really did complete.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ isnt($pd_checksum, 0, 'block 0 of t on disk carries a checksum');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition is complete on disk, so the cluster has to come back
+ # "on".
+ test_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+}
+reset_scenario($node);
+
+# Scenario 4: a crash between the checkpoint an online enable requests and
+# the control file write that follows it must not lose the transition. The
+# last steps of SetDataChecksumsOn() are
+#
+# WAL record -> shmem -> barrier -> checkpoint -> persist
+#
+# The checkpoint is what makes "on" safe to persist, but it also moves the
+# redo point above the XLOG2_CHECKSUMS record that carries the new state.
+# Crash recovery started from that checkpoint therefore never replays the
+# record, so without the state the checkpoint itself persists, a crash in
+# the remaining window would bring the cluster back at "inprogress-on",
+# which StartupXLOG() resolves to "off", discarding a transition whose
+# pages are all on disk with a checksum. The previous scenario exercises
+# the same window through a checkpoint requested by another session; this
+# one closes it from the other side: the transition's own checkpoint is the
+# one that moves the redo point, so persisting the state from
+# CreateCheckPoint() is what has to save it, not the write below the
+# injection point.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Hold the launcher after the checkpoint that licenses "on" has
+ # completed and before the state reaches the control file.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-after-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-after-checkpoint');
+
+ # The transition is complete as far as the running cluster is concerned.
+ test_checksum_state($node, 'on');
+
+ # The checkpoint has already recorded it, so the pending write below the
+ # injection point has nothing left to do.
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($version) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($version, '1',
+ 'the requested checkpoint recorded "on" in the control file');
+
+ # Crash before the control file write that follows the checkpoint.
+ $node->stop('immediate');
+
+ $node->start;
+
+ # Every page on disk carries a checksum and the checkpoint that
+ # flushed them completed, so the cluster has to come back verifying
+ # them.
+ test_checksum_state($node, 'on');
+
+ is($node->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '10000', 'relation readable after the crash');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ # The data directory must be one the offline tools accept.
+ $node->stop;
+ $node->command_ok([ 'pg_checksums', '--check', '-D', $node->data_dir ],
+ 'pg_checksums accepts the data directory');
+
+ # Leave the node running for the reset.
+ $node->start;
+}
+reset_scenario($node);
+
+# Scenario 5: a checkpoint racing SetDataChecksumsOn() between the insertion
+# of the XLOG2_CHECKSUMS("on") record and the shared memory update must not
+# insert an XLOG_CHECKPOINT_REDO record that follows the transition in WAL
+# order while still carrying "inprogress-on". Recovery resuming from such a
+# redo point never replays the preceding transition record, comes up in
+# "inprogress-on" and resolves the finished transition as interrupted.
+# XLogChecksums() closes the window by inserting the record and publishing
+# the new state under DataChecksumTransitionLock, which CreateCheckPoint()
+# takes around sampling the state and inserting the redo record. Here the
+# launcher is held between the two steps at the
+# datachecksums-on-before-publish injection point, and a concurrent
+# CHECKPOINT has to block on the lock instead of completing inside the
+# window.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # The datachecksums-on-before-publish point fires inside a critical
+ # section, where the wait machinery must not allocate. Waiting once
+ # at the datachecksums-enable-checksums-delay point, which the
+ # launcher runs outside the critical section, initializes it; see
+ # 050_redo_segment_missing.pl for the same recipe around
+ # create-checkpoint-run.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-enable-checksums-delay','wait');"
+ );
+
+ # Hold the launcher after the XLOG2_CHECKSUMS("on") record is in WAL but
+ # before the new state is published in shared memory.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-publish','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-enable-checksums-delay');
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-enable-checksums-delay');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-enable-checksums-delay');");
+
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-publish');
+
+ # The record is in WAL, the published state is still the old one.
+ test_checksum_state($node, 'inprogress-on');
+
+ # A concurrent checkpoint. It must block on DataChecksumTransitionLock
+ # before inserting its redo record rather than complete inside the window.
+ my $checkpointer = IPC::Run::start(
+ [ 'psql', '-XAtq', '-d', $node->connstr('postgres'), '-c', 'CHECKPOINT;' ],
+ '<' => '/dev/null',
+ '>' => '/dev/null',
+ '2>' => '/dev/null',
+ IPC::Run::timer($PostgreSQL::Test::Utils::timeout_default));
+
+ ok( $node->poll_query_until(
+ 'postgres',
+ "SELECT wait_event = 'DataChecksumTransition' "
+ . "FROM pg_stat_activity WHERE backend_type = 'checkpointer';"),
+ 'concurrent checkpoint blocks on DataChecksumTransitionLock');
+
+ # Release the launcher; the checkpoint then samples the published "on" and
+ # its redo record follows the transition record.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-publish');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-publish');");
+
+ $checkpointer->finish;
+
+ # Crash while the launcher's own checkpoint may still be in flight.
+ # Recovery resumes from the concurrent checkpoint's redo point, which
+ # now lies above the transition record and carries "on".
+ $node->stop('immediate');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition completed, so the cluster has to come back "on".
+ wait_for_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ $node->stop;
+}
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
new file mode 100644
index 00000000000..aaa35f8f7dd
--- /dev/null
+++ b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
@@ -0,0 +1,126 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A standby stopped after replaying an online enable, but before the
+# checkpoint record that follows it on the primary.
+#
+# The shutdown restartpoint has no new checkpoint record to work from
+# and is skipped, so the control file would keep the in-progress state
+# the transition passed through, even though replay left the node at
+# "on" and rewrote every page. pg_checksums reads that field and
+# refuses to run on an in-progress state, so the shutdown has to flush
+# and catch it up instead.
+#
+# The caught-up control file then carries the watermark of the "on"
+# record, which the second half of this test relies on: an offline
+# disable made while the standby is down must survive the re-replay of
+# that record on the next startup, or the change would be silently
+# reverted.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+wal_keep_size = 1GB
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Anchor the standby's last restartpoint here, so that nothing replayed
+# from now on gives the shutdown restartpoint a newer checkpoint record
+# to use.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Hold the enable right after the state change record, before the
+# checkpoint it requests once the transition is complete.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+enable_data_checksums($primary);
+$primary->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+wait_for_checksum_state($primary, 'on');
+
+# The state change record is already flushed and streams on its own;
+# no checkpoint record follows while the launcher is held.
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+$standby->stop('fast');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($state, '1',
+ 'shutdown catches the control file up with the replayed state');
+
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'pg_checksums verifies the stopped standby');
+
+# Disable checksums offline while the standby is down. The restart
+# below resumes replay under the last restartpoint, re-reading the "on"
+# record; the watermark persisted with the catch-up above marks it as
+# already applied, so the offline change must survive.
+$standby->checksum_disable_offline;
+
+# Let the enable finish on the primary.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-checkpoint');");
+$primary->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-checkpoint');");
+
+# The launcher exits only after the checkpoint it requests completes,
+# so a checkpoint record now follows the transition in WAL.
+$primary->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_catalog.pg_stat_activity "
+ . "WHERE backend_type = 'datachecksums launcher';");
+
+my $log_offset = -s $standby->logfile;
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+
+# The nodes legitimately diverged, which replayed checkpoints report.
+$standby->wait_for_log(
+ qr/data checksum state "off" of this node does not match the state "on" in the replayed WAL/,
+ $log_offset);
+
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/020_cascade_divergence.pl b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
new file mode 100644
index 00000000000..b4b474fa583
--- /dev/null
+++ b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
@@ -0,0 +1,115 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# An offline data checksum state change made on one node of a cascading setup
+# is reported all the way down the chain.
+#
+# CheckReplayedDataChecksumState() is reached from xlog_redo() when a
+# checkpoint record (XLOG_CHECKPOINT_SHUTDOWN or XLOG_CHECKPOINT_REDO) is
+# replayed after consistency has been reached. Those records are written by
+# the root primary only and are relayed verbatim by every intermediate
+# standby, so a cascaded standby that was changed offline learns about the
+# divergence too. An intermediate standby's own restartpoints are not
+# WAL-logged and cannot surface it.
+#
+# Note that this makes detection checkpoint-driven: an idle or read-only
+# primary surfaces the divergence no sooner than its next checkpoint, so up to
+# checkpoint_timeout may pass before the operator sees anything. Closing that
+# window would require comparing the state when a walreceiver connects.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# checkpoint_timeout is deliberately long so that the only checkpoint record in
+# this test is the one requested explicitly below.
+my $conf = qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+);
+
+my $primary = PostgreSQL::Test::Cluster->new('cascade_primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', $conf);
+$primary->start;
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,1000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('cascade_standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf('postgresql.conf', $conf);
+$standby->start;
+
+$standby->backup('backup');
+my $cascade = PostgreSQL::Test::Cluster->new('cascade_cascade');
+$cascade->init_from_backup($standby, 'backup', has_streaming => 1);
+$cascade->append_conf('postgresql.conf', $conf);
+$cascade->start;
+
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+test_checksum_state($cascade, 'off');
+
+# Change the cascaded standby offline, and only it.
+$cascade->stop;
+$cascade->command_ok([ 'pg_checksums', '--enable', '-D', $cascade->data_dir ],
+ 'pg_checksums enables checksums on the cascaded standby only');
+
+my $logstart = -s $cascade->logfile;
+$cascade->start;
+
+test_checksum_state($cascade, 'on');
+
+# Ordinary WAL traffic carries no state, so the cascaded standby keeps
+# streaming from a chain whose state is "off" without noticing.
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (1);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+is($cascade->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '1001', 'cascaded standby keeps streaming from a divergent chain');
+
+# Restartpoints on the intermediate standby are not WAL-logged, so this does
+# not surface anything on the cascaded standby either.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (2);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+my $log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+unlike(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'an intermediate restartpoint does not surface the divergence');
+
+# A checkpoint on the root primary does: the record is relayed down the whole
+# chain, so both the direct and the cascaded standby compare it against their
+# own state. wait_for_catchup() below waits for replay, not just receipt, so
+# the record has been through xlog_redo() on both nodes once it returns.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+$log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+like(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'the cascaded standby reports the divergence on the primary checkpoint');
+
+$cascade->stop;
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
new file mode 100644
index 00000000000..a6c8a1091ee
--- /dev/null
+++ b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
@@ -0,0 +1,193 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind across online data checksum transitions on both sides of a
+# divergence, with the two nodes trading roles between the scenarios.
+#
+# Scenario 1: after a switchover the new primary enables checksums
+# online, while the old primary restarts on its old timeline, advances
+# its WAL beyond the enable records and runs an online enable/disable
+# cycle of its own. The old primary ends "off" with a checksum
+# watermark numerically above every checksum record the new primary has
+# written. pg_rewind clamps the watermark it keeps to the divergence
+# point; without the clamp, replay on the rewound node would skip the
+# source's enable as already applied and stay "off" under an "on"
+# primary.
+#
+# Scenario 2: checksums are disabled again, and after another
+# switchover both nodes enable them online independently, so the
+# divergence checkpoint carries "off" while both control files say
+# "on". The rewind is allowed, the target keeps its own "on" state,
+# and replay re-walks the source's enable onto the already enabled
+# node.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+sub controldata_watermark
+{
+ my ($node) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $node->data_dir ]);
+ $stdout =~ /^Data checksum watermark:\s*([0-9A-F]+)\/([0-9A-F]+)$/m
+ or die "watermark missing from pg_controldata output";
+ return (hex($1) << 32) + hex($2);
+}
+
+# Wait until the standby has replayed the shutdown checkpoint of the
+# stopped primary, so that a subsequent promotion diverges after it and
+# the shutdown checkpoint becomes the last common checkpoint.
+sub wait_for_shutdown_checkpoint_replay
+{
+ my ($primary, $standby) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $primary->data_dir ]);
+ $stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+ my $shutdown_checkpoint = $1;
+
+ $standby->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_checkpoint'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+}
+
+# Old primary, checksums off. wal_log_hints is required by pg_rewind
+# on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+# Clean switchover to B; enable checksums online on it.
+$node_a->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_a, $node_b);
+$node_b->promote;
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+$node_b->stop('fast');
+my $watermark_b = controldata_watermark($node_b);
+$node_b->start;
+
+# Accidental restart of the old primary on the old timeline. Advance
+# its WAL beyond the enable watermark of B, then run an online enable
+# and disable cycle: the node ends "off" with a watermark above every
+# checksum record B has written.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t_pad AS SELECT generate_series(1,200000) AS a;");
+enable_data_checksums($node_a, wait => 'on');
+disable_data_checksums($node_a, wait => 'off');
+test_checksum_state($node_a, 'off');
+$node_a->stop('fast');
+
+my $watermark_a = controldata_watermark($node_a);
+die "test broken: target watermark not above the source's enable"
+ unless $watermark_a > $watermark_b;
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the promoted node');
+
+# Start the rewound node as a standby of B. Replay from the last
+# common checkpoint runs through B's online enable, which the target
+# never saw, so the rewound node must converge to "on".
+my $connstr_b = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr_b application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+# Scenario 2, reusing the pair with the roles reversed. Disable
+# checksums online so the next divergence point carries "off", and let
+# A replay the change.
+disable_data_checksums($node_b, wait => 'off');
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'off');
+
+# Clean switchover back to A; enable checksums online on it.
+$node_b->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_b, $node_a);
+$node_a->promote;
+enable_data_checksums($node_a, wait => 'on');
+test_checksum_state($node_a, 'on');
+my $source_enable_watermark = controldata_watermark($node_a);
+
+# The old primary restarts on its old timeline and enables checksums
+# online independently: both control files say "on", the divergence
+# checkpoint says "off".
+$node_b->start;
+$node_b->safe_psql('postgres', "INSERT INTO t_div VALUES (2);");
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+$node_b->stop('fast');
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_b->data_dir,
+ '--source-server' => $node_a->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled on both nodes');
+
+my $connstr_a = $node_a->connstr;
+$node_b->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_b->port]}
+primary_conninfo = '$connstr_a application_name=@{[$node_b->name]}'
+]);
+$node_b->set_standby_mode;
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_b, 'on');
+
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_b->stop('fast');
+is(controldata_watermark($node_b), $source_enable_watermark,
+ 'rewound node replayed the source checksum transition');
+command_ok([ 'pg_checksums', '--check', '-D', $node_b->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the twice-rewound node');
+
+done_testing();
--
2.55.0
[application/octet-stream] v10-0002-pg_checksums-Refuse-interrupted-transitions-note.patch (7.7K, ../../CAN4CZFOOHZmVnL-B2D+V7ODmVFgkYucgQD+-qegDTKZgJ3dtQg@mail.gmail.com/5-v10-0002-pg_checksums-Refuse-interrupted-transitions-note.patch)
download | inline diff:
From 645b5b2a55b3a36626455af163b162a23dc2b6fa Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:19 +0000
Subject: [PATCH v10 2/4] pg_checksums: Refuse interrupted transitions, note
change is local
An inprogress-on or inprogress-off control file means an online
transition was cut short; running the offline tool on top of it mixes
two procedures. Refuse it in every mode, with a hint pointing at the
start-stop cycle (or, on a standby, at letting replication finish)
that resets the state. Also print that the change applies to one
data directory only, with standby-aware wording, since in a
replication setup the same change must be made on every node.
On a primary, inprogress-on cannot survive a graceful stop: the
checksums launcher resolves it from its exit cleanup, and a crashed
primary is already rejected by the existing "cluster must be shut
down" check. inprogress-off can, however: it is set by the backend
running pg_disable_data_checksums(), and a fast shutdown arriving
between its two barriers leaves it behind in a cleanly shut down
control file, where the next start-stop cycle resolves it at end of
recovery. A standby can be stopped with either state, since it has
no launcher and only carries forward whatever state the last replayed
WAL record left it in, with a restartpoint persisting that as-is.
The standby is also the deterministic way to reach the new guard,
which is why the test coverage uses one; the primary window would
need an injection point between the two barriers.
The pg_checksums page said the tool still processes all relation files
regardless of an interrupted online transition; describe the refusal
instead.
---
doc/src/sgml/ref/pg_checksums.sgml | 11 ++++---
src/bin/pg_checksums/pg_checksums.c | 21 +++++++++++++
.../test_checksums/t/012_offline_standby.pl | 31 ++++++++++++++-----
3 files changed, 51 insertions(+), 12 deletions(-)
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index 969497586f4..020f6f82c53 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -49,11 +49,12 @@ PostgreSQL documentation
</para>
<para>
- When enabling checksums with <application>pg_checksums</application>, if
- checksums were in the process of being enabled using
- <xref linkend="checksums-online-enable-disable"/> when the cluster was shut
- down, <application>pg_checksums</application> will still process all
- relation files regardless of the progress of online checksum processing.
+ If checksums were in the process of being enabled or disabled using
+ <xref linkend="checksums-online-enable-disable"/> when the cluster was
+ shut down, the control file still records that interrupted state, and
+ <application>pg_checksums</application> refuses to run in any mode.
+ Start the cluster and shut it down cleanly to reset the state, then
+ retry; on a standby, let replication complete the transition first.
</para>
<para>
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 8c4ae59222d..95757d0424b 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -586,6 +586,21 @@ main(int argc, char *argv[])
ControlFile->state != DB_SHUTDOWNED_IN_RECOVERY)
pg_fatal("cluster must be shut down");
+ /*
+ * An inprogress state means an online transition was cut short. A
+ * standby stopped mid-transition carries either state; a cleanly shut
+ * down primary can still carry inprogress-off, which a fast shutdown
+ * during pg_disable_data_checksums() leaves behind, while inprogress-on
+ * is always resolved by the launcher's exit cleanup.
+ */
+ if (ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted");
+ pg_log_error_hint("Start and cleanly shut down the cluster once to reset the data checksum state, then retry. On a standby, let replication complete the transition first.");
+ exit(1);
+ }
+
if (ControlFile->data_checksum_version != PG_DATA_CHECKSUM_VERSION &&
mode == PG_MODE_CHECK)
pg_fatal("data checksums are not enabled in cluster");
@@ -673,6 +688,12 @@ main(int argc, char *argv[])
printf(_("Checksums enabled in cluster\n"));
else
printf(_("Checksums disabled in cluster\n"));
+
+ printf(_("This change applies to this data directory only.\n"));
+ if (ControlFile->state == DB_SHUTDOWNED_IN_RECOVERY)
+ printf(_("This node appears to be a standby; apply the same change to the primary and all other standbys.\n"));
+ else
+ printf(_("In a replication setup, apply the same change to every node while all are stopped, before restarting any of them.\n"));
}
return 0;
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
index 241a1282ae9..5efe11f993e 100644
--- a/src/test/modules/test_checksums/t/012_offline_standby.pl
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -149,7 +149,10 @@ $newnode->stop('immediate');
# Converge the cluster: enable offline on the standby too.
$standby->stop;
-$standby->checksum_enable_offline;
+command_checks_all(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ 0, [qr/appears to be a standby/],
+ [], 'standby-role notice on offline enable');
$standby->start;
test_checksum_state($standby, 'on');
$primary->wait_for_catchup($standby);
@@ -217,11 +220,11 @@ unlike(
);
# Scenario 4: a standby stopped while replaying an interrupted online
-# transition keeps the interrupted state in its own control file. A
-# primary is never caught this way, as its checksums launcher resolves
-# inprogress-on back to off from its exit cleanup. A standby has no
-# launcher; it carries forward whatever the last replayed record left
-# it in.
+# transition keeps the interrupted state in its own control file, and
+# pg_checksums must refuse to touch it. A primary is never caught this
+# way, as its checksums launcher resolves inprogress-on back to off from
+# its exit cleanup. A standby has no launcher; it carries forward
+# whatever the last replayed record left it in.
# Block an online enable on the primary at inprogress-on with a
# blocking temp table, same trick as in 004_offline.pl.
@@ -235,6 +238,20 @@ wait_for_checksum_state($standby, 'inprogress-on');
# Stop the standby cleanly; its restartpoint persists inprogress-on to
# its own control file, since nothing on a standby resolves it away.
$standby->stop;
+
+command_fails_like(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --enable refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --check refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--disable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --disable refuses a standby stopped mid-transition');
+
$standby->start;
wait_for_checksum_state($standby, 'inprogress-on');
@@ -257,7 +274,7 @@ wait_for_checksum_state($primary, 'on');
$primary->wait_for_catchup($standby);
wait_for_checksum_state($standby, 'on');
-is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
'10001', 'standby readable once the transition completes');
# Scenario 4 continued: the backup taken mid-transition must also
--
2.55.0
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-02 14:34 Zsolt Parragi <zsolt.parragi@percona.com>
parent: Zsolt Parragi <zsolt.parragi@percona.com>
0 siblings, 1 reply; 43+ messages in thread
From: Zsolt Parragi @ 2026-09-02 14:34 UTC (permalink / raw)
To: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; pgsql-hackers@lists.postgresql.org
v11 adds a few more edits based on Daniel's feedback.
Attachments:
[application/octet-stream] v11-0003-pg_rewind-Check-the-data-checksum-states-of-sour.patch (21.6K, ../../CAN4CZFP_-vMOtHEj0u_v9OYoihxmi1QkfP_wiP_ytdg9FE=KfA@mail.gmail.com/2-v11-0003-pg_rewind-Check-the-data-checksum-states-of-sour.patch)
download | inline diff:
From 311414e23d0d84a4d0eed775c3f4ca6eadc6788b Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 18:40:22 +0000
Subject: [PATCH v11 3/4] pg_rewind: Check the data checksum states of source
and target
pg_rewind installs the source's control file on the target, but every
block it does not copy keeps the target's content. With checksums
enabled on the target and disabled on the source, the rewound server
claims enabled checksums while the blocks copied from the source have
none, and fails checksum verification as soon as it reads them; in the
worst case every connection attempt dies on an unverifiable catalog
page.
Comparing the control files is not enough. Replay on the rewound
server resumes from the last common checkpoint and adopts the data
checksum state recorded there, so an offline disable on the source
after the divergence leaves both control files saying "off" while the
rewound server still resumes with verification enabled. Check the
state carried by the divergence checkpoint as well, which pg_rewind
already reads.
Refuse both cases, and an interrupted online transition on either
side, same as pg_checksums does. The opposite mismatch, checksums on
the source only, is allowed with a warning: recovery under the backup
label adopts the divergence state, so the rewound server keeps
checksums disabled (or converges through the replayed WAL if the
source enabled them online), and no page can fail verification.
---
doc/src/sgml/ref/pg_rewind.sgml | 14 ++
src/bin/pg_rewind/parsexlog.c | 4 +-
src/bin/pg_rewind/pg_rewind.c | 93 +++++++++++-
src/bin/pg_rewind/pg_rewind.h | 1 +
src/test/modules/test_checksums/meson.build | 2 +
.../test_checksums/t/022_rewind_state.pl | 133 +++++++++++++++++
.../t/023_rewind_standby_target.pl | 141 ++++++++++++++++++
7 files changed, 386 insertions(+), 2 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/022_rewind_state.pl
create mode 100644 src/test/modules/test_checksums/t/023_rewind_standby_target.pl
diff --git a/doc/src/sgml/ref/pg_rewind.sgml b/doc/src/sgml/ref/pg_rewind.sgml
index b95e3868e9d..ac3d0c9328f 100644
--- a/doc/src/sgml/ref/pg_rewind.sgml
+++ b/doc/src/sgml/ref/pg_rewind.sgml
@@ -124,6 +124,20 @@ PostgreSQL documentation
<literal>on</literal>, but is enabled by default.
</para>
+ <para>
+ The data checksum states of the source and target server must be
+ compatible. <application>pg_rewind</application> refuses to run if an
+ online checksum state transition was interrupted on either server, or
+ if data checksums are enabled on the target server, or were enabled at
+ the point of divergence, while the source server runs without them: the
+ rewound server would fail checksum verification on the blocks copied
+ from the source. The opposite combination is allowed with a warning;
+ the rewound server keeps data checksums disabled unless the source
+ server enabled them online after the point of divergence. See
+ <xref linkend="app-pgchecksums"/> for how to change the state
+ consistently across a replication setup.
+ </para>
+
<warning>
<title>Warning: Failures While Rewinding</title>
<para>
diff --git a/src/bin/pg_rewind/parsexlog.c b/src/bin/pg_rewind/parsexlog.c
index 023e23b063c..6e87b00f8c2 100644
--- a/src/bin/pg_rewind/parsexlog.c
+++ b/src/bin/pg_rewind/parsexlog.c
@@ -167,7 +167,8 @@ readOneRecord(const char *datadir, XLogRecPtr ptr, int tliIndex,
void
findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
- XLogRecPtr *lastchkptredo, const char *restoreCommand)
+ XLogRecPtr *lastchkptredo, uint32 *lastchkptdatachecksums,
+ const char *restoreCommand)
{
/* Walk backwards, starting from the given record */
XLogRecord *record;
@@ -255,6 +256,7 @@ findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
*lastchkptrec = searchptr;
*lastchkpttli = checkPoint.ThisTimeLineID;
*lastchkptredo = checkPoint.redo;
+ *lastchkptdatachecksums = checkPoint.dataChecksumState;
break;
}
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 466f4223501..9b99e628d4e 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -145,6 +145,7 @@ main(int argc, char **argv)
XLogRecPtr chkptrec;
TimeLineID chkpttli;
XLogRecPtr chkptredo;
+ uint32 chkptdatachecksums;
TimeLineID source_tli;
TimeLineID target_tli;
XLogRecPtr target_wal_endrec;
@@ -472,10 +473,51 @@ main(int argc, char **argv)
keepwal_init();
findLastCheckpoint(datadir_target, divergerec, lastcommontliIndex,
- &chkptrec, &chkpttli, &chkptredo, restore_command);
+ &chkptrec, &chkpttli, &chkptredo, &chkptdatachecksums,
+ restore_command);
pg_log_info("rewinding from last common checkpoint at %X/%08X on timeline %u",
LSN_FORMAT_ARGS(chkptrec), chkpttli);
+ /*
+ * Replay on the rewound server resumes from the last common checkpoint
+ * and adopts the data checksum state recorded there, not the state in the
+ * control file installed from the source. If checksums were enabled at
+ * the divergence point but the source runs without them, the rewound
+ * server would verify checksums while the blocks copied from the source
+ * have none. The control files cannot reveal this: an offline disable on
+ * the source after the divergence leaves both of them saying "off".
+ *
+ * Only the fully enabled state needs checking. An in-progress state at
+ * the divergence point means a transition was still running there, and
+ * all of its page rewrites are logged after that point, so replay brings
+ * the rewound server to whatever state the copied WAL ends in.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the source server or recreate the target server from a base backup.");
+ exit(1);
+ }
+
+ /*
+ * The same comparison is needed against the target: every block the
+ * rewind does not copy keeps the target's content. The control files
+ * cannot reveal this case either, in the other direction: a standby's
+ * checkpoints are written by its upstream primary, so a standby whose
+ * checksums were disabled offline still has "on" checkpoints in its WAL,
+ * while both control files may agree.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_target.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the target server");
+ pg_log_error_detail("Blocks kept from the target would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the target server with pg_checksums or recreate the target server from a base backup.");
+ exit(1);
+ }
+
/* Initialize the hash table to track the status of each file */
filehash_init();
@@ -799,6 +841,55 @@ sanityChecks(void)
pg_fatal("target server needs to use either data checksums or \"wal_log_hints = on\"");
}
+ /*
+ * The rewound target keeps its control file fields from the source, but
+ * every block the rewind does not copy keeps the target's content. If
+ * checksums are enabled on the target and disabled on the source, the
+ * result would claim enabled checksums while the blocks copied from the
+ * source have none, and the target would fail checksum verification as
+ * soon as it reads them. Refuse that combination, and refuse an
+ * interrupted online transition on either side, same as pg_checksums.
+ *
+ * The opposite mismatch is allowed: recovery under the backup label
+ * written by pg_rewind adopts the checksum state as of the divergence
+ * point, so a target whose own checkpoints carry no checksums stays
+ * without them, or converges through the replayed WAL if the source
+ * enabled them online. Only warn about it, so that an offline change on
+ * the source is not overlooked.
+ *
+ * That reasoning does not hold when the target is a standby, whose
+ * checkpoints were written by its upstream primary; see the checks after
+ * findLastCheckpoint().
+ */
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted on the target server");
+ pg_log_error_hint("Start the server and shut it down cleanly to reset the state, then retry; on a standby, let replication complete the transition first.");
+ exit(1);
+ }
+ if (ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition is incomplete on the source server");
+ pg_log_error_hint("Let the transition complete, or reset the state with a clean restart, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_OFF)
+ {
+ pg_log_error("data checksums are enabled on the target server but disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, and the target would fail checksum verification after the rewind.");
+ pg_log_error_hint("Either disable data checksums on the target server or enable them on the source server, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_OFF &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_warning("data checksums are disabled on the target server but enabled on the source server");
+ pg_log_warning_detail("The rewound server will keep data checksums disabled unless the source server enabled them online after the point of divergence.");
+ }
+
/*
* Target cluster better not be running. This doesn't guard against
* someone starting the cluster concurrently. Also, this is probably more
diff --git a/src/bin/pg_rewind/pg_rewind.h b/src/bin/pg_rewind/pg_rewind.h
index 9a981f7f246..9187c3bf72f 100644
--- a/src/bin/pg_rewind/pg_rewind.h
+++ b/src/bin/pg_rewind/pg_rewind.h
@@ -39,6 +39,7 @@ extern void findLastCheckpoint(const char *datadir, XLogRecPtr forkptr,
int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
XLogRecPtr *lastchkptredo,
+ uint32 *lastchkptdatachecksums,
const char *restoreCommand);
extern XLogRecPtr readOneRecord(const char *datadir, XLogRecPtr ptr,
int tliIndex, const char *restoreCommand);
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index e5d38fafb7d..7d07c757052 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -45,6 +45,8 @@ tests += {
't/019_standby_shutdown_catchup.pl',
't/020_cascade_divergence.pl',
't/021_rewind_divergent_transitions.pl',
+ 't/022_rewind_state.pl',
+ 't/023_rewind_standby_target.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/022_rewind_state.pl b/src/test/modules/test_checksums/t/022_rewind_state.pl
new file mode 100644
index 00000000000..b7adc968c4d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/022_rewind_state.pl
@@ -0,0 +1,133 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind checks the data checksum states of source and target.
+# Checksums enabled on the target, or at the point of divergence, with
+# a source running without them are refused: the rewound server would
+# verify checksums on blocks copied from a source that has none. The
+# opposite mismatch only warns; replay keeps the target's own state.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# Scenarios 1 and 2: checksums on from initdb. wal_log_hints keeps the
+# target eligible for pg_rewind after checksums are disabled on it.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b);
+
+# Failover to B, and divergence on A.
+$node_b->promote;
+$node_b->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_a->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_a->stop('fast');
+
+# Offline disable on the source only: enabled target, disabled source.
+$node_b->stop;
+$node_b->checksum_disable_offline;
+$node_b->start;
+test_checksum_state($node_b, 'off');
+
+my @rewind_cmd = (
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'));
+
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums are enabled on the target server but disabled on the source server/,
+ 'refuses an enabled target with a disabled source');
+
+# The refusal happens before any modification: the target still starts
+# on its own timeline.
+$node_a->start;
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'target untouched by the refused rewind');
+$node_a->stop('fast');
+
+# Scenario 2: disabling the target too makes the control files match,
+# but checksums were still enabled at the point of divergence, which is
+# the state replay on the rewound server would resume with.
+$node_a->checksum_disable_offline;
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums were enabled at the point of divergence but are disabled on the source server/,
+ 'refuses when checksums were enabled at the point of divergence');
+
+# Scenario 3: fresh pair without checksums; offline enable on the
+# source only. Allowed with a warning, and the rewound server keeps
+# checksums disabled.
+my $node_c = PostgreSQL::Test::Cluster->new('node_c');
+$node_c->init(allows_streaming => 1, no_data_checksums => 1);
+$node_c->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_c->start;
+$node_c->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_c->backup('backup');
+my $node_d = PostgreSQL::Test::Cluster->new('node_d');
+$node_d->init_from_backup($node_c, 'backup', has_streaming => 1);
+$node_d->start;
+$node_c->wait_for_catchup($node_d);
+
+$node_d->promote;
+$node_d->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_c->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_c->stop('fast');
+
+$node_d->stop;
+$node_d->checksum_enable_offline;
+$node_d->start;
+test_checksum_state($node_d, 'on');
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_c->data_dir,
+ '--source-server' => $node_d->connstr('postgres'),
+ ]);
+like(
+ $stderr,
+ qr/data checksums are disabled on the target server but enabled on the source server/,
+ 'warns for a disabled target with an enabled source');
+like($stderr, qr/Done!/, 'rewind completed despite the warning');
+
+# The rewound server follows D and keeps its own state.
+$node_c->append_conf('postgresql.conf', 'port = ' . $node_c->port);
+$node_c->enable_streaming($node_d);
+$node_c->set_standby_mode;
+$node_c->start;
+$node_d->wait_for_catchup($node_c);
+is($node_c->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+test_checksum_state($node_c, 'off');
+
+$node_c->stop;
+$node_d->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/023_rewind_standby_target.pl b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
new file mode 100644
index 00000000000..3f5a3c6be8b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
@@ -0,0 +1,141 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_rewind compares the data checksum state at the point of
+# divergence against the target as well as the source.
+#
+# A standby's checkpoints are written by its upstream primary, so a standby
+# whose checksums were disabled offline still has "on" checkpoints in its
+# WAL. Rewinding it onto a promoted sibling passes every control file
+# check (target off + source on only warns), yet recovery after the rewind
+# adopts the divergence checkpoint's state and the rewound server would
+# verify checksums over the pages it wrote while it was locally "off".
+# pg_rewind must refuse; after an offline enable on the target, which
+# rewrites all of its pages, the rewind goes through.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t0 AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+
+my $lagging = PostgreSQL::Test::Cluster->new('lagging');
+$lagging->init_from_backup($primary, 'backup', has_streaming => 1);
+$lagging->start;
+
+my $failover = PostgreSQL::Test::Cluster->new('failover');
+$failover->init_from_backup($primary, 'backup', has_streaming => 1);
+$failover->start;
+
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+
+test_checksum_state($primary, 'on');
+test_checksum_state($lagging, 'on');
+test_checksum_state($failover, 'on');
+
+# Offline-disable checksums on the lagging standby only. This is a
+# divergence 012_offline_standby.pl declares survivable: the standby keeps
+# its own state, warns, and stays readable.
+$lagging->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $lagging->data_dir);
+$lagging->start;
+test_checksum_state($lagging, 'off');
+
+# Give the lagging standby plenty of pages to write while it is locally
+# "off", and force a restartpoint so they reach disk without checksums.
+# The other standby replays the same WAL with checksums on.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+$lagging->safe_psql('postgres', "CHECKPOINT;");
+
+# Take the "failover" standby out of the picture. It is promoted from
+# this position later, so everything replayed from here on is the part of
+# the WAL that diverges and that pg_rewind will roll back. t1 is already
+# replayed everywhere, so its blocks are *not* rolled back.
+$failover->stop;
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$primary->wait_for_catchup($lagging);
+
+# Failover: promote the other standby, which forks the timeline behind the
+# replay position of the lagging standby.
+$primary->stop('fast');
+$failover->start;
+$failover->promote;
+$failover->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$failover->safe_psql('postgres', "CHECKPOINT;");
+
+# Re-attaching the lagging standby with pg_rewind must be refused: the
+# divergence checkpoint says "on" while the target is "off", so the rewound
+# server would verify checksums over the blocks it keeps.
+$lagging->stop;
+
+command_fails_like(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ],
+ qr/data checksums were enabled at the point of divergence but are disabled on the target server/,
+ 'pg_rewind refuses a target whose divergence checkpoint has checksums');
+
+# Enabling checksums offline on the target rewrites all of its pages, after
+# which the rewind is safe.
+system_or_bail('pg_checksums', '--enable', '--pgdata', $lagging->data_dir);
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ]);
+like($stderr, qr/Done!/, 'pg_rewind completes after the offline enable');
+
+$lagging->append_conf('postgresql.conf', 'port = ' . $lagging->port);
+$lagging->enable_streaming($failover);
+$lagging->set_standby_mode;
+$lagging->start;
+
+my ($rc, $out, $err) =
+ $lagging->psql('postgres',
+ "SELECT setting FROM pg_settings " . "WHERE name = 'data_checksums';");
+is($rc, 0, 'rewound standby accepts connections') or diag("stderr: $err");
+is($out, 'on', 'rewound standby resumes with checksums on');
+
+($rc, $out, $err) = $lagging->psql('postgres', "SELECT count(*) FROM t1;");
+is($rc, 0, 'blocks kept from the target stay readable')
+ or diag("stderr: $err");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($lagging->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the rewound standby');
+
+$lagging->stop('immediate');
+$failover->stop;
+done_testing();
--
2.55.0
[application/octet-stream] v11-0001-Do-not-adopt-data-checksum-state-from-another-no.patch (133.6K, ../../CAN4CZFP_-vMOtHEj0u_v9OYoihxmi1QkfP_wiP_ytdg9FE=KfA@mail.gmail.com/3-v11-0001-Do-not-adopt-data-checksum-state-from-another-no.patch)
download | inline diff:
From e6db355c8075ced39d3b1eec5c802805725fbff9 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:02 +0000
Subject: [PATCH v11 1/4] Do not adopt data checksum state from another node
during replay
Offline pg_checksums changes are local to one node, but replay adopted
the data checksum state carried by checkpoint records unconditionally.
After an offline enable on the primary, a standby whose pages were
never rewritten started verifying checksums it does not have; after an
offline change on a standby, the next replayed checkpoint silently
reverted it.
To fix, make the control file track this node's state alone, and have
replay cross-check the replayed state against it instead of adopting
it, warning once per divergent value and reporting when the states
agree again. pg_control gains a watermark, normally the end LSN of
the newest XLOG2_CHECKSUMS record the node has written or applied, so
that replay can skip transition records whose effect the control file
already contains, and a flag marking a state last written by
pg_checksums, which recovery must never overwrite with a replayed one.
Since the control file may only claim "on" once every page on disk
carries a checksum, persisting the state is tied to flushes:
XLOG2_CHECKSUMS replay persists every state but "on" immediately and
leaves "on" to the next restartpoint, checkpoints and restartpoints
only persist a state their flush ran under from beginning to end, and
transitions publish their state under the new
DataChecksumTransitionLock so that the states carried by WAL records
match their WAL order. The comments in xlog.c spell out the
individual rules.
Recovery from a base backup is the exception to not adopting: its
control file was copied at an arbitrary moment, so the state carried
by the starting checkpoint is the one the WAL from there on was
written under. pg_rewind keeps the target's own state and clamps the
watermark to the divergence point, since a watermark set on the
target's abandoned fork could numerically cover transition records the
source wrote after the divergence.
Bump PG_CONTROL_VERSION.
Document the offline procedure for replication setups, the lockstep
one: stop all nodes, run pg_checksums on each of them, and only then
restart. An offline change writes no WAL and has no ordering against
WAL a node has not replayed yet, so a node must not stop before
replaying all WAL of its upstream, or the change is overridden on
restart.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Author: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
---
doc/src/sgml/ref/pg_checksums.sgml | 87 ++-
doc/src/sgml/wal.sgml | 23 +
src/backend/access/transam/xlog.c | 570 ++++++++++++++++--
src/backend/postmaster/datachecksum_state.c | 26 +-
.../utils/activity/wait_event_names.txt | 1 +
src/bin/pg_checksums/pg_checksums.c | 10 +
src/bin/pg_controldata/pg_controldata.c | 4 +
src/bin/pg_resetwal/pg_resetwal.c | 7 +
src/bin/pg_rewind/pg_rewind.c | 35 +-
src/bin/pg_upgrade/controldata.c | 2 +-
src/include/catalog/pg_control.h | 29 +-
src/include/storage/lwlocklist.h | 1 +
src/test/modules/test_checksums/Makefile | 2 +-
src/test/modules/test_checksums/meson.build | 10 +
.../test_checksums/t/012_offline_standby.pl | 289 +++++++++
.../modules/test_checksums/t/013_rewind.pl | 201 ++++++
.../modules/test_checksums/t/014_lockstep.pl | 182 ++++++
.../t/015_standby_crash_after_disable.pl | 125 ++++
.../t/016_promote_enable_crash.pl | 134 ++++
.../test_checksums/t/017_restartpoint_race.pl | 138 +++++
.../t/018_enable_crash_windows.pl | 496 +++++++++++++++
.../t/019_standby_shutdown_catchup.pl | 126 ++++
.../t/020_cascade_divergence.pl | 115 ++++
.../t/021_rewind_divergent_transitions.pl | 193 ++++++
24 files changed, 2722 insertions(+), 84 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/012_offline_standby.pl
create mode 100644 src/test/modules/test_checksums/t/013_rewind.pl
create mode 100644 src/test/modules/test_checksums/t/014_lockstep.pl
create mode 100644 src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
create mode 100644 src/test/modules/test_checksums/t/016_promote_enable_crash.pl
create mode 100644 src/test/modules/test_checksums/t/017_restartpoint_race.pl
create mode 100644 src/test/modules/test_checksums/t/018_enable_crash_windows.pl
create mode 100644 src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
create mode 100644 src/test/modules/test_checksums/t/020_cascade_divergence.pl
create mode 100644 src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index ae66fad3f0f..1cb68187e51 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -242,15 +242,84 @@ PostgreSQL documentation
data directory must not be started or else data loss may occur.
</para>
<para>
- When using a replication setup with tools which perform direct copies
- of relation file blocks (for example <xref linkend="app-pgrewind"/>),
- enabling or disabling checksums can lead to page corruptions in the
- shape of incorrect checksums if the operation is not done consistently
- across all nodes. When enabling or disabling checksums in a replication
- setup, it is thus recommended to stop all the clusters before switching
- them all consistently. Destroying all standbys, performing the operation
- on the primary and finally recreating the standbys from scratch is also
- safe.
+ Enabling or disabling checksums with
+ <application>pg_checksums</application> changes only the local data
+ directory; the new state is not replicated to any other node. In a
+ replication setup the same change must be applied to every node:
+ </para>
+
+ <procedure>
+ <step>
+ <title>Shut down all nodes</title>
+ <para>
+ All nodes participating in the replication must be stopped with a
+ clean shutdown; <application>pg_checksums</application> refuses to run
+ on a data directory left behind by an immediate shutdown. Before
+ stopping a standby, make sure it has replayed all WAL of the primary:
+ stop the primary first, read its <quote>Latest checkpoint
+ location</quote> with <xref linkend="app-pgcontroldata"/>, and check
+ that <function>pg_last_wal_replay_lsn()</function> on the standby has
+ advanced past it. Comparing the replay position with
+ <function>pg_last_wal_receive_lsn()</function> is not enough, as it
+ only shows that the WAL the standby received has been replayed.
+ </para>
+ </step>
+
+ <step>
+ <title>Enable or disable data checksums on each node</title>
+ <para>
+ Run <application>pg_checksums</application> on the data directory of
+ each node in the replication setup. Nodes can be processed in
+ parallel as they are shut down. Processing must have ended
+ successfully on all nodes before continuing.
+ </para>
+ </step>
+
+ <step>
+ <title>Restart all nodes</title>
+ <para>
+ Start the nodes normally, verify that
+ <xref linkend="guc-data-checksums"/> matches on all of them, and
+ monitor the logs of the standbys for data checksum state mismatch
+ warnings.
+ </para>
+ </step>
+ </procedure>
+
+ <para>
+ Tools that copy relation file blocks directly between nodes, such as
+ <xref linkend="app-pgrewind"/>, likewise require all nodes to be in
+ the same data checksum state.
+ </para>
+ <para>
+ The replay requirement exists because an offline change is recorded
+ only in the control file and has no defined ordering against WAL the
+ node has not replayed yet; see
+ <xref linkend="checksums-offline-enable-disable"/>. A node stopped
+ before replaying an online checksum state change applies that change
+ when it is restarted, overriding the offline change, and the states of
+ the nodes silently diverge until a later checkpoint record triggers
+ the warning described below. Because of this it is best not to mix
+ the two mechanisms: change the state of a replication setup either
+ with the offline procedure above or with an online transition, and
+ make sure the previous change has reached every node before starting
+ the next one.
+ </para>
+ <para>
+ If the change is applied inconsistently, each node keeps its own state,
+ and a standby logs a warning when the state recorded in the replayed WAL
+ differs from its own. The same warning can appear transiently while a
+ standby catches up over WAL written before a consistent change; it stops
+ once a checkpoint record carrying the new state has been replayed.
+ </para>
+ <para>
+ A standby whose data directory was never checksummed must not have
+ checksums enabled by catching up this way. Converge the cluster by
+ running <application>pg_checksums</application> on it while stopped as
+ outlined above, by enabling checksums online, or by recreating it from
+ a base backup. Note
+ that an online enable only starts from a primary whose checksums are off,
+ so if they are already enabled there, disable them online first.
</para>
<para>
If <application>pg_checksums</application> is aborted or killed while
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index ec62d17fbcc..dbbfdae5e4f 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -317,6 +317,29 @@
verify checksums, on an offline cluster.
</para>
+ <para>
+ An offline change provides durability differently from an
+ <link linkend="checksums-online-enable-disable">online change</link>.
+ An online transition is WAL-logged: it is ordered against all other
+ WAL records, it is replayed after a crash, and it propagates to
+ standbys. An offline change is recorded only in the cluster's
+ control file: it writes no WAL, it is invisible to replication, and
+ it has no defined ordering against WAL the node has not replayed
+ yet. When a node later replays WAL that contains an online checksum
+ state change, that change takes effect on the node even if it was
+ written before the offline change was made.
+ </para>
+
+ <para>
+ An offline change only affects the data directory it is run on; the
+ new state does not propagate over replication. In a replication setup
+ the same change must be applied to all nodes while all of them are
+ stopped, as described in <xref linkend="app-pgchecksums"/>. A standby
+ whose state diverges logs a warning but keeps its local setting. The
+ mismatch persists until the states are brought together again, with
+ the offline procedure or with an online transition; do this promptly.
+ </para>
+
</sect2>
<sect2 id="checksums-online-enable-disable" xreflabel="Online Enabling of Checksums">
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index de4c96e135f..7e877db22a2 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -556,14 +556,22 @@ typedef struct XLogCtlData
*/
XLogRecPtr lastFpwDisableRecPtr;
- /* last data_checksum_version we've seen */
+ /* current data checksum state of this node */
uint32 data_checksum_version;
+ /*
+ * Copies of control file fields with the same names, see pg_control.h for
+ * an in-depth description of these fields. Must be updated together with
+ * data_checksum_version under info_lck.
+ */
+ XLogRecPtr data_checksum_lsn;
+ bool data_checksum_is_local;
+
slock_t info_lck; /* locks shared variables shown above */
/*
* lastChecksumChangeRecPtr points to the end of the last XLOG2_CHECKSUMS
- * record inserted or replayed, i.e. the last change of
+ * record inserted or replayed which corresponds to the last change of
* data_checksum_version. InvalidXLogRecPtr if the state hasn't changed
* since the server started.
*/
@@ -690,6 +698,20 @@ static ChecksumStateType LocalDataChecksumState = 0;
*/
int data_checksums = 0;
+/*
+ * Whether replay of the next checkpoint-family record must adopt the data
+ * checksum state it carries. Set when recovery starts from a base backup,
+ * where the state at the redo point takes precedence over the control file
+ * copied with the backup later.
+ */
+static bool adoptChecksumStateFromNextCheckpoint = false;
+
+/*
+ * Sentinel for the data checksum mismatch warning tracking in
+ * CheckReplayedDataChecksumState(): no warning is outstanding.
+ */
+#define NO_WARNING_ISSUED PG_UINT32_MAX
+
/* For WALInsertLockAcquire/Release functions */
static int MyLockNo = 0;
static bool holdingAllLocks = false;
@@ -730,6 +752,8 @@ static void ValidateXLOGDirectoryStructure(void);
static void CleanupBackupHistory(void);
static void UpdateMinRecoveryPoint(XLogRecPtr lsn, bool force);
static bool PerformRecoveryXLogAction(void);
+static void CheckReplayedDataChecksumState(uint32 replayed_version);
+static void AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn);
static void InitControlFile(uint64 sysidentifier, uint32 data_checksum_version);
static void WriteControlFile(void);
static void ReadControlFile(void);
@@ -757,7 +781,7 @@ static void WALInsertLockAcquireExclusive(void);
static void WALInsertLockRelease(void);
static void WALInsertLockUpdateInsertingAt(XLogRecPtr insertingAt);
-static void XLogChecksums(uint32 new_type);
+static XLogRecPtr XLogChecksums(uint32 new_type);
/*
* Insert an XLOG record represented by an already-constructed chain of data
@@ -4292,6 +4316,8 @@ InitControlFile(uint64 sysidentifier, uint32 data_checksum_version)
ControlFile->track_commit_timestamp = track_commit_timestamp;
ControlFile->data_checksum_version = data_checksum_version;
ControlFile->data_checksum_version_init = data_checksum_version;
+ ControlFile->data_checksum_is_local = false;
+ ControlFile->data_checksum_lsn = InvalidXLogRecPtr;
/*
* Set the data_checksum_version value into XLogCtl, which is where all
@@ -4774,6 +4800,7 @@ void
SetDataChecksumsOnInProgress(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
/*
* The state transition is performed in a critical section with
@@ -4782,14 +4809,12 @@ SetDataChecksumsOnInProgress(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4827,6 +4852,8 @@ void
SetDataChecksumsOn(void)
{
uint64 barrier;
+ bool persist;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4847,23 +4874,11 @@ SetDataChecksumsOn(void)
SpinLockRelease(&XLogCtl->info_lck);
INJECTION_POINT("datachecksums-enable-checksums-delay", NULL);
+ INJECTION_POINT_LOAD("datachecksums-on-before-publish");
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_VERSION);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- SpinLockRelease(&XLogCtl->info_lck);
-
- /*
- * Update the controlfile before waiting since if we have an immediate
- * shutdown while waiting we want to come back up with checksums enabled.
- */
- LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- UpdateControlFile();
- LWLockRelease(ControlFileLock);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_VERSION);
barrier = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_CHECKSUM_ON);
@@ -4873,6 +4888,40 @@ SetDataChecksumsOn(void)
INJECTION_POINT("datachecksums-on-before-checkpoint", NULL);
RequestCheckpoint(CHECKPOINT_FORCE | CHECKPOINT_WAIT | CHECKPOINT_FAST);
+
+ INJECTION_POINT("datachecksums-on-after-checkpoint", NULL);
+
+ /*
+ * Persist "on" only now that the checkpoint has flushed the pages the
+ * transition rewrote. Crash recovery initializes verification from the
+ * control file but resumes from a checkpoint that can predate the
+ * transition, so an "on" persisted earlier would verify pages during
+ * replay whose rewrite never reached the disk: the pages found by the data
+ * checksums worker in shared buffers are not written back by its ring
+ * buffer.
+ *
+ * The checkpoint above normally persists the state itself; this covers the
+ * case where it started before the record written above and left the field
+ * alone. Crashing before this point is safe, as replay then re-establishes
+ * "on" from the full page images of the rewrite. Skip the write if the
+ * state moved on meanwhile, since whatever moved it persists its own.
+ * Compare the watermark rather than the state: a state comparison could
+ * not tell our transition from a later round trip back to "on".
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ persist = (XLogCtl->data_checksum_lsn == recptr);
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (persist)
+ {
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+ }
+
WaitForProcSignalBarrier(barrier);
}
@@ -4893,6 +4942,7 @@ void
SetDataChecksumsOff(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4918,14 +4968,12 @@ SetDataChecksumsOff(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4956,14 +5004,12 @@ SetDataChecksumsOff(void)
/* Ensure that we don't incur a checkpoint during disabling checksums */
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -5001,6 +5047,138 @@ SetLocalDataChecksumState(uint32 data_checksum_version)
data_checksums = data_checksum_version;
}
+/*
+ * CheckReplayedDataChecksumState
+ * Cross-check the data checksum state carried by a replayed checkpoint
+ * record against the state of this node.
+ *
+ * The state in the record belongs to whichever node wrote the WAL, and must
+ * not be adopted: an offline state change made with pg_checksums on one node
+ * of a replication set generates no WAL, and adopting would leak it into the
+ * other nodes through replay. Backup label recovery is the one exception,
+ * see AdoptReplayedDataChecksumState(). XLOG_CHECKPOINT_ONLINE needs no call
+ * here, as an online checkpoint's state already traveled in the preceding
+ * XLOG_CHECKPOINT_REDO record.
+ *
+ * Only archive recovery can see a lasting mismatch, as only there can the WAL
+ * and the control file come from different nodes or different times. In
+ * crash recovery a mismatch means replay resumed from a restartpoint
+ * predating an already-applied XLOG2_CHECKSUMS record, and replaying forward
+ * re-establishes the same state.
+ */
+static void
+CheckReplayedDataChecksumState(uint32 replayed_version)
+{
+ /*
+ * Warn once per remote value, so a lasting mismatch does not flood the
+ * log. Matching states re-arm the warning. Backend-local state is
+ * enough: replay only runs in the startup process, and restarting it
+ * re-arms as well.
+ */
+ static uint32 last_warned_version = NO_WARNING_ISSUED;
+ uint32 local_version;
+
+ if (!ArchiveRecoveryRequested)
+ return;
+
+ /*
+ * Re-replayed WAL below the consistency point was already cross-checked
+ * before minRecoveryPoint was last persisted, and the persisted state can
+ * legitimately be newer than what checkpoint records there carry:
+ * XLOG2_CHECKSUMS replay persists most states ahead of the restartpoint
+ * horizon. In particular the checkpoint record recovery restarts from is
+ * such a re-replay.
+ */
+ if (!reachedConsistency)
+ return;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ local_version = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (replayed_version == local_version)
+ {
+ /*
+ * Report convergence if this process warned before. Nothing else
+ * tells the operator that running pg_checksums on the other nodes, or
+ * a rebuild, took effect.
+ */
+ if (last_warned_version != NO_WARNING_ISSUED)
+ ereport(LOG,
+ errmsg("data checksum state \"%s\" of this node now agrees with the replayed WAL",
+ get_checksum_state_string(local_version)));
+
+ last_warned_version = NO_WARNING_ISSUED;
+ return;
+ }
+
+ /* the nodes legitimately differ while an online transition runs */
+ if (replayed_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ replayed_version == PG_DATA_CHECKSUM_INPROGRESS_OFF ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ return;
+
+ if (replayed_version == last_warned_version)
+ return;
+ last_warned_version = replayed_version;
+
+ ereport(WARNING,
+ errmsg("data checksum state \"%s\" of this node does not match the state \"%s\" in the replayed WAL",
+ get_checksum_state_string(local_version),
+ get_checksum_state_string(replayed_version)),
+ errdetail("The data checksum state was most likely changed with pg_checksums on another node."),
+ errhint("Apply the same change with pg_checksums on the primary and all standby servers, or rebuild this server from a base backup."));
+}
+
+/*
+ * AdoptReplayedDataChecksumState
+ * Adopt the data checksum state at the redo point of backup label
+ * recovery.
+ *
+ * The state is persisted immediately so that a crash before the first
+ * restartpoint does not resurrect the state copied with the backup; a crash
+ * at this point restarts from the same redo point, so the control file does
+ * not run ahead of the replay position. If the value is unchanged the
+ * control file already carries it, so both the barrier and the persist are
+ * skipped.
+ *
+ * lsn is the location of the checkpoint-family record the state was taken
+ * from and becomes the new watermark: the adopted state covers everything
+ * below the redo point, which replay never revisits. For the same reason
+ * the old watermark can stay when the value is unchanged: the records
+ * between the two positions are never replayed, so nothing depends on
+ * which one is recorded.
+ */
+static void
+AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn)
+{
+ bool changed = false;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (XLogCtl->data_checksum_version != new_version)
+ {
+ XLogCtl->data_checksum_version = new_version;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(new_version);
+ changed = true;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (!changed)
+ return;
+
+ EmitAndWaitDataChecksumsBarrier(new_version);
+
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = new_version;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+}
+
/* guc hook */
const char *
show_data_checksums(void)
@@ -5454,6 +5632,8 @@ XLOGShmemInit(void *arg)
/* Use the checksum info from control file */
XLogCtl->data_checksum_version = ControlFile->data_checksum_version;
+ XLogCtl->data_checksum_lsn = ControlFile->data_checksum_lsn;
+ XLogCtl->data_checksum_is_local = ControlFile->data_checksum_is_local;
SetLocalDataChecksumState(XLogCtl->data_checksum_version);
SpinLockInit(&XLogCtl->Insert.insertpos_lck);
@@ -6031,6 +6211,52 @@ StartupXLOG(void)
SetCommitTsLimit(checkPoint.oldestCommitTsXid,
checkPoint.newestCommitTsXid);
+ /*
+ * When recovery starts from a base backup, the control file was copied at
+ * an arbitrary moment and its data checksum state may differ from the
+ * state at the redo point, which is what the WAL from there on was
+ * written under. Adopt the state of the starting checkpoint: a shutdown
+ * checkpoint is not replayed, so take it from the record read above; the
+ * redo point of an online checkpoint is its CHECKPOINT_REDO record, so
+ * let the replay of that record adopt it. Check backupStartPoint in
+ * addition to the label: on a crash restart during backup recovery the
+ * label file is already renamed away, but the start point persists until
+ * the backup end record.
+ *
+ * Not for a base backup taken from a standby, though. Its starting
+ * checkpoint is the standby's last restartpoint, a record written by the
+ * upstream primary, whose state is not the one the copied files were
+ * written under. The copied control file is already correct: a standby
+ * persists its state only at restartpoint horizons and never claims
+ * more than what reached disk. Such backups are recognized by
+ * backupEndPoint together with backupEndRequired; backupEndPoint is only
+ * set for "BACKUP FROM: standby" labels and persists across a crash
+ * restart. pg_rewind writes a standby label as well, but no
+ * backupEndPoint, and its recovery keeps adopting: the control file it
+ * installs carries the target's own checksum state, which can lag the
+ * redo point of the last common checkpoint the same way a restartpoint
+ * horizon can.
+ *
+ * Never adopt over a state the control file's watermark or local flag
+ * marks as newer than the starting checkpoint. A pg_checksums change is
+ * local to the node and generates no WAL, so nothing in the replayed WAL
+ * could ever restore it once overwritten; and a watermark above the redo
+ * point means the control file already contains the effect of every
+ * transition record up to there.
+ */
+ if ((haveBackupLabel || XLogRecPtrIsValid(ControlFile->backupStartPoint)) &&
+ !(XLogRecPtrIsValid(ControlFile->backupEndPoint) &&
+ ControlFile->backupEndRequired) &&
+ !ControlFile->data_checksum_is_local &&
+ checkPoint.redo > ControlFile->data_checksum_lsn)
+ {
+ if (wasShutdown)
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ checkPoint.redo);
+ else
+ adoptChecksumStateFromNextCheckpoint = true;
+ }
+
/*
* Clear out any old relcache cache files. This is *necessary* if we do
* any WAL replay, since that would probably result in the cache files
@@ -6633,11 +6859,7 @@ StartupXLOG(void)
if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
ereport(WARNING,
@@ -6654,11 +6876,7 @@ StartupXLOG(void)
else if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
}
@@ -6683,6 +6901,8 @@ StartupXLOG(void)
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
XLogCtl->SharedRecoveryState = RECOVERY_STATE_DONE;
SpinLockRelease(&XLogCtl->info_lck);
@@ -6814,6 +7034,23 @@ static bool
PerformRecoveryXLogAction(void)
{
bool promoted = false;
+ bool flushForChecksums;
+ uint32 checksum_state;
+
+ /*
+ * The end-of-recovery record persists the data checksum state without
+ * flushing the buffer pool, but the control file may only claim "on" once
+ * every page on disk carries a checksum. If replay entered that state
+ * without a restartpoint following it, the pages rewritten by the
+ * transition are still only in the buffer pool, so take the full
+ * checkpoint below instead of the lightweight record.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ flushForChecksums = (checksum_state == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile->data_checksum_version != checksum_state);
/*
* Perform a checkpoint to update all our recovery activity to disk.
@@ -6829,7 +7066,7 @@ PerformRecoveryXLogAction(void)
* fully out of recovery mode and already accepting queries.
*/
if (ArchiveRecoveryRequested && IsUnderPostmaster &&
- PromoteIsTriggered())
+ PromoteIsTriggered() && !flushForChecksums)
{
promoted = true;
@@ -7436,6 +7673,7 @@ CreateCheckPoint(int flags)
uint32 freespace;
XLogRecPtr PriorRedoPtr;
XLogRecPtr last_important_lsn;
+ XLogRecPtr checksumLsn;
VirtualTransactionId *vxids;
int nvxids;
int oldXLogAllowed = 0;
@@ -7548,11 +7786,14 @@ CreateCheckPoint(int flags)
checkPoint.wal_level = wal_level;
/*
- * Get the current data_checksum_version value from xlogctl, valid at the
- * time of the checkpoint.
+ * Get the current data_checksum_version value from xlogctl. This is
+ * final only for a shutdown checkpoint, where no concurrent transition
+ * is possible; an online checkpoint resamples it together with the redo
+ * record below.
*/
SpinLockAcquire(&XLogCtl->info_lck);
checkPoint.dataChecksumState = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
if (shutdown)
@@ -7609,10 +7850,21 @@ CreateCheckPoint(int flags)
{
xl_checkpoint_redo redo_rec;
+ /*
+ * Sample the data checksum state and insert the redo record under
+ * DataChecksumTransitionLock, so that a concurrent transition cannot
+ * insert its XLOG2_CHECKSUMS record between the sampling and the
+ * insertion below. Without this, the redo record could follow the
+ * transition record in WAL while carrying the pre-transition state,
+ * and recovery resuming here would never learn about the transition.
+ * See XLogChecksums().
+ */
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
WALInsertLockAcquire();
redo_rec.wal_level = wal_level;
SpinLockAcquire(&XLogCtl->info_lck);
redo_rec.data_checksum_version = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
WALInsertLockRelease();
@@ -7620,6 +7872,14 @@ CreateCheckPoint(int flags)
XLogBeginInsert();
XLogRegisterData(&redo_rec, sizeof(xl_checkpoint_redo));
(void) XLogInsert(RM_XLOG_ID, XLOG_CHECKPOINT_REDO);
+ LWLockRelease(DataChecksumTransitionLock);
+
+ /*
+ * The checkpoint record must carry the same state as the redo record
+ * just inserted: the sample taken before redo determination can be
+ * stale by now, and the pair would otherwise disagree.
+ */
+ checkPoint.dataChecksumState = redo_rec.data_checksum_version;
/*
* XLogInsertRecord will have updated XLogCtl->Insert.RedoRecPtr in
@@ -7823,6 +8083,40 @@ CreateCheckPoint(int flags)
ControlFile->minRecoveryPoint = InvalidXLogRecPtr;
ControlFile->minRecoveryPointTLI = 0;
+ /*
+ * Persist the data checksum state this node runs under. Only the
+ * top-level field tracks this node; ControlFile->checkPointCopy above is
+ * a historical record used to resume replay.
+ *
+ * checkPoint.dataChecksumState was sampled under
+ * DataChecksumTransitionLock together with the redo record, so it is the
+ * state in effect at the redo point. If it was "on",
+ * the XLOG2_CHECKSUMS record announcing that precedes the redo point and
+ * every page the transition rewrote was dirtied before it, so
+ * CheckPointGuts() has just written all of them out. Recording the state
+ * here is what keeps a finished transition from being resolved as
+ * interrupted when this checkpoint is the one crash recovery resumes
+ * from: replay never sees the record announcing it.
+ *
+ * Persist it only if the state did not change while the flush was in
+ * progress. If it changed in between, the pages written out straddle two
+ * states, and the newer one could claim checksums that pages already on
+ * disk do not carry; leave the field to the next checkpoint then, the
+ * transition itself has already persisted every state that is safe
+ * without a flush.
+ * Compare the watermark rather than the state: record positions are
+ * unique, so a full round trip back to the sampled state cannot alias,
+ * while its flushed pages straddle the intermediate states all the same.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksumLsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checkPoint.dataChecksumState;
+ ControlFile->data_checksum_lsn = checksumLsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
/*
* Persist unloggedLSN value. It's reset on crash recovery, so this goes
* unused on non-shutdown checkpoints, but seems useful to store it always
@@ -7967,9 +8261,11 @@ CreateEndOfRecoveryRecord(void)
ControlFile->minRecoveryPoint = recptr;
ControlFile->minRecoveryPointTLI = xlrec.ThisTimeLineID;
- /* start with the latest checksum version (as of the end of recovery) */
+ /* persist the data checksum state this node ended recovery with */
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
@@ -8174,6 +8470,9 @@ CreateRestartPoint(int flags)
XLogRecPtr endptr;
XLogSegNo _logSegNo;
TimestampTz xtime;
+ uint32 checksum_state;
+ XLogRecPtr checksum_lsn;
+ bool checksum_is_local;
/* Concurrent checkpoint/restartpoint cannot happen */
Assert(!IsUnderPostmaster || MyBackendType == B_CHECKPOINTER);
@@ -8220,8 +8519,47 @@ CreateRestartPoint(int flags)
UpdateMinRecoveryPoint(InvalidXLogRecPtr, true);
if (flags & CHECKPOINT_IS_SHUTDOWN)
{
+ bool catchUpChecksums;
+
+ /*
+ * There is no new restartpoint to persist the data checksum state
+ * with, but a cleanly stopped node should not leave the control
+ * file behind the state replay reached: pg_checksums and
+ * pg_rewind read it, and an in-progress state there makes them
+ * refuse to run. Catching it up needs the same guarantee a
+ * restartpoint gives: that every page on disk carries a checksum,
+ * so flush the buffer pool first. Only a transition to
+ * "on" that no restartpoint followed can get here; the other
+ * states are already persisted by XLOG2_CHECKSUMS replay. Replay
+ * has ended by now, so the state cannot change under us.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ checksum_is_local = XLogCtl->data_checksum_is_local;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockAcquire(ControlFileLock, LW_SHARED);
+ catchUpChecksums =
+ (checksum_lsn != ControlFile->data_checksum_lsn &&
+ XLogRecPtrIsValid(lastCheckPointRecPtr));
+ LWLockRelease(ControlFileLock);
+
+ if (catchUpChecksums)
+ {
+ MemSet(&CheckpointStats, 0, sizeof(CheckpointStats));
+ CheckpointStats.ckpt_start_t = GetCurrentTimestamp();
+ CheckPointGuts(lastCheckPoint.redo, flags);
+ }
+
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
+ if (catchUpChecksums)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = checksum_is_local;
+ }
UpdateControlFile();
LWLockRelease(ControlFileLock);
}
@@ -8262,6 +8600,17 @@ CreateRestartPoint(int flags)
/* Update the process title */
update_checkpoint_display(flags, true, false);
+ /*
+ * Note the data checksum state the flush below starts under. Replay runs
+ * concurrently and can change the state while the flush is in progress,
+ * in which case the flush covers pages written under both states; see
+ * where the state is persisted further down.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
CheckPointGuts(lastCheckPoint.redo, flags);
/*
@@ -8321,8 +8670,26 @@ CreateRestartPoint(int flags)
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
}
- /* we shall start with the latest checksum version */
- ControlFile->data_checksum_version = lastCheckPoint.dataChecksumState;
+ /*
+ * Persist the data checksum state of this node. Not the state of the
+ * replayed checkpoint: that one belongs to the node that wrote it and
+ * may differ after an offline change on either side.
+ * ControlFile->checkPointCopy above keeps the replayed value on
+ * purpose, being a historical record used to resume replay rather
+ * than a tracker of node state.
+ *
+ * Persist only if the flush above ran under one state throughout; see
+ * CreateCheckPoint() for why, including why this compares the
+ * watermark and not the state.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksum_lsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
}
@@ -8763,9 +9130,22 @@ XLogReportParameters(void)
}
/*
- * Log the new state of checksums
+ * XLogChecksums
+ * Log and publish the new state of checksums
+ *
+ * Inserting the record and publishing the new state must be atomic with
+ * respect to a checkpoint sampling the state for its XLOG_CHECKPOINT_REDO
+ * record: without that, a checkpoint could read the old state after the
+ * record is already in WAL and insert a redo record that both precedes the
+ * transition in WAL order and carries the pre-transition state. Recovery
+ * resuming from such a redo point would never replay the transition record
+ * and resolve the finished transition as interrupted.
+ * DataChecksumTransitionLock serializes the two; see CreateCheckPoint().
+ *
+ * Returns the end LSN of the inserted record, which the caller persists
+ * together with the new state as the data checksum watermark.
*/
-static void
+static XLogRecPtr
XLogChecksums(uint32 new_type)
{
xl_checksum_state xlrec;
@@ -8773,12 +9153,28 @@ XLogChecksums(uint32 new_type)
xlrec.new_checksum_state = new_type;
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
+
XLogBeginInsert();
XLogRegisterData((char *) &xlrec, sizeof(xl_checksum_state));
recptr = XLogInsert(RM_XLOG2_ID, XLOG2_CHECKSUMS);
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, recptr);
+
+ /* only loaded by SetDataChecksumsOn(), a no-op for the other callers */
+ INJECTION_POINT_CACHED("datachecksums-on-before-publish", NULL);
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ XLogCtl->data_checksum_version = new_type;
+ XLogCtl->data_checksum_lsn = recptr;
+ XLogCtl->data_checksum_is_local = false;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockRelease(DataChecksumTransitionLock);
+
XLogFlush(recptr);
+
+ return recptr;
}
/*
@@ -8966,11 +9362,19 @@ xlog_redo(XLogReaderState *record)
/* ControlFile->checkPointCopy always tracks the latest ckpt XID */
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->checkPointCopy.nextXid = checkPoint.nextXid;
- ControlFile->data_checksum_version = checkPoint.dataChecksumState;
UpdateControlFile();
LWLockRelease(ControlFileLock);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(checkPoint.dataChecksumState);
+
/*
* We should've already switched to the new TLI before replaying this
* record.
@@ -9206,19 +9610,17 @@ xlog_redo(XLogReaderState *record)
else if (info == XLOG_CHECKPOINT_REDO)
{
xl_checkpoint_redo redo_rec;
- bool new_state = false;
memcpy(&redo_rec, XLogRecGetData(record), sizeof(xl_checkpoint_redo));
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = redo_rec.data_checksum_version;
- SetLocalDataChecksumState(redo_rec.data_checksum_version);
- if (redo_rec.data_checksum_version != ControlFile->data_checksum_version)
- new_state = true;
- SpinLockRelease(&XLogCtl->info_lck);
-
- if (new_state)
- EmitAndWaitDataChecksumsBarrier(redo_rec.data_checksum_version);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(redo_rec.data_checksum_version,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(redo_rec.data_checksum_version);
}
else if (info == XLOG_LOGICAL_DECODING_STATUS_CHANGE)
{
@@ -9280,25 +9682,63 @@ xlog2_redo(XLogReaderState *record)
{
xl_checksum_state state;
XLogRecPtr lsn = record->EndRecPtr;
+ XLogRecPtr watermark;
memcpy(&state, XLogRecGetData(record), sizeof(xl_checksum_state));
+ SpinLockAcquire(&XLogCtl->info_lck);
+ watermark = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ /*
+ * Skip records this node has already applied. The control file
+ * carries the watermark, so this holds across restarts: recovery
+ * resuming below a record whose effect the control file already
+ * contains must not re-apply it, or it would revert a state change
+ * made with pg_checksums in between, which moves the state without
+ * writing any record of its own.
+ */
+ if (lsn <= watermark)
+ return;
+
/* advertise the location before the new state becomes visible */
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, lsn);
SpinLockAcquire(&XLogCtl->info_lck);
XLogCtl->data_checksum_version = state.new_checksum_state;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(state.new_checksum_state);
SpinLockRelease(&XLogCtl->info_lck);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = state.new_checksum_state;
+
+ /*
+ * Persist the new state, except when it is "on". Only "on" verifies
+ * checksums during reads, and between the last restartpoint and this
+ * record there may be pages on disk flushed under the old state; a
+ * crash-restart initializes verification from the control file and
+ * replay reads those pages back, so the control file may only say
+ * "on" once everything written under the transition has been flushed,
+ * as restartpoints and the end of recovery do.
+ * The opposite direction cannot wait for the restartpoint: once this
+ * record is replayed, evicted pages are written without checksums,
+ * and a control file still saying "on" would fail verification on
+ * exactly those pages after a crash.
+ */
+ if (state.new_checksum_state != PG_DATA_CHECKSUM_VERSION)
+ {
+ ControlFile->data_checksum_version = state.new_checksum_state;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ }
/*
* Update minRecoveryPoint to ensure that if recovery is aborted, we
* recover back up to this point before allowing hot standby again.
- * The new state is durable in pg_control while its location is only
- * tracked in shared memory; a standby becoming consistent below this
- * record would let base backups resume checksum verification with the
+ * The change location is only tracked in shared memory and is lost
+ * over a restart; a standby becoming consistent below this record
+ * would let base backups resume checksum verification with the
* location unknown. The local copies cannot be updated as long as
* crash recovery is happening and we expect all the WAL to be
* replayed.
diff --git a/src/backend/postmaster/datachecksum_state.c b/src/backend/postmaster/datachecksum_state.c
index f69258bc33d..2bdb3c75d9d 100644
--- a/src/backend/postmaster/datachecksum_state.c
+++ b/src/backend/postmaster/datachecksum_state.c
@@ -94,9 +94,9 @@
*
* If processing is started in an online cluster then all backends are in Bd.
* If processing was halted by the cluster shutting down (due to a crash or
- * intentional restart), the controlfile state "inprogress-on" will be observed
- * on system startup and all backends will be placed in Bd. The controlfile
- * state will also be set to "off".
+ * intentional restart), the control file state "inprogress-on" will be
+ * observed on system startup and all backends will be placed in Bd. The
+ * control file state will also be set to "off".
*
* Backends transition Bd -> Bi via a procsignalbarrier which is emitted by the
* DataChecksumsWorkerLauncherMain. When all backends have acknowledged the
@@ -146,7 +146,25 @@
* stop writing data checksums as no backend is enforcing data checksum
* validation any longer.
*
- * 4. Future opportunities for optimizations
+ * 4. Interaction with offline data checksum changes
+ * -------------------------------------------------
+ * Enabling or disabling checksums offline with pg_checksums uses none of the
+ * machinery in this file, but the two mechanisms share the state kept in the
+ * control file, so their interaction is documented here.
+ *
+ * pg_checksums writes the new state to the control file and sets
+ * data_checksum_is_local, marking a state that no WAL record accounts for.
+ * Recovery then does not adopt the state carried by a replayed checkpoint
+ * record over it. The control file also carries a watermark, the end LSN of
+ * the newest XLOG2_CHECKSUMS record this node has written or applied. Replay
+ * skips transition records at or below the watermark, as their effect is
+ * already contained in the control file, and applies records above it as
+ * usual, whether they were written before or after an offline change. This
+ * is why an offline change in a replicated setup must be made on every node
+ * while all of them are stopped and caught up; see the pg_checksums
+ * documentation for the procedure.
+ *
+ * 5. Future opportunities for optimizations
* -----------------------------------------
* Below are some potential optimizations and improvements which were brought
* up during reviews of this feature, but which weren't implemented in the
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 256b3a3c02e..77fa543f9d8 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -371,6 +371,7 @@ WaitLSN "Waiting to read or update shared Wait-for-LSN state."
LogicalDecodingControl "Waiting to read or update logical decoding status information."
DataChecksumsWorker "Waiting for data checksums worker."
AioWorkerControl "Waiting to update AIO worker information."
+DataChecksumTransition "Waiting for a data checksum state transition to be written to WAL."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 3b3ae23f1a6..8c4ae59222d 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -648,6 +648,16 @@ main(int argc, char *argv[])
ControlFile->data_checksum_version =
(mode == PG_MODE_ENABLE) ? PG_DATA_CHECKSUM_VERSION : PG_DATA_CHECKSUM_OFF;
+ /*
+ * Mark the state as changed locally, without a WAL record. Recovery
+ * then does not let a replayed checkpoint overwrite it, as no record
+ * could restore the change afterwards. The watermark is left alone:
+ * XLOG2_CHECKSUMS records at or below it stay covered, while records
+ * above it, which this node has not applied yet, still take effect
+ * on replay no matter when they were written.
+ */
+ ControlFile->data_checksum_is_local = true;
+
if (do_sync)
{
pg_log_info("syncing data directory");
diff --git a/src/bin/pg_controldata/pg_controldata.c b/src/bin/pg_controldata/pg_controldata.c
index 6fc87ed114d..c363ce3dbb7 100644
--- a/src/bin/pg_controldata/pg_controldata.c
+++ b/src/bin/pg_controldata/pg_controldata.c
@@ -349,6 +349,10 @@ main(int argc, char *argv[])
(ControlFile->float8ByVal ? _("by value") : _("by reference")));
printf(_("Data page checksum version: %u\n"),
ControlFile->data_checksum_version);
+ printf(_("Data checksum watermark: %X/%08X\n"),
+ LSN_FORMAT_ARGS(ControlFile->data_checksum_lsn));
+ printf(_("Data checksum state is node-local: %s\n"),
+ (ControlFile->data_checksum_is_local ? _("yes") : _("no")));
printf(_("Default char data signedness: %s\n"),
(ControlFile->default_char_signedness ? _("signed") : _("unsigned")));
printf(_("Mock authentication nonce: %s\n"),
diff --git a/src/bin/pg_resetwal/pg_resetwal.c b/src/bin/pg_resetwal/pg_resetwal.c
index 1542a56ca4b..cdfb9898066 100644
--- a/src/bin/pg_resetwal/pg_resetwal.c
+++ b/src/bin/pg_resetwal/pg_resetwal.c
@@ -923,6 +923,13 @@ RewriteControlFile(void)
ControlFile.backupEndPoint = InvalidXLogRecPtr;
ControlFile.backupEndRequired = false;
+ /*
+ * The old WAL is gone and the new position may lie below the old
+ * watermark, which would make replay ignore future checksum transition
+ * records. The state itself is kept.
+ */
+ ControlFile.data_checksum_lsn = InvalidXLogRecPtr;
+
/*
* Force the defaults for max_* settings. The values don't really matter
* as long as wal_level='minimal'; the postmaster will reset these fields
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 2e86fd158d0..466f4223501 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -37,7 +37,8 @@ static void usage(const char *progname);
static void perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo);
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec);
static void createBackupLabel(XLogRecPtr startpoint, TimeLineID starttli,
XLogRecPtr checkpointloc);
@@ -531,7 +532,8 @@ main(int argc, char **argv)
* This is the point of no return. Once we start copying things, there is
* no turning back!
*/
- perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo);
+ perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo,
+ divergerec);
if (showprogress)
pg_log_info("syncing target data directory");
@@ -566,7 +568,8 @@ static void
perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo)
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec)
{
XLogRecPtr endrec;
TimeLineID endtli;
@@ -738,6 +741,32 @@ perform_rewind(filemap_t *filemap, rewind_source *source,
ControlFile_new.minRecoveryPoint = endrec;
ControlFile_new.minRecoveryPointTLI = endtli;
ControlFile_new.state = DB_IN_ARCHIVE_RECOVERY;
+
+ /*
+ * Keep the target's own data checksum state. Most of the data directory
+ * is still the target's: only blocks it changed since the divergence
+ * were copied from the source, so the source's state says nothing about
+ * the pages that stay. Replay from the last common checkpoint applies
+ * any WAL-logged transition the target has not seen (the watermark tells
+ * them apart), which converges the rewound server to the source's state
+ * whenever the WAL carries it.
+ */
+ ControlFile_new.data_checksum_version = ControlFile_target.data_checksum_version;
+ ControlFile_new.data_checksum_lsn = ControlFile_target.data_checksum_lsn;
+ ControlFile_new.data_checksum_is_local = ControlFile_target.data_checksum_is_local;
+
+ /*
+ * The watermark is only meaningful within the history the node replays.
+ * Records at or below the divergence point are common to both histories
+ * and stay covered, but a watermark above it was set by a transition
+ * record on the target's own abandoned fork: numerically it can cover
+ * transition records the source wrote after the divergence, and replay
+ * would skip them as already applied. Clamp it to the divergence point,
+ * so that every transition record on the source's history takes effect.
+ */
+ if (ControlFile_new.data_checksum_lsn > divergerec)
+ ControlFile_new.data_checksum_lsn = divergerec;
+
if (!dry_run)
update_controlfile(datadir_target, &ControlFile_new, do_sync);
}
diff --git a/src/bin/pg_upgrade/controldata.c b/src/bin/pg_upgrade/controldata.c
index b3bd4ccde83..de539479bdc 100644
--- a/src/bin/pg_upgrade/controldata.c
+++ b/src/bin/pg_upgrade/controldata.c
@@ -431,7 +431,7 @@ get_control_data(ClusterInfo *cluster)
cluster->controldata.date_is_int = strstr(p, "64-bit integers") != NULL;
got_date_is_int = true;
}
- else if ((p = strstr(bufin, "checksum")) != NULL)
+ else if ((p = strstr(bufin, "Data page checksum version:")) != NULL)
{
p = strchr(p, ':');
diff --git a/src/include/catalog/pg_control.h b/src/include/catalog/pg_control.h
index 7b5404460ec..5365f4e8288 100644
--- a/src/include/catalog/pg_control.h
+++ b/src/include/catalog/pg_control.h
@@ -22,7 +22,7 @@
/* Version identifier for this pg_control format */
-#define PG_CONTROL_VERSION 1903
+#define PG_CONTROL_VERSION 1904
/* Nonce key length, see below */
#define MOCK_AUTH_NONCE_LEN 32
@@ -237,6 +237,33 @@ typedef struct ControlFileData
/* Current data checksums state */
uint32 data_checksum_version;
+ /*
+ * WAL position through which data checksum transitions are covered.
+ * Replay ignores XLOG2_CHECKSUMS records ending at or below this point:
+ * their effect is already contained in data_checksum_version, or an
+ * offline pg_checksums change made after they were first applied
+ * supersedes them. Ordinarily this is the end of the newest such record
+ * this node has written or applied, but a tool may store any position
+ * that covers the same set of records. If the node has never written or
+ * applied such a record this field shall be set to InvalidXLogRecPtr.
+ *
+ * The comparison has no timeline context, so the value is only valid
+ * within the WAL history this node replays. A tool that moves the node
+ * to another history must clamp the watermark to the point where the
+ * histories fork, as pg_rewind does, or reset it, as pg_resetwal does.
+ */
+ XLogRecPtr data_checksum_lsn;
+
+ /*
+ * True when data_checksum_version was last set by pg_checksums in an
+ * offline operation rather than by an online, WAL-logged transition.
+ * Such a state is local to this node and not derived from WAL, so
+ * recovery must not replace it with a state taken from a checkpoint
+ * record; nothing in the WAL could restore the change once it is
+ * overwritten. Cleared by the next WAL-logged transition.
+ */
+ bool data_checksum_is_local;
+
/*
* True if the default signedness of char is "signed" on a platform where
* the cluster is initialized.
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index d7eb648bd27..8d858be9927 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -89,6 +89,7 @@ PG_LWLOCK(54, WaitLSN)
PG_LWLOCK(55, LogicalDecodingControl)
PG_LWLOCK(56, DataChecksumsWorker)
PG_LWLOCK(57, AioWorkerControl)
+PG_LWLOCK(58, DataChecksumTransition)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/test/modules/test_checksums/Makefile b/src/test/modules/test_checksums/Makefile
index 71455cd5577..80f54bf6d8c 100644
--- a/src/test/modules/test_checksums/Makefile
+++ b/src/test/modules/test_checksums/Makefile
@@ -9,7 +9,7 @@
#
#-------------------------------------------------------------------------
-EXTRA_INSTALL = src/test/modules/injection_points
+EXTRA_INSTALL = contrib/pg_buffercache src/test/modules/injection_points
export enable_injection_points
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index fb7129d796f..e5d38fafb7d 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -35,6 +35,16 @@ tests += {
't/009_fpi.pl',
't/010_backup_straddle.pl',
't/011_standby_straddle.pl',
+ 't/012_offline_standby.pl',
+ 't/013_rewind.pl',
+ 't/014_lockstep.pl',
+ 't/015_standby_crash_after_disable.pl',
+ 't/016_promote_enable_crash.pl',
+ 't/017_restartpoint_race.pl',
+ 't/018_enable_crash_windows.pl',
+ 't/019_standby_shutdown_catchup.pl',
+ 't/020_cascade_divergence.pl',
+ 't/021_rewind_divergent_transitions.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
new file mode 100644
index 00000000000..241a1282ae9
--- /dev/null
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -0,0 +1,289 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Offline checksum changes with pg_checksums are local to one node. A
+# standby must neither adopt the state of the primary from replayed
+# checkpoint records, nor lose its own offline change to them.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', 'autovacuum = off');
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Scenario 1: enable offline on the primary only. The standby must
+# stay off, warn about the mismatch, and remain readable.
+$standby->stop;
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+test_checksum_state($primary, 'on');
+test_checksum_state($standby, 'off');
+
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after offline enable on the primary');
+
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+my @warnings = $log =~ /(does not match the state)/g;
+is(scalar(@warnings), 1, 'mismatch warned once per remote value');
+
+# Matching states re-arm the warning: undo the divergence on the primary,
+# then diverge again to the same value, all without restarting the standby.
+$primary->stop;
+$primary->checksum_disable_offline;
+$logstart = -s $standby->logfile;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no warning while the states match again');
+like($log, qr/now agrees with the replayed WAL/,
+ 'convergence is reported once the states match again');
+
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+test_checksum_state($standby, 'off');
+
+# The local state survives both clean and immediate restarts.
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+# Still in scenario 1's divergence: take a base backup *from the
+# standby*. A backup taken on a standby uses the last restartpoint as
+# its starting checkpoint (do_pg_backup_start()), so the record at the
+# redo point was written by the upstream primary and carries the
+# primary's state. Recovery from such a backup must not adopt that
+# state: the files were copied from the standby, and the new node
+# would come up verifying checksums its files do not have.
+
+# Make sure the standby has written pages under its own "off" state, so
+# its files really do lack checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,50000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', "SELECT count(*) FROM t2;");
+
+$standby->backup('from_standby');
+my $newnode = PostgreSQL::Test::Cluster->new('newnode');
+$newnode->init_from_backup($standby, 'from_standby');
+
+# Stream from the primary, not from the standby the files were copied
+# from, so the new node replays the "on" primary's records.
+$newnode->enable_streaming($primary);
+$newnode->start;
+
+# The new node's files all came from a cluster running with checksums
+# off. Anything but "off" here means it adopted the primary's state
+# through the checkpoint record at the redo point.
+my ($rc, $stdout, $stderr) = $newnode->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'the node copied from the standby accepts connections')
+ or diag("stderr: $stderr");
+is($stdout, 'off', 'backup of an "off" standby comes up with checksums off');
+
+# A checkpoint record streamed from the "on" primary must not flip the
+# state either.
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($newnode);
+test_checksum_state($newnode, 'off');
+
+# And it must be able to read the pages the standby wrote without
+# checksums.
+(undef, $stdout, $stderr) =
+ $newnode->psql('postgres', "SELECT count(*) FROM t2;");
+is($stdout, '50000', 'pages copied from the standby are readable')
+ or diag("stderr: $stderr");
+
+my $newnode_log = PostgreSQL::Test::Utils::slurp_file($newnode->logfile);
+unlike(
+ $newnode_log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the node copied from the standby');
+
+$newnode->stop('immediate');
+
+# Converge the cluster: enable offline on the standby too.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+$primary->wait_for_catchup($standby);
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after converging');
+
+# Scenario 2: disable offline on the standby only. The replayed
+# checkpoint records of the still-enabled primary must not override it.
+$standby->stop;
+$standby->checksum_disable_offline;
+$standby->start;
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'off');
+
+# Restartpoints must persist the local state, not the replayed copy.
+$standby->safe_psql('postgres', "CHECKPOINT;");
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable with checksums disabled locally');
+
+# Scenario 3: crash-restart right after an online transition, before the
+# next restartpoint. Replay then resumes from an older restartpoint whose
+# checkpoint records still carry the pre-transition state. Those must
+# still match the state seeded from the control file, and the transition
+# itself must be re-established by re-replaying the XLOG2_CHECKSUMS record,
+# without a spurious mismatch warning along the way.
+
+# Converge first: bring the standby back to "on" offline.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+disable_data_checksums($primary, wait => 'off');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+# Crash-restart the standby immediately, before any restartpoint has had a
+# chance to persist the new state to its control file.
+$logstart = -s $standby->logfile;
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001',
+ 'standby readable after crash-restart across an online transition');
+
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no spurious mismatch warning after crash-restart across an online transition'
+);
+
+# Scenario 4: a standby stopped while replaying an interrupted online
+# transition keeps the interrupted state in its own control file. A
+# primary is never caught this way, as its checksums launcher resolves
+# inprogress-on back to off from its exit cleanup. A standby has no
+# launcher; it carries forward whatever the last replayed record left
+# it in.
+
+# Block an online enable on the primary at inprogress-on with a
+# blocking temp table, same trick as in 004_offline.pl.
+my $bsession = $primary->background_psql('postgres');
+$bsession->query_safe('CREATE TEMPORARY TABLE tt (a integer);');
+enable_data_checksums($primary, wait => 'inprogress-on');
+
+# The standby picks up the in-progress state from the XLOG2_CHECKSUMS record.
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# Stop the standby cleanly; its restartpoint persists inprogress-on to
+# its own control file, since nothing on a standby resolves it away.
+$standby->stop;
+$standby->start;
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# The primary still sits at inprogress-on, so a base backup taken now
+# copies a control file and a redo point that both carry the
+# in-progress state; replay of the XLOG2_CHECKSUMS records completes
+# the transition on the new standby.
+$primary->backup('inprogress_backup');
+my $standby2 = PostgreSQL::Test::Cluster->new('standby2');
+$standby2->init_from_backup($primary, 'inprogress_backup',
+ has_streaming => 1);
+$standby2->start;
+
+# Backup label recovery adopts the redo point's state immediately,
+# before any further WAL is replayed.
+wait_for_checksum_state($standby2, 'inprogress-on');
+
+$bsession->quit;
+wait_for_checksum_state($primary, 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable once the transition completes');
+
+# Scenario 4 continued: the backup taken mid-transition must also
+# complete the transition and stay readable.
+$primary->wait_for_catchup($standby2);
+wait_for_checksum_state($standby2, 'on');
+
+is($standby2->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'mid-transition backup standby readable after the transition');
+
+$standby2->stop;
+
+# Nothing along this standby's replay can legitimately disagree: it
+# started out at the same inprogress-on state as the primary.
+my $standby2_log = PostgreSQL::Test::Utils::slurp_file($standby2->logfile);
+unlike(
+ $standby2_log,
+ qr/does not match the state/,
+ 'no mismatch warning on a backup taken mid-transition');
+
+# No extra CHECKPOINT needed: the transition's completion checkpoint
+# wrote every page with a checksum, and the shutdown restartpoint
+# flushes the rest.
+command_ok([ 'pg_checksums', '--check', '-D', $standby2->data_dir ],
+ 'checksums valid on the mid-transition backup standby');
+
+$standby->stop;
+$primary->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/013_rewind.pl b/src/test/modules/test_checksums/t/013_rewind.pl
new file mode 100644
index 00000000000..a791e24317d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/013_rewind.pl
@@ -0,0 +1,201 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test pg_rewind across an online data checksum enable.
+#
+# A clean switchover leaves the shutdown checkpoint of the old primary
+# as the last common checkpoint between the two nodes. When data
+# checksums are enabled online on the new primary before the old one is
+# rewound, pg_rewind installs the control file of the new primary, which
+# already claims checksums are fully enabled, while replay begins at the
+# shutdown checkpoint whose record still carries the old state. Replay
+# of the WAL stretch from before the enable must run with checksums off,
+# as recorded in the checkpoint, else it would verify pages which never
+# had checksums written and fail recovery.
+#
+# The new primary requests a checkpoint right after promotion, and
+# replaying its CHECKPOINT_REDO record would repair the state before
+# any interesting WAL is reached. Hold the checkpointer on the standby
+# in a restartpoint over the promotion, like in the recovery test
+# 041_checkpoint_at_promote, so that the post-promotion writes end up
+# in WAL before the first checkpoint of the new timeline.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Old primary. full_page_writes is off so that the updates done on the
+# promoted node do not carry full page images, forcing replay on the
+# rewound node to read the pages from disk. wal_log_hints is required
+# by pg_rewind on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+full_page_writes = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+log_checkpoints = on
+]);
+$node_a->start;
+
+if (!$node_a->check_extension('injection_points'))
+{
+ plan skip_all => 'Extension injection_points not installed';
+}
+
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+$node_a->safe_psql('postgres', "CREATE EXTENSION injection_points;");
+
+# Set the hint bits on t before taking the backup, so that reads on the
+# promoted node do not emit full page images for its pages later.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->safe_psql('postgres', "SELECT count(*) FROM t;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_a, 'off');
+test_checksum_state($node_b, 'off');
+
+# Hold the next restartpoint on the standby.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point', 'wait');");
+
+# Give the restartpoint a checkpoint record to work on, then start it
+# in a background session; it will block on the injection point with
+# the checkpointer busy until released.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+my $bg_psql = $node_b->background_psql('postgres', on_error_stop => 0);
+$bg_psql->query_until(
+ qr/starting_restartpoint/, q(
+ \echo starting_restartpoint
+ CHECKPOINT;
+));
+$node_b->wait_for_event('checkpointer', 'create-restart-point');
+
+# Clean switchover: the shutdown checkpoint of A streams to B and
+# becomes the last common checkpoint.
+$node_a->stop('fast');
+
+my ($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+$stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+my $shutdown_ckpt = $1;
+$node_b->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_ckpt'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+
+my $logstart = -s $node_b->logfile;
+$node_b->promote;
+
+# Accidental restart of the old primary, diverging its timeline.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->stop('fast');
+
+# Updates on the new primary before its first checkpoint; replay of
+# these on the rewound node has to read the pages from disk.
+$node_b->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 25 = 0;");
+ok( !$node_b->log_contains("checkpoint complete", $logstart),
+ "no checkpoint on the new timeline before the updates");
+
+# Release the checkpointer; the queued post-promotion checkpoint runs
+# after the updates.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+$bg_psql->quit;
+
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+# pg_rewind refuses to run with full_page_writes disabled on the
+# source; the updates it was disabled for are already in WAL.
+$node_b->safe_psql('postgres', "ALTER SYSTEM SET full_page_writes = on;");
+$node_b->safe_psql('postgres', "SELECT pg_reload_conf();");
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the new primary');
+
+# Replay on the rewound node must start at the shutdown checkpoint of
+# the switchover, with the control file of the new primary.
+my $backup_label = slurp_file($node_a->data_dir . '/backup_label');
+$backup_label =~ /^CHECKPOINT LOCATION: ([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from backup_label";
+is($1, $shutdown_ckpt, 'replay starts at the switchover checkpoint');
+
+($stdout, $stderr) = run_command(
+ [
+ 'pg_waldump',
+ '-p' => $node_a->data_dir . '/pg_wal',
+ '-t' => 1,
+ '-s' => $shutdown_ckpt,
+ '-n' => 1,
+ ]);
+like($stdout, qr/CHECKPOINT_SHUTDOWN/,
+ 'last common checkpoint is a shutdown checkpoint');
+
+# pg_rewind keeps the target's own checksum state in the control file it
+# writes; the online enable reaches the rewound node through WAL replay
+# below, not through the copied control file.
+($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+like(
+ $stdout,
+ qr/^Data page checksum version:\s*0$/m,
+ 'rewound node keeps its own checksum state in the control file');
+
+# Start the rewound node as a standby of the new primary. Replay runs
+# through the pre-enable WAL stretch and the online enable.
+#
+# The rewind replaced the configuration files with those of the new
+# primary, so put the port back.
+my $connstr = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_b->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/014_lockstep.pl b/src/test/modules/test_checksums/t/014_lockstep.pl
new file mode 100644
index 00000000000..713f37fb22c
--- /dev/null
+++ b/src/test/modules/test_checksums/t/014_lockstep.pl
@@ -0,0 +1,182 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# The lockstep procedure for offline checksum changes in a replication
+# setup: stop all nodes, run pg_checksums on all of them, restart.
+# Replay of checkpoint records written before the change must not
+# revert the state of the standby.
+#
+# The same pair then runs the procedure in the disable direction, and
+# finally across a divergence: after a failover both nodes get their
+# checksums enabled offline, while the last common checkpoint still
+# carries "off". pg_rewind must keep the offline "on" on the rewound
+# node, since no record in the replayed WAL could ever restore it.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# wal_log_hints keeps the pair eligible for pg_rewind without checksums.
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Part 1: the lockstep procedure in the enable direction. Stop the
+# standby first: the WAL written after this point is replayed only
+# after the offline switch, and every checkpoint record in it still
+# carries the old state.
+$standby->stop;
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 0;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 1;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+# The lockstep procedure.
+$primary->checksum_enable_offline;
+$standby->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to off.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "off" in the replayed WAL/,
+ 0);
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep enable');
+
+# Crash the standby and replay the same stretch again.
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'on');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after crash restart');
+
+# Once a post-switch checkpoint has been replayed the states match and
+# no warning may be logged.
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no mismatch warning once the states match');
+
+# Every page the standby wrote in this window must carry a checksum.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$standby->stop;
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'checksums valid on the standby');
+
+# Part 2: the lockstep procedure in the other direction. The standby
+# is already stopped; give it a pre-switch WAL stretch to replay, with
+# every checkpoint record in it still carrying "on" -- an explicit
+# checkpoint plus the shutdown checkpoint written by stop().
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 2;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+$primary->checksum_disable_offline;
+$standby->checksum_disable_offline;
+
+my $disable_logstart = -s $standby->logfile;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to on.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $disable_logstart);
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep disable');
+
+# Part 3: failover, divergence and pg_rewind across offline enables.
+# The roles swap from here on: the standby becomes the new primary and
+# the old primary is rewound to follow it, but the variables keep the
+# names they had above.
+#
+# Failover: promote the standby, then diverge the old primary.
+$standby->promote;
+$standby->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$primary->stop;
+
+# The lockstep procedure across the divergence: both nodes get their
+# checksums enabled offline.
+$primary->checksum_enable_offline;
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+
+# The states match, so the rewind proceeds without complaint.
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $primary->data_dir,
+ '--source-server' => $standby->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled offline on both nodes');
+
+# The rewind replaced the configuration files with those of the source,
+# so put the port back. The copied file also carries the source's own
+# stale primary_conninfo; enable_streaming() below appends ours, which
+# wins as the later entry.
+$primary->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$primary->port]}
+]);
+$primary->enable_streaming($standby);
+
+my $rewind_logstart = -s $primary->logfile;
+$primary->start;
+$standby->wait_for_catchup($primary);
+
+is($primary->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+
+# The offline enable survives: replay from the common checkpoint must
+# not resurrect the pre-divergence "off".
+test_checksum_state($primary, 'on');
+
+$log =
+ PostgreSQL::Test::Utils::slurp_file($primary->logfile, $rewind_logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no divergence reported between the rewound server and its source');
+
+$primary->stop;
+$standby->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
new file mode 100644
index 00000000000..315b27b93d4
--- /dev/null
+++ b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
@@ -0,0 +1,125 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that a standby crashing after a replayed online disable, but before
+# its next restartpoint, restarts cleanly.
+#
+# Pages dirtied before the XLOG2_CHECKSUMS record and evicted after it are
+# written out under the new "off" state, without checksums. The control
+# file must follow the record immediately in this direction: a crash-restart
+# initializes checksum verification from the control file, and one still
+# saying "on" would fail verification on exactly those pages while replaying
+# records older than the state change.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$primary->start;
+
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,1000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+
+# Small buffer pool so that replaying a bulk load evicts the pages we care
+# about, and no restartpoints so the control file stays where it is.
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 1MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_delay = 10000
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'on');
+
+# No full page images, so that replay has to read the pages of "t" back from
+# disk instead of overwriting them from the WAL.
+$primary->append_conf('postgresql.conf', 'full_page_writes = off');
+$primary->reload;
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" on the standby, *before* the state change. They are
+# not flushed: the standby has no restartpoint from here on.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+# Online disable. The standby replays it and switches to "off", but its
+# control file is not updated.
+$primary->safe_psql('postgres', 'SELECT pg_disable_data_checksums();');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+my ($ctl_before) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($ctl_state) = $ctl_before =~ /Data page checksum version:\s+(\d+)/;
+note(
+ "standby control file data_checksum_version after the disable: "
+ . "$ctl_state (0 = off, 1 = on)");
+
+# Force the standby to evict the dirty pages of "t" now that it is "off":
+# they are written back without checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE filler AS SELECT generate_series(1,300000) AS a;");
+$primary->wait_for_catchup($standby);
+
+# Crash the standby before it gets a chance to run a restartpoint.
+$standby->stop('immediate');
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'standby restarts after crashing between a checksum state '
+ . 'change and the next restartpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'the pages written while "off" are readable')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("standby log tail:\n" . join("\n", @lines[ -12 .. -1 ]))
+ if @lines >= 12;
+ fail('the pages written while "off" are readable');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/016_promote_enable_crash.pl b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
new file mode 100644
index 00000000000..97f84da5685
--- /dev/null
+++ b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a standby promoted right after replaying an online enable, before any
+# restartpoint has flushed the rewritten pages.
+#
+# The end-of-recovery record persists the checksum state without flushing the
+# buffer pool, while the control file still points at a restartpoint older than
+# the transition. Persisting "on" there would make a crash before the
+# post-promotion checkpoint verify checksums over the pages that were flushed
+# while the state was still "off"; promotion must take the full checkpoint
+# instead.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Dirty the pages of "t" without full page images, then push them out to disk
+# on the standby while checksums are still off.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Online enable on the primary; the standby replays it into its own buffers,
+# where the rewritten pages stay dirty (no restartpoint, no bgwriter).
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+note("standby control file before the promotion: $before");
+
+# Promote. The end-of-recovery record persists the live state without any
+# buffer flush; the checkpoint requested afterwards is not immediate.
+$standby->promote;
+$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery();')
+ or die 'timed out waiting for the promotion';
+$standby->stop('immediate');
+
+($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($ckpt) = $ctl =~ /Latest checkpoint location:\s+(\S+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note(
+ "standby control file after the promotion: $after, checkpoint $ckpt, redo $redo"
+);
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'promoted node restarts after crashing before its first checkpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/017_restartpoint_race.pl b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
new file mode 100644
index 00000000000..65aa834f366
--- /dev/null
+++ b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
@@ -0,0 +1,138 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a restartpoint whose flush races the replay of an online enable.
+#
+# Replay keeps running while CheckPointGuts() writes out the buffer pool, so
+# the state at the end of the flush can be newer than the one the flush ran
+# under. Persisting it would claim checksums for pages the flush wrote out
+# before the transition, against a redo pointer that predates it.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# C1: the checkpoint the pending restartpoint will target.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" without full page images and push them to disk on
+# the standby while it is still "off".
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Start a restartpoint for C1 and hold it right after CheckPointGuts().
+$standby->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point','wait');");
+my $bg = $standby->background_psql('postgres');
+$bg->query_until(qr//, "\\echo restartpoint\nCHECKPOINT;\n");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint injection point';
+
+# The startup process keeps replaying while the checkpointer is held: the
+# whole online enable lands, and the rewritten pages stay dirty.
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+# Let the restartpoint finish; it persists the state it samples now.
+$standby->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint to finish';
+$standby->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+
+$standby->stop('immediate');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note("standby control file after the restartpoint: $after, redo $redo");
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started, 'standby restarts after a restartpoint that raced the enable');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/018_enable_crash_windows.pl b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
new file mode 100644
index 00000000000..17f96ff6272
--- /dev/null
+++ b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
@@ -0,0 +1,496 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Crash and checkpoint windows inside an online data checksum enable on
+# a single node. Each scenario starts from "off" on the same cluster,
+# runs the enable into a held injection point, and checks what a crash
+# or a concurrent checkpoint in that window may and may not do to the
+# control file and the pages on disk.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+use IPC::Run;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Bring the node back to "off" with no leftover table between
+# scenarios. The crash restarts already resolve an interrupted enable
+# to "off", so only disable when something is left to disable. Those
+# restarts also clear any attached injection points. The checkpoint
+# bounds the WAL the next scenario's crash recovery has to replay.
+sub reset_scenario
+{
+ my ($node) = @_;
+
+ BAIL_OUT('node is not running after the previous scenario')
+ unless $node->is_alive;
+
+ my $state = $node->safe_psql('postgres', 'SHOW data_checksums;');
+ disable_data_checksums($node, wait => 'off') if $state ne 'off';
+ $node->safe_psql('postgres', 'DROP TABLE IF EXISTS t;');
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+ test_checksum_state($node, 'off');
+}
+
+my $node = PostgreSQL::Test::Cluster->new('crash_windows');
+$node->init(no_data_checksums => 1);
+
+# Scenario 1 needs the pages of "t" to reach disk without a full page
+# image, hence full_page_writes and wal_log_hints off. Scenario 2
+# needs them to stay dirty in shared buffers until a checkpoint, hence
+# no bgwriter and a large shared_buffers. wal_level is pinned because
+# init() writes "minimal" for a non-streaming node. The rest merely
+# tolerate the settings.
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+wal_level = replica
+max_wal_size = 10GB
+full_page_writes = off
+wal_log_hints = off
+shared_buffers = 512MB
+bgwriter_lru_maxpages = 0
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$node->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+
+test_checksum_state($node, 'off');
+
+# Scenario 1: a primary crashes inside SetDataChecksumsOn(), just before the
+# forced checkpoint that flushes the rewritten pages, with crash recovery
+# resuming from a checkpoint older than the transition. The control file
+# must still say "inprogress-on" there: replay does not adopt the state of
+# the checkpoint record it resumes from, so an "on" written before the flush
+# would stay in effect while replay reads pages whose rewrite never reached
+# disk. Here the pages went out through the rewriting worker's ring buffer,
+# so only the control file state discriminates; see the next scenario for
+# the case where they did not.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Establish the checkpoint that crash recovery will resume from.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty the pages of "t" without full page images, then push them out
+ # to disk while checksums are still off.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+ my $evicted = $node->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+ note("evict_relation on t: $evicted, relpath $relpath");
+
+ # Stop the enable right after the control file has been updated to
+ # "on" but before the checkpoint that flushes the rewritten pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+ $node->safe_psql('postgres', 'SELECT pg_enable_data_checksums();');
+
+ $node->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'datachecksums-on-before-checkpoint';"
+ ) or die 'timed out waiting for the injection point';
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ $node->stop('immediate');
+
+ # Show the on-disk checksum field of the first page of "t".
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 2: the same crash as above, but with the rewritten pages resident
+# in shared buffers instead of gone out through the ring. The rewriting
+# worker reads through a BAS_VACUUM ring, which writes pages back as the
+# ring recycles, but a page already resident in shared buffers is not read
+# through the ring: ReadBufferExtended() hands back the existing buffer, and
+# it stays dirty until a checkpoint. The control file may therefore not
+# say "on" before that checkpoint has run, or crash recovery would resume
+# from a checkpoint older than the transition with verification already
+# enabled, and replay records that read those still-unchecksummed pages.
+# full_page_writes is off so that the records do not simply overwrite the
+# pages with a full page image.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,100000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Establish the checkpoint that crash recovery will resume from, with the
+ # pages of "t" written out while checksums are still off.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty those pages again without emitting full page images, and leave
+ # them in shared buffers. Replay of these records has to read the
+ # pages from disk.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+
+ my $dirty_before = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_before, '>', 0,
+ 'pages of t are resident and dirty before enabling checksums');
+
+ # Hold the enabling right before the checkpoint that flushes the rewritten
+ # pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ # The rewritten pages must still be sitting dirty in shared buffers, or
+ # the window this test is about does not exist.
+ my $dirty_after = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_after, '>', 0,
+ 'rewritten pages of t are still dirty before the checkpoint');
+
+ $node->stop('immediate');
+
+ # The on-disk copy is the one written before the transition, without a
+ # checksum.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ is($pd_checksum, 0, 'block 0 of t on disk carries no checksum');
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 3: a checkpoint that runs between the XLOG2_CHECKSUMS("on")
+# record and the control file write at the end of SetDataChecksumsOn() must
+# not make a crash throw the completed transition away. SetDataChecksumsOn()
+# writes the record, flips shared memory to "on", emits the barrier and only
+# then requests the checkpoint that flushes the rewritten pages; the control
+# file is written after that checkpoint returns. A crash in that window is
+# harmless only as long as recovery still starts before the record. Any
+# checkpoint completing in the window moves the redo point past the record,
+# so recovery would never see it, would come up with the control file's
+# "inprogress-on" and StartupXLOG() would demote that to "off", even though
+# every page on disk carries a checksum by then. CreateCheckPoint()
+# therefore persists the state the checkpoint ran under, the same way
+# CreateRestartPoint() does. Here the window is made deterministic by
+# holding the launcher at the datachecksums-on-before-checkpoint injection
+# point and checkpointing from another session.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Hold the launcher after the record, the shared memory flip and the
+ # barrier, but before the checkpoint SetDataChecksumsOn() requests
+ # itself.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ # Every backend already sees "on" and writes checksums.
+ test_checksum_state($node, 'on');
+
+ # ... while the control file still says "inprogress-on".
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($before, '3', 'control file says "inprogress-on" inside the window');
+
+ # A concurrent checkpoint. It flushes the rewritten pages and moves
+ # the redo point past the XLOG2_CHECKSUMS("on") record, so it has to
+ # record the "on" state in the control file as well.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($after, '1',
+ 'the concurrent checkpoint records "on" in the control file');
+
+ $node->stop('immediate');
+
+ # The checkpoint flushed the rewritten pages, so they carry a checksum on
+ # disk: the transition really did complete.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ isnt($pd_checksum, 0, 'block 0 of t on disk carries a checksum');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition is complete on disk, so the cluster has to come back
+ # "on".
+ test_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+}
+reset_scenario($node);
+
+# Scenario 4: a crash between the checkpoint an online enable requests and
+# the control file write that follows it must not lose the transition. The
+# last steps of SetDataChecksumsOn() are
+#
+# WAL record -> shmem -> barrier -> checkpoint -> persist
+#
+# The checkpoint is what makes "on" safe to persist, but it also moves the
+# redo point above the XLOG2_CHECKSUMS record that carries the new state.
+# Crash recovery started from that checkpoint therefore never replays the
+# record, so without the state the checkpoint itself persists, a crash in
+# the remaining window would bring the cluster back at "inprogress-on",
+# which StartupXLOG() resolves to "off", discarding a transition whose
+# pages are all on disk with a checksum. The previous scenario exercises
+# the same window through a checkpoint requested by another session; this
+# one closes it from the other side: the transition's own checkpoint is the
+# one that moves the redo point, so persisting the state from
+# CreateCheckPoint() is what has to save it, not the write below the
+# injection point.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Hold the launcher after the checkpoint that licenses "on" has
+ # completed and before the state reaches the control file.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-after-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-after-checkpoint');
+
+ # The transition is complete as far as the running cluster is concerned.
+ test_checksum_state($node, 'on');
+
+ # The checkpoint has already recorded it, so the pending write below the
+ # injection point has nothing left to do.
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($version) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($version, '1',
+ 'the requested checkpoint recorded "on" in the control file');
+
+ # Crash before the control file write that follows the checkpoint.
+ $node->stop('immediate');
+
+ $node->start;
+
+ # Every page on disk carries a checksum and the checkpoint that
+ # flushed them completed, so the cluster has to come back verifying
+ # them.
+ test_checksum_state($node, 'on');
+
+ is($node->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '10000', 'relation readable after the crash');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ # The data directory must be one the offline tools accept.
+ $node->stop;
+ $node->command_ok([ 'pg_checksums', '--check', '-D', $node->data_dir ],
+ 'pg_checksums accepts the data directory');
+
+ # Leave the node running for the reset.
+ $node->start;
+}
+reset_scenario($node);
+
+# Scenario 5: a checkpoint racing SetDataChecksumsOn() between the insertion
+# of the XLOG2_CHECKSUMS("on") record and the shared memory update must not
+# insert an XLOG_CHECKPOINT_REDO record that follows the transition in WAL
+# order while still carrying "inprogress-on". Recovery resuming from such a
+# redo point never replays the preceding transition record, comes up in
+# "inprogress-on" and resolves the finished transition as interrupted.
+# XLogChecksums() closes the window by inserting the record and publishing
+# the new state under DataChecksumTransitionLock, which CreateCheckPoint()
+# takes around sampling the state and inserting the redo record. Here the
+# launcher is held between the two steps at the
+# datachecksums-on-before-publish injection point, and a concurrent
+# CHECKPOINT has to block on the lock instead of completing inside the
+# window.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # The datachecksums-on-before-publish point fires inside a critical
+ # section, where the wait machinery must not allocate. Waiting once
+ # at the datachecksums-enable-checksums-delay point, which the
+ # launcher runs outside the critical section, initializes it; see
+ # 050_redo_segment_missing.pl for the same recipe around
+ # create-checkpoint-run.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-enable-checksums-delay','wait');"
+ );
+
+ # Hold the launcher after the XLOG2_CHECKSUMS("on") record is in WAL but
+ # before the new state is published in shared memory.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-publish','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-enable-checksums-delay');
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-enable-checksums-delay');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-enable-checksums-delay');");
+
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-publish');
+
+ # The record is in WAL, the published state is still the old one.
+ test_checksum_state($node, 'inprogress-on');
+
+ # A concurrent checkpoint. It must block on DataChecksumTransitionLock
+ # before inserting its redo record rather than complete inside the window.
+ my $checkpointer = IPC::Run::start(
+ [ 'psql', '-XAtq', '-d', $node->connstr('postgres'), '-c', 'CHECKPOINT;' ],
+ '<' => '/dev/null',
+ '>' => '/dev/null',
+ '2>' => '/dev/null',
+ IPC::Run::timer($PostgreSQL::Test::Utils::timeout_default));
+
+ ok( $node->poll_query_until(
+ 'postgres',
+ "SELECT wait_event = 'DataChecksumTransition' "
+ . "FROM pg_stat_activity WHERE backend_type = 'checkpointer';"),
+ 'concurrent checkpoint blocks on DataChecksumTransitionLock');
+
+ # Release the launcher; the checkpoint then samples the published "on" and
+ # its redo record follows the transition record.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-publish');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-publish');");
+
+ $checkpointer->finish;
+
+ # Crash while the launcher's own checkpoint may still be in flight.
+ # Recovery resumes from the concurrent checkpoint's redo point, which
+ # now lies above the transition record and carries "on".
+ $node->stop('immediate');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition completed, so the cluster has to come back "on".
+ wait_for_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ $node->stop;
+}
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
new file mode 100644
index 00000000000..aaa35f8f7dd
--- /dev/null
+++ b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
@@ -0,0 +1,126 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A standby stopped after replaying an online enable, but before the
+# checkpoint record that follows it on the primary.
+#
+# The shutdown restartpoint has no new checkpoint record to work from
+# and is skipped, so the control file would keep the in-progress state
+# the transition passed through, even though replay left the node at
+# "on" and rewrote every page. pg_checksums reads that field and
+# refuses to run on an in-progress state, so the shutdown has to flush
+# and catch it up instead.
+#
+# The caught-up control file then carries the watermark of the "on"
+# record, which the second half of this test relies on: an offline
+# disable made while the standby is down must survive the re-replay of
+# that record on the next startup, or the change would be silently
+# reverted.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+wal_keep_size = 1GB
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Anchor the standby's last restartpoint here, so that nothing replayed
+# from now on gives the shutdown restartpoint a newer checkpoint record
+# to use.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Hold the enable right after the state change record, before the
+# checkpoint it requests once the transition is complete.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+enable_data_checksums($primary);
+$primary->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+wait_for_checksum_state($primary, 'on');
+
+# The state change record is already flushed and streams on its own;
+# no checkpoint record follows while the launcher is held.
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+$standby->stop('fast');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($state, '1',
+ 'shutdown catches the control file up with the replayed state');
+
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'pg_checksums verifies the stopped standby');
+
+# Disable checksums offline while the standby is down. The restart
+# below resumes replay under the last restartpoint, re-reading the "on"
+# record; the watermark persisted with the catch-up above marks it as
+# already applied, so the offline change must survive.
+$standby->checksum_disable_offline;
+
+# Let the enable finish on the primary.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-checkpoint');");
+$primary->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-checkpoint');");
+
+# The launcher exits only after the checkpoint it requests completes,
+# so a checkpoint record now follows the transition in WAL.
+$primary->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_catalog.pg_stat_activity "
+ . "WHERE backend_type = 'datachecksums launcher';");
+
+my $log_offset = -s $standby->logfile;
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+
+# The nodes legitimately diverged, which replayed checkpoints report.
+$standby->wait_for_log(
+ qr/data checksum state "off" of this node does not match the state "on" in the replayed WAL/,
+ $log_offset);
+
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/020_cascade_divergence.pl b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
new file mode 100644
index 00000000000..b4b474fa583
--- /dev/null
+++ b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
@@ -0,0 +1,115 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# An offline data checksum state change made on one node of a cascading setup
+# is reported all the way down the chain.
+#
+# CheckReplayedDataChecksumState() is reached from xlog_redo() when a
+# checkpoint record (XLOG_CHECKPOINT_SHUTDOWN or XLOG_CHECKPOINT_REDO) is
+# replayed after consistency has been reached. Those records are written by
+# the root primary only and are relayed verbatim by every intermediate
+# standby, so a cascaded standby that was changed offline learns about the
+# divergence too. An intermediate standby's own restartpoints are not
+# WAL-logged and cannot surface it.
+#
+# Note that this makes detection checkpoint-driven: an idle or read-only
+# primary surfaces the divergence no sooner than its next checkpoint, so up to
+# checkpoint_timeout may pass before the operator sees anything. Closing that
+# window would require comparing the state when a walreceiver connects.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# checkpoint_timeout is deliberately long so that the only checkpoint record in
+# this test is the one requested explicitly below.
+my $conf = qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+);
+
+my $primary = PostgreSQL::Test::Cluster->new('cascade_primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', $conf);
+$primary->start;
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,1000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('cascade_standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf('postgresql.conf', $conf);
+$standby->start;
+
+$standby->backup('backup');
+my $cascade = PostgreSQL::Test::Cluster->new('cascade_cascade');
+$cascade->init_from_backup($standby, 'backup', has_streaming => 1);
+$cascade->append_conf('postgresql.conf', $conf);
+$cascade->start;
+
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+test_checksum_state($cascade, 'off');
+
+# Change the cascaded standby offline, and only it.
+$cascade->stop;
+$cascade->command_ok([ 'pg_checksums', '--enable', '-D', $cascade->data_dir ],
+ 'pg_checksums enables checksums on the cascaded standby only');
+
+my $logstart = -s $cascade->logfile;
+$cascade->start;
+
+test_checksum_state($cascade, 'on');
+
+# Ordinary WAL traffic carries no state, so the cascaded standby keeps
+# streaming from a chain whose state is "off" without noticing.
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (1);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+is($cascade->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '1001', 'cascaded standby keeps streaming from a divergent chain');
+
+# Restartpoints on the intermediate standby are not WAL-logged, so this does
+# not surface anything on the cascaded standby either.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (2);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+my $log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+unlike(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'an intermediate restartpoint does not surface the divergence');
+
+# A checkpoint on the root primary does: the record is relayed down the whole
+# chain, so both the direct and the cascaded standby compare it against their
+# own state. wait_for_catchup() below waits for replay, not just receipt, so
+# the record has been through xlog_redo() on both nodes once it returns.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+$log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+like(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'the cascaded standby reports the divergence on the primary checkpoint');
+
+$cascade->stop;
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
new file mode 100644
index 00000000000..a6c8a1091ee
--- /dev/null
+++ b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
@@ -0,0 +1,193 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind across online data checksum transitions on both sides of a
+# divergence, with the two nodes trading roles between the scenarios.
+#
+# Scenario 1: after a switchover the new primary enables checksums
+# online, while the old primary restarts on its old timeline, advances
+# its WAL beyond the enable records and runs an online enable/disable
+# cycle of its own. The old primary ends "off" with a checksum
+# watermark numerically above every checksum record the new primary has
+# written. pg_rewind clamps the watermark it keeps to the divergence
+# point; without the clamp, replay on the rewound node would skip the
+# source's enable as already applied and stay "off" under an "on"
+# primary.
+#
+# Scenario 2: checksums are disabled again, and after another
+# switchover both nodes enable them online independently, so the
+# divergence checkpoint carries "off" while both control files say
+# "on". The rewind is allowed, the target keeps its own "on" state,
+# and replay re-walks the source's enable onto the already enabled
+# node.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+sub controldata_watermark
+{
+ my ($node) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $node->data_dir ]);
+ $stdout =~ /^Data checksum watermark:\s*([0-9A-F]+)\/([0-9A-F]+)$/m
+ or die "watermark missing from pg_controldata output";
+ return (hex($1) << 32) + hex($2);
+}
+
+# Wait until the standby has replayed the shutdown checkpoint of the
+# stopped primary, so that a subsequent promotion diverges after it and
+# the shutdown checkpoint becomes the last common checkpoint.
+sub wait_for_shutdown_checkpoint_replay
+{
+ my ($primary, $standby) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $primary->data_dir ]);
+ $stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+ my $shutdown_checkpoint = $1;
+
+ $standby->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_checkpoint'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+}
+
+# Old primary, checksums off. wal_log_hints is required by pg_rewind
+# on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+# Clean switchover to B; enable checksums online on it.
+$node_a->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_a, $node_b);
+$node_b->promote;
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+$node_b->stop('fast');
+my $watermark_b = controldata_watermark($node_b);
+$node_b->start;
+
+# Accidental restart of the old primary on the old timeline. Advance
+# its WAL beyond the enable watermark of B, then run an online enable
+# and disable cycle: the node ends "off" with a watermark above every
+# checksum record B has written.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t_pad AS SELECT generate_series(1,200000) AS a;");
+enable_data_checksums($node_a, wait => 'on');
+disable_data_checksums($node_a, wait => 'off');
+test_checksum_state($node_a, 'off');
+$node_a->stop('fast');
+
+my $watermark_a = controldata_watermark($node_a);
+die "test broken: target watermark not above the source's enable"
+ unless $watermark_a > $watermark_b;
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the promoted node');
+
+# Start the rewound node as a standby of B. Replay from the last
+# common checkpoint runs through B's online enable, which the target
+# never saw, so the rewound node must converge to "on".
+my $connstr_b = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr_b application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+# Scenario 2, reusing the pair with the roles reversed. Disable
+# checksums online so the next divergence point carries "off", and let
+# A replay the change.
+disable_data_checksums($node_b, wait => 'off');
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'off');
+
+# Clean switchover back to A; enable checksums online on it.
+$node_b->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_b, $node_a);
+$node_a->promote;
+enable_data_checksums($node_a, wait => 'on');
+test_checksum_state($node_a, 'on');
+my $source_enable_watermark = controldata_watermark($node_a);
+
+# The old primary restarts on its old timeline and enables checksums
+# online independently: both control files say "on", the divergence
+# checkpoint says "off".
+$node_b->start;
+$node_b->safe_psql('postgres', "INSERT INTO t_div VALUES (2);");
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+$node_b->stop('fast');
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_b->data_dir,
+ '--source-server' => $node_a->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled on both nodes');
+
+my $connstr_a = $node_a->connstr;
+$node_b->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_b->port]}
+primary_conninfo = '$connstr_a application_name=@{[$node_b->name]}'
+]);
+$node_b->set_standby_mode;
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_b, 'on');
+
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_b->stop('fast');
+is(controldata_watermark($node_b), $source_enable_watermark,
+ 'rewound node replayed the source checksum transition');
+command_ok([ 'pg_checksums', '--check', '-D', $node_b->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the twice-rewound node');
+
+done_testing();
--
2.55.0
[application/octet-stream] v11-0002-pg_checksums-Refuse-interrupted-transitions-note.patch (7.7K, ../../CAN4CZFP_-vMOtHEj0u_v9OYoihxmi1QkfP_wiP_ytdg9FE=KfA@mail.gmail.com/4-v11-0002-pg_checksums-Refuse-interrupted-transitions-note.patch)
download | inline diff:
From cee98345803d425d1fa368b949fecb3f1e6faf5d Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:19 +0000
Subject: [PATCH v11 2/4] pg_checksums: Refuse interrupted transitions, note
change is local
An inprogress-on or inprogress-off control file means an online
transition was cut short; running the offline tool on top of it mixes
two procedures. Refuse it in every mode, with a hint pointing at the
start-stop cycle (or, on a standby, at letting replication finish)
that resets the state. Also print that the change applies to one
data directory only, with standby-aware wording, since in a
replication setup the same change must be made on every node.
On a primary, inprogress-on cannot survive a graceful stop: the
checksums launcher resolves it from its exit cleanup, and a crashed
primary is already rejected by the existing "cluster must be shut
down" check. inprogress-off can, however: it is set by the backend
running pg_disable_data_checksums(), and a fast shutdown arriving
between its two barriers leaves it behind in a cleanly shut down
control file, where the next start-stop cycle resolves it at end of
recovery. A standby can be stopped with either state, since it has
no launcher and only carries forward whatever state the last replayed
WAL record left it in, with a restartpoint persisting that as-is.
The standby is also the deterministic way to reach the new guard,
which is why the test coverage uses one; the primary window would
need an injection point between the two barriers.
The pg_checksums page said the tool still processes all relation files
regardless of an interrupted online transition; describe the refusal
instead.
---
doc/src/sgml/ref/pg_checksums.sgml | 11 ++++---
src/bin/pg_checksums/pg_checksums.c | 21 +++++++++++++
.../test_checksums/t/012_offline_standby.pl | 31 ++++++++++++++-----
3 files changed, 51 insertions(+), 12 deletions(-)
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index 1cb68187e51..cfc6518e740 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -49,11 +49,12 @@ PostgreSQL documentation
</para>
<para>
- When enabling checksums with <application>pg_checksums</application>, if
- checksums were in the process of being enabled using
- <xref linkend="checksums-online-enable-disable"/> when the cluster was shut
- down, <application>pg_checksums</application> will still process all
- relation files regardless of the progress of online checksum processing.
+ If checksums were in the process of being enabled or disabled using
+ <xref linkend="checksums-online-enable-disable"/> when the cluster was
+ shut down, the control file still records that interrupted state, and
+ <application>pg_checksums</application> refuses to run in any mode.
+ Start the cluster and shut it down cleanly to reset the state, then
+ retry; on a standby, let replication complete the transition first.
</para>
<para>
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 8c4ae59222d..95757d0424b 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -586,6 +586,21 @@ main(int argc, char *argv[])
ControlFile->state != DB_SHUTDOWNED_IN_RECOVERY)
pg_fatal("cluster must be shut down");
+ /*
+ * An inprogress state means an online transition was cut short. A
+ * standby stopped mid-transition carries either state; a cleanly shut
+ * down primary can still carry inprogress-off, which a fast shutdown
+ * during pg_disable_data_checksums() leaves behind, while inprogress-on
+ * is always resolved by the launcher's exit cleanup.
+ */
+ if (ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted");
+ pg_log_error_hint("Start and cleanly shut down the cluster once to reset the data checksum state, then retry. On a standby, let replication complete the transition first.");
+ exit(1);
+ }
+
if (ControlFile->data_checksum_version != PG_DATA_CHECKSUM_VERSION &&
mode == PG_MODE_CHECK)
pg_fatal("data checksums are not enabled in cluster");
@@ -673,6 +688,12 @@ main(int argc, char *argv[])
printf(_("Checksums enabled in cluster\n"));
else
printf(_("Checksums disabled in cluster\n"));
+
+ printf(_("This change applies to this data directory only.\n"));
+ if (ControlFile->state == DB_SHUTDOWNED_IN_RECOVERY)
+ printf(_("This node appears to be a standby; apply the same change to the primary and all other standbys.\n"));
+ else
+ printf(_("In a replication setup, apply the same change to every node while all are stopped, before restarting any of them.\n"));
}
return 0;
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
index 241a1282ae9..5efe11f993e 100644
--- a/src/test/modules/test_checksums/t/012_offline_standby.pl
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -149,7 +149,10 @@ $newnode->stop('immediate');
# Converge the cluster: enable offline on the standby too.
$standby->stop;
-$standby->checksum_enable_offline;
+command_checks_all(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ 0, [qr/appears to be a standby/],
+ [], 'standby-role notice on offline enable');
$standby->start;
test_checksum_state($standby, 'on');
$primary->wait_for_catchup($standby);
@@ -217,11 +220,11 @@ unlike(
);
# Scenario 4: a standby stopped while replaying an interrupted online
-# transition keeps the interrupted state in its own control file. A
-# primary is never caught this way, as its checksums launcher resolves
-# inprogress-on back to off from its exit cleanup. A standby has no
-# launcher; it carries forward whatever the last replayed record left
-# it in.
+# transition keeps the interrupted state in its own control file, and
+# pg_checksums must refuse to touch it. A primary is never caught this
+# way, as its checksums launcher resolves inprogress-on back to off from
+# its exit cleanup. A standby has no launcher; it carries forward
+# whatever the last replayed record left it in.
# Block an online enable on the primary at inprogress-on with a
# blocking temp table, same trick as in 004_offline.pl.
@@ -235,6 +238,20 @@ wait_for_checksum_state($standby, 'inprogress-on');
# Stop the standby cleanly; its restartpoint persists inprogress-on to
# its own control file, since nothing on a standby resolves it away.
$standby->stop;
+
+command_fails_like(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --enable refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --check refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--disable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --disable refuses a standby stopped mid-transition');
+
$standby->start;
wait_for_checksum_state($standby, 'inprogress-on');
@@ -257,7 +274,7 @@ wait_for_checksum_state($primary, 'on');
$primary->wait_for_catchup($standby);
wait_for_checksum_state($standby, 'on');
-is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
'10001', 'standby readable once the transition completes');
# Scenario 4 continued: the backup taken mid-transition must also
--
2.55.0
[application/octet-stream] v11-0004-pg_combinebackup-Refuse-mixed-data-checksum-stat.patch (9.5K, ../../CAN4CZFP_-vMOtHEj0u_v9OYoihxmi1QkfP_wiP_ytdg9FE=KfA@mail.gmail.com/5-v11-0004-pg_combinebackup-Refuse-mixed-data-checksum-stat.patch)
download | inline diff:
From 8ee5dc5b0f0047aa44799275e5f7d39880418636 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Tue, 25 Aug 2026 15:06:58 +0000
Subject: [PATCH v11 4/4] pg_combinebackup: Refuse mixed data checksum states
in a backup chain
check_control_files() detected a chain whose backups were taken under
different data checksum states, warned, and proceeded. The output
directory keeps the last backup's control file, so a full backup taken
with checksums off combined with an incremental taken after an offline
enable produces a cluster whose control file says "on" while most of
its blocks carry no checksums; it starts, and then every connection
dies on the first unchecksummed catalog page. An offline enable
between two backups of a chain is all it takes, since it rewrites
every page without logging anything, so the incremental backup does
not re-ship the pages.
Turn the warning into an error, matching what pg_rewind does for the
equivalent combinations. The check stays asymmetric on purpose: when
the last backup was taken with checksums off, stale checksums from an
earlier backup are never verified, and that chain remains usable.
---
doc/src/sgml/ref/pg_combinebackup.sgml | 19 +--
src/bin/pg_combinebackup/pg_combinebackup.c | 14 +-
src/test/modules/test_checksums/meson.build | 1 +
.../t/024_combinebackup_mixed.pl | 134 ++++++++++++++++++
4 files changed, 155 insertions(+), 13 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
diff --git a/doc/src/sgml/ref/pg_combinebackup.sgml b/doc/src/sgml/ref/pg_combinebackup.sgml
index 9a6d201e0b8..f5d7e177b36 100644
--- a/doc/src/sgml/ref/pg_combinebackup.sgml
+++ b/doc/src/sgml/ref/pg_combinebackup.sgml
@@ -306,18 +306,19 @@ PostgreSQL documentation
<para>
<literal>pg_combinebackup</literal> does not recompute page checksums when
- writing the output directory. Therefore, if any of the backups used for
- reconstruction were taken with checksums disabled, but the final backup was
- taken with checksums enabled, the resulting directory may contain pages
- with invalid checksums.
+ writing the output directory. It therefore refuses a chain in which some
+ of the backups used for reconstruction were taken with checksums disabled
+ but the final backup was taken with checksums enabled: the resulting
+ directory would contain pages with invalid checksums, and the cluster
+ would fail checksum verification as soon as it reads them. Take a new
+ full backup after enabling data checksums with
+ <xref linkend="app-pgchecksums"/>.
</para>
<para>
- To avoid this problem, taking a new full backup after changing the checksum
- state of the cluster using <xref linkend="app-pgchecksums"/> is
- recommended. Otherwise, you can disable and then optionally reenable
- checksums on the directory produced by <literal>pg_combinebackup</literal>
- in order to correct the problem.
+ The reverse case is accepted: when the final backup was taken with
+ checksums disabled, stale checksums copied from an earlier backup are
+ never verified.
</para>
</refsect1>
diff --git a/src/bin/pg_combinebackup/pg_combinebackup.c b/src/bin/pg_combinebackup/pg_combinebackup.c
index 86e2ee37c40..254a27b125b 100644
--- a/src/bin/pg_combinebackup/pg_combinebackup.c
+++ b/src/bin/pg_combinebackup/pg_combinebackup.c
@@ -670,13 +670,19 @@ check_control_files(int n_backups, char **backup_dirs)
pg_log_debug("system identifier is %" PRIu64, system_identifier);
/*
- * Warn the user if not all backups are in the same state with regards to
- * checksums.
+ * Reject a chain whose backups were not all in the same state with
+ * regards to checksums. The loop above only flags that when the last
+ * backup has checksums enabled: the blocks taken from the older backups
+ * have no checksums, and pg_combinebackup does not recompute them, so the
+ * combined cluster would fail verification as soon as it reads them. The
+ * other direction is harmless, since stale checksums are never verified
+ * while checksums are disabled.
*/
if (data_checksum_mismatch)
{
- pg_log_warning("only some backups have checksums enabled");
- pg_log_warning_hint("Disable, and optionally reenable, checksums on the output directory to avoid failures.");
+ pg_log_error("only some backups have checksums enabled");
+ pg_log_error_hint("Take a new full backup after changing the data checksum state with pg_checksums.");
+ exit(1);
}
return system_identifier;
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index 7d07c757052..7eccd5156b9 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -47,6 +47,7 @@ tests += {
't/021_rewind_divergent_transitions.pl',
't/022_rewind_state.pl',
't/023_rewind_standby_target.pl',
+ 't/024_combinebackup_mixed.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
new file mode 100644
index 00000000000..ce9dee9413b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_combinebackup refuses a chain whose backups were taken under
+# different data checksum states when the final state has them enabled.
+#
+# The output directory keeps the last backup's control file, so a full
+# backup taken with checksums off plus an incremental taken after an
+# offline enable would produce a directory whose control file says "on"
+# while most of its blocks have no checksums. An offline enable between
+# the two backups is enough to get there: it rewrites every page but logs
+# nothing, so the incremental backup does not re-ship the pages.
+#
+# The reverse order stays allowed: with the final backup taken with
+# checksums off, stale checksums from an earlier backup are never verified.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $node = PostgreSQL::Test::Cluster->new('node');
+$node->init(
+ no_data_checksums => 1,
+ has_archiving => 1,
+ allows_streaming => 1);
+$node->append_conf('postgresql.conf', 'summarize_wal = on');
+$node->append_conf('postgresql.conf', 'autovacuum = off');
+$node->start;
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+test_checksum_state($node, 'off');
+
+$node->backup('full');
+
+# Offline enable: rewrites every page, logs nothing.
+$node->stop;
+system_or_bail('pg_checksums', '--enable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'on');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full/backup_manifest',
+ ],
+ 'incremental backup with checksums on');
+
+# Combining the chain must be refused: the result would say "on" while the
+# blocks inherited from the full backup have no checksums.
+command_fails_like(
+ [
+ 'pg_combinebackup',
+ $node->backup_dir . '/full',
+ $node->backup_dir . '/incr',
+ '--output' => $node->backup_dir . '/combined',
+ ],
+ qr/only some backups have checksums enabled/,
+ 'pg_combinebackup refuses a chain crossing an offline enable');
+
+# The other direction: full backup with checksums on, offline disable, then
+# an incremental. The last backup wins, so the combined cluster comes up off.
+$node->backup('full2');
+
+$node->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr2',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full2/backup_manifest',
+ ],
+ 'incremental backup with checksums off');
+
+my $restored = PostgreSQL::Test::Cluster->new('restored');
+$restored->init_from_backup(
+ $node, 'incr2',
+ combine_with_prior => ['full2'],
+ has_restoring => 1,
+ standby => 0);
+
+$restored->start;
+
+my ($rc, $stdout, $stderr) = $restored->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'combined backup accepts connections') or diag("stderr: $stderr");
+is($stdout, 'off', 'combined backup follows the last backup state');
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t3;');
+is($rc, 0, 'blocks from the incremental backup are readable')
+ or diag("stderr: $stderr");
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t1;');
+is($rc, 0, 'blocks inherited from the full backup are readable')
+ or diag("stderr: $stderr");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($restored->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures in the combined backup');
+
+$restored->stop('immediate');
+$node->stop;
+
+done_testing();
--
2.55.0
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-03 04:08 Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
parent: Zsolt Parragi <zsolt.parragi@percona.com>
0 siblings, 1 reply; 43+ messages in thread
From: Bertrand Drouvot @ 2026-09-03 04:08 UTC (permalink / raw)
To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; pgsql-hackers@lists.postgresql.org
Hi,
On Wed, Sep 02, 2026 at 03:34:40PM +0100, Zsolt Parragi wrote:
> v11 adds a few more edits based on Daniel's feedback.
So, I compared v9 and v11, and the additional C changes look good to me (persisting
the complete checksum state at the end of recovery and protecting the control file
read with ControlFileLock).
I just have a few wording comments on 0001:
=== 1
+ * record over it. The control file also carries a watermark, the end LSN of
+ * the newest XLOG2_CHECKSUMS record this node has written or applied.
After pg_rewind, the watermark can be the divergence point rather than the end
of an XLOG2_CHECKSUMS record. Maybe use the wording from pg_control.h here?
=== 2
+ each node in the replication setup. Nodes can be processed in
+ parallel as they are shut down. Processing must have ended
s/as they are shut down/while they are shut down/?
s/must have ended/must complete/?
And also, one that was already in v9:
=== 3
+ * record: without that, a checkpoint could read the old state after the
+ * record is already in WAL and insert a redo record that both precedes the
+ * transition in WAL order and carries the pre-transition state.
I think that should be s/precedes/follows/. This would also match the wording in
CreateCheckPoint().
In the commit message:
=== 4
"
Recovery from a base backup is the exception to not adopting
"
I think this is not true for every base backup. A backup taken from a standby
keeps the copied control file state. Otherwise, adoption only happens when the
state is not node local and its watermark is below the starting checkpoint.
Maybe s/is the exception/may be an exception/, with a short mention of those
conditions?
Other than that, v11-0001 looks good to me.
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-03 11:06 Zsolt Parragi <zsolt.parragi@percona.com>
parent: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 1 reply; 43+ messages in thread
From: Zsolt Parragi @ 2026-09-03 11:06 UTC (permalink / raw)
To: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; pgsql-hackers@lists.postgresql.org
Thanks!
v12 addresses these, otherwise it is unchanged to compared 11.
Attachments:
[application/octet-stream] v12-0003-pg_rewind-Check-the-data-checksum-states-of-sour.patch (21.6K, ../../CAN4CZFNG57F8DY5NF=y=bdbCju6+_JxrBnUVYg5Xz3LOMCrSug@mail.gmail.com/2-v12-0003-pg_rewind-Check-the-data-checksum-states-of-sour.patch)
download | inline diff:
From ea29d0dc20d1e7b939c3b8716b33b4c7be9af2da Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 18:40:22 +0000
Subject: [PATCH v12 3/4] pg_rewind: Check the data checksum states of source
and target
pg_rewind installs the source's control file on the target, but every
block it does not copy keeps the target's content. With checksums
enabled on the target and disabled on the source, the rewound server
claims enabled checksums while the blocks copied from the source have
none, and fails checksum verification as soon as it reads them; in the
worst case every connection attempt dies on an unverifiable catalog
page.
Comparing the control files is not enough. Replay on the rewound
server resumes from the last common checkpoint and adopts the data
checksum state recorded there, so an offline disable on the source
after the divergence leaves both control files saying "off" while the
rewound server still resumes with verification enabled. Check the
state carried by the divergence checkpoint as well, which pg_rewind
already reads.
Refuse both cases, and an interrupted online transition on either
side, same as pg_checksums does. The opposite mismatch, checksums on
the source only, is allowed with a warning: recovery under the backup
label adopts the divergence state, so the rewound server keeps
checksums disabled (or converges through the replayed WAL if the
source enabled them online), and no page can fail verification.
---
doc/src/sgml/ref/pg_rewind.sgml | 14 ++
src/bin/pg_rewind/parsexlog.c | 4 +-
src/bin/pg_rewind/pg_rewind.c | 93 +++++++++++-
src/bin/pg_rewind/pg_rewind.h | 1 +
src/test/modules/test_checksums/meson.build | 2 +
.../test_checksums/t/022_rewind_state.pl | 133 +++++++++++++++++
.../t/023_rewind_standby_target.pl | 141 ++++++++++++++++++
7 files changed, 386 insertions(+), 2 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/022_rewind_state.pl
create mode 100644 src/test/modules/test_checksums/t/023_rewind_standby_target.pl
diff --git a/doc/src/sgml/ref/pg_rewind.sgml b/doc/src/sgml/ref/pg_rewind.sgml
index b95e3868e9d..ac3d0c9328f 100644
--- a/doc/src/sgml/ref/pg_rewind.sgml
+++ b/doc/src/sgml/ref/pg_rewind.sgml
@@ -124,6 +124,20 @@ PostgreSQL documentation
<literal>on</literal>, but is enabled by default.
</para>
+ <para>
+ The data checksum states of the source and target server must be
+ compatible. <application>pg_rewind</application> refuses to run if an
+ online checksum state transition was interrupted on either server, or
+ if data checksums are enabled on the target server, or were enabled at
+ the point of divergence, while the source server runs without them: the
+ rewound server would fail checksum verification on the blocks copied
+ from the source. The opposite combination is allowed with a warning;
+ the rewound server keeps data checksums disabled unless the source
+ server enabled them online after the point of divergence. See
+ <xref linkend="app-pgchecksums"/> for how to change the state
+ consistently across a replication setup.
+ </para>
+
<warning>
<title>Warning: Failures While Rewinding</title>
<para>
diff --git a/src/bin/pg_rewind/parsexlog.c b/src/bin/pg_rewind/parsexlog.c
index 023e23b063c..6e87b00f8c2 100644
--- a/src/bin/pg_rewind/parsexlog.c
+++ b/src/bin/pg_rewind/parsexlog.c
@@ -167,7 +167,8 @@ readOneRecord(const char *datadir, XLogRecPtr ptr, int tliIndex,
void
findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
- XLogRecPtr *lastchkptredo, const char *restoreCommand)
+ XLogRecPtr *lastchkptredo, uint32 *lastchkptdatachecksums,
+ const char *restoreCommand)
{
/* Walk backwards, starting from the given record */
XLogRecord *record;
@@ -255,6 +256,7 @@ findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
*lastchkptrec = searchptr;
*lastchkpttli = checkPoint.ThisTimeLineID;
*lastchkptredo = checkPoint.redo;
+ *lastchkptdatachecksums = checkPoint.dataChecksumState;
break;
}
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 466f4223501..9b99e628d4e 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -145,6 +145,7 @@ main(int argc, char **argv)
XLogRecPtr chkptrec;
TimeLineID chkpttli;
XLogRecPtr chkptredo;
+ uint32 chkptdatachecksums;
TimeLineID source_tli;
TimeLineID target_tli;
XLogRecPtr target_wal_endrec;
@@ -472,10 +473,51 @@ main(int argc, char **argv)
keepwal_init();
findLastCheckpoint(datadir_target, divergerec, lastcommontliIndex,
- &chkptrec, &chkpttli, &chkptredo, restore_command);
+ &chkptrec, &chkpttli, &chkptredo, &chkptdatachecksums,
+ restore_command);
pg_log_info("rewinding from last common checkpoint at %X/%08X on timeline %u",
LSN_FORMAT_ARGS(chkptrec), chkpttli);
+ /*
+ * Replay on the rewound server resumes from the last common checkpoint
+ * and adopts the data checksum state recorded there, not the state in the
+ * control file installed from the source. If checksums were enabled at
+ * the divergence point but the source runs without them, the rewound
+ * server would verify checksums while the blocks copied from the source
+ * have none. The control files cannot reveal this: an offline disable on
+ * the source after the divergence leaves both of them saying "off".
+ *
+ * Only the fully enabled state needs checking. An in-progress state at
+ * the divergence point means a transition was still running there, and
+ * all of its page rewrites are logged after that point, so replay brings
+ * the rewound server to whatever state the copied WAL ends in.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the source server or recreate the target server from a base backup.");
+ exit(1);
+ }
+
+ /*
+ * The same comparison is needed against the target: every block the
+ * rewind does not copy keeps the target's content. The control files
+ * cannot reveal this case either, in the other direction: a standby's
+ * checkpoints are written by its upstream primary, so a standby whose
+ * checksums were disabled offline still has "on" checkpoints in its WAL,
+ * while both control files may agree.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_target.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the target server");
+ pg_log_error_detail("Blocks kept from the target would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the target server with pg_checksums or recreate the target server from a base backup.");
+ exit(1);
+ }
+
/* Initialize the hash table to track the status of each file */
filehash_init();
@@ -799,6 +841,55 @@ sanityChecks(void)
pg_fatal("target server needs to use either data checksums or \"wal_log_hints = on\"");
}
+ /*
+ * The rewound target keeps its control file fields from the source, but
+ * every block the rewind does not copy keeps the target's content. If
+ * checksums are enabled on the target and disabled on the source, the
+ * result would claim enabled checksums while the blocks copied from the
+ * source have none, and the target would fail checksum verification as
+ * soon as it reads them. Refuse that combination, and refuse an
+ * interrupted online transition on either side, same as pg_checksums.
+ *
+ * The opposite mismatch is allowed: recovery under the backup label
+ * written by pg_rewind adopts the checksum state as of the divergence
+ * point, so a target whose own checkpoints carry no checksums stays
+ * without them, or converges through the replayed WAL if the source
+ * enabled them online. Only warn about it, so that an offline change on
+ * the source is not overlooked.
+ *
+ * That reasoning does not hold when the target is a standby, whose
+ * checkpoints were written by its upstream primary; see the checks after
+ * findLastCheckpoint().
+ */
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted on the target server");
+ pg_log_error_hint("Start the server and shut it down cleanly to reset the state, then retry; on a standby, let replication complete the transition first.");
+ exit(1);
+ }
+ if (ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition is incomplete on the source server");
+ pg_log_error_hint("Let the transition complete, or reset the state with a clean restart, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_OFF)
+ {
+ pg_log_error("data checksums are enabled on the target server but disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, and the target would fail checksum verification after the rewind.");
+ pg_log_error_hint("Either disable data checksums on the target server or enable them on the source server, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_OFF &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_warning("data checksums are disabled on the target server but enabled on the source server");
+ pg_log_warning_detail("The rewound server will keep data checksums disabled unless the source server enabled them online after the point of divergence.");
+ }
+
/*
* Target cluster better not be running. This doesn't guard against
* someone starting the cluster concurrently. Also, this is probably more
diff --git a/src/bin/pg_rewind/pg_rewind.h b/src/bin/pg_rewind/pg_rewind.h
index 9a981f7f246..9187c3bf72f 100644
--- a/src/bin/pg_rewind/pg_rewind.h
+++ b/src/bin/pg_rewind/pg_rewind.h
@@ -39,6 +39,7 @@ extern void findLastCheckpoint(const char *datadir, XLogRecPtr forkptr,
int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
XLogRecPtr *lastchkptredo,
+ uint32 *lastchkptdatachecksums,
const char *restoreCommand);
extern XLogRecPtr readOneRecord(const char *datadir, XLogRecPtr ptr,
int tliIndex, const char *restoreCommand);
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index e5d38fafb7d..7d07c757052 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -45,6 +45,8 @@ tests += {
't/019_standby_shutdown_catchup.pl',
't/020_cascade_divergence.pl',
't/021_rewind_divergent_transitions.pl',
+ 't/022_rewind_state.pl',
+ 't/023_rewind_standby_target.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/022_rewind_state.pl b/src/test/modules/test_checksums/t/022_rewind_state.pl
new file mode 100644
index 00000000000..b7adc968c4d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/022_rewind_state.pl
@@ -0,0 +1,133 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind checks the data checksum states of source and target.
+# Checksums enabled on the target, or at the point of divergence, with
+# a source running without them are refused: the rewound server would
+# verify checksums on blocks copied from a source that has none. The
+# opposite mismatch only warns; replay keeps the target's own state.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# Scenarios 1 and 2: checksums on from initdb. wal_log_hints keeps the
+# target eligible for pg_rewind after checksums are disabled on it.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b);
+
+# Failover to B, and divergence on A.
+$node_b->promote;
+$node_b->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_a->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_a->stop('fast');
+
+# Offline disable on the source only: enabled target, disabled source.
+$node_b->stop;
+$node_b->checksum_disable_offline;
+$node_b->start;
+test_checksum_state($node_b, 'off');
+
+my @rewind_cmd = (
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'));
+
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums are enabled on the target server but disabled on the source server/,
+ 'refuses an enabled target with a disabled source');
+
+# The refusal happens before any modification: the target still starts
+# on its own timeline.
+$node_a->start;
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'target untouched by the refused rewind');
+$node_a->stop('fast');
+
+# Scenario 2: disabling the target too makes the control files match,
+# but checksums were still enabled at the point of divergence, which is
+# the state replay on the rewound server would resume with.
+$node_a->checksum_disable_offline;
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums were enabled at the point of divergence but are disabled on the source server/,
+ 'refuses when checksums were enabled at the point of divergence');
+
+# Scenario 3: fresh pair without checksums; offline enable on the
+# source only. Allowed with a warning, and the rewound server keeps
+# checksums disabled.
+my $node_c = PostgreSQL::Test::Cluster->new('node_c');
+$node_c->init(allows_streaming => 1, no_data_checksums => 1);
+$node_c->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_c->start;
+$node_c->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_c->backup('backup');
+my $node_d = PostgreSQL::Test::Cluster->new('node_d');
+$node_d->init_from_backup($node_c, 'backup', has_streaming => 1);
+$node_d->start;
+$node_c->wait_for_catchup($node_d);
+
+$node_d->promote;
+$node_d->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_c->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_c->stop('fast');
+
+$node_d->stop;
+$node_d->checksum_enable_offline;
+$node_d->start;
+test_checksum_state($node_d, 'on');
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_c->data_dir,
+ '--source-server' => $node_d->connstr('postgres'),
+ ]);
+like(
+ $stderr,
+ qr/data checksums are disabled on the target server but enabled on the source server/,
+ 'warns for a disabled target with an enabled source');
+like($stderr, qr/Done!/, 'rewind completed despite the warning');
+
+# The rewound server follows D and keeps its own state.
+$node_c->append_conf('postgresql.conf', 'port = ' . $node_c->port);
+$node_c->enable_streaming($node_d);
+$node_c->set_standby_mode;
+$node_c->start;
+$node_d->wait_for_catchup($node_c);
+is($node_c->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+test_checksum_state($node_c, 'off');
+
+$node_c->stop;
+$node_d->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/023_rewind_standby_target.pl b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
new file mode 100644
index 00000000000..3f5a3c6be8b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
@@ -0,0 +1,141 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_rewind compares the data checksum state at the point of
+# divergence against the target as well as the source.
+#
+# A standby's checkpoints are written by its upstream primary, so a standby
+# whose checksums were disabled offline still has "on" checkpoints in its
+# WAL. Rewinding it onto a promoted sibling passes every control file
+# check (target off + source on only warns), yet recovery after the rewind
+# adopts the divergence checkpoint's state and the rewound server would
+# verify checksums over the pages it wrote while it was locally "off".
+# pg_rewind must refuse; after an offline enable on the target, which
+# rewrites all of its pages, the rewind goes through.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t0 AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+
+my $lagging = PostgreSQL::Test::Cluster->new('lagging');
+$lagging->init_from_backup($primary, 'backup', has_streaming => 1);
+$lagging->start;
+
+my $failover = PostgreSQL::Test::Cluster->new('failover');
+$failover->init_from_backup($primary, 'backup', has_streaming => 1);
+$failover->start;
+
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+
+test_checksum_state($primary, 'on');
+test_checksum_state($lagging, 'on');
+test_checksum_state($failover, 'on');
+
+# Offline-disable checksums on the lagging standby only. This is a
+# divergence 012_offline_standby.pl declares survivable: the standby keeps
+# its own state, warns, and stays readable.
+$lagging->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $lagging->data_dir);
+$lagging->start;
+test_checksum_state($lagging, 'off');
+
+# Give the lagging standby plenty of pages to write while it is locally
+# "off", and force a restartpoint so they reach disk without checksums.
+# The other standby replays the same WAL with checksums on.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+$lagging->safe_psql('postgres', "CHECKPOINT;");
+
+# Take the "failover" standby out of the picture. It is promoted from
+# this position later, so everything replayed from here on is the part of
+# the WAL that diverges and that pg_rewind will roll back. t1 is already
+# replayed everywhere, so its blocks are *not* rolled back.
+$failover->stop;
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$primary->wait_for_catchup($lagging);
+
+# Failover: promote the other standby, which forks the timeline behind the
+# replay position of the lagging standby.
+$primary->stop('fast');
+$failover->start;
+$failover->promote;
+$failover->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$failover->safe_psql('postgres', "CHECKPOINT;");
+
+# Re-attaching the lagging standby with pg_rewind must be refused: the
+# divergence checkpoint says "on" while the target is "off", so the rewound
+# server would verify checksums over the blocks it keeps.
+$lagging->stop;
+
+command_fails_like(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ],
+ qr/data checksums were enabled at the point of divergence but are disabled on the target server/,
+ 'pg_rewind refuses a target whose divergence checkpoint has checksums');
+
+# Enabling checksums offline on the target rewrites all of its pages, after
+# which the rewind is safe.
+system_or_bail('pg_checksums', '--enable', '--pgdata', $lagging->data_dir);
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ]);
+like($stderr, qr/Done!/, 'pg_rewind completes after the offline enable');
+
+$lagging->append_conf('postgresql.conf', 'port = ' . $lagging->port);
+$lagging->enable_streaming($failover);
+$lagging->set_standby_mode;
+$lagging->start;
+
+my ($rc, $out, $err) =
+ $lagging->psql('postgres',
+ "SELECT setting FROM pg_settings " . "WHERE name = 'data_checksums';");
+is($rc, 0, 'rewound standby accepts connections') or diag("stderr: $err");
+is($out, 'on', 'rewound standby resumes with checksums on');
+
+($rc, $out, $err) = $lagging->psql('postgres', "SELECT count(*) FROM t1;");
+is($rc, 0, 'blocks kept from the target stay readable')
+ or diag("stderr: $err");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($lagging->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the rewound standby');
+
+$lagging->stop('immediate');
+$failover->stop;
+done_testing();
--
2.55.0
[application/octet-stream] v12-0002-pg_checksums-Refuse-interrupted-transitions-note.patch (7.7K, ../../CAN4CZFNG57F8DY5NF=y=bdbCju6+_JxrBnUVYg5Xz3LOMCrSug@mail.gmail.com/3-v12-0002-pg_checksums-Refuse-interrupted-transitions-note.patch)
download | inline diff:
From 012eb75252399bf8d6f9342b9781547a78dd0637 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:19 +0000
Subject: [PATCH v12 2/4] pg_checksums: Refuse interrupted transitions, note
change is local
An inprogress-on or inprogress-off control file means an online
transition was cut short; running the offline tool on top of it mixes
two procedures. Refuse it in every mode, with a hint pointing at the
start-stop cycle (or, on a standby, at letting replication finish)
that resets the state. Also print that the change applies to one
data directory only, with standby-aware wording, since in a
replication setup the same change must be made on every node.
On a primary, inprogress-on cannot survive a graceful stop: the
checksums launcher resolves it from its exit cleanup, and a crashed
primary is already rejected by the existing "cluster must be shut
down" check. inprogress-off can, however: it is set by the backend
running pg_disable_data_checksums(), and a fast shutdown arriving
between its two barriers leaves it behind in a cleanly shut down
control file, where the next start-stop cycle resolves it at end of
recovery. A standby can be stopped with either state, since it has
no launcher and only carries forward whatever state the last replayed
WAL record left it in, with a restartpoint persisting that as-is.
The standby is also the deterministic way to reach the new guard,
which is why the test coverage uses one; the primary window would
need an injection point between the two barriers.
The pg_checksums page said the tool still processes all relation files
regardless of an interrupted online transition; describe the refusal
instead.
---
doc/src/sgml/ref/pg_checksums.sgml | 11 ++++---
src/bin/pg_checksums/pg_checksums.c | 21 +++++++++++++
.../test_checksums/t/012_offline_standby.pl | 31 ++++++++++++++-----
3 files changed, 51 insertions(+), 12 deletions(-)
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index aa6d30e436a..44bec1aa605 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -49,11 +49,12 @@ PostgreSQL documentation
</para>
<para>
- When enabling checksums with <application>pg_checksums</application>, if
- checksums were in the process of being enabled using
- <xref linkend="checksums-online-enable-disable"/> when the cluster was shut
- down, <application>pg_checksums</application> will still process all
- relation files regardless of the progress of online checksum processing.
+ If checksums were in the process of being enabled or disabled using
+ <xref linkend="checksums-online-enable-disable"/> when the cluster was
+ shut down, the control file still records that interrupted state, and
+ <application>pg_checksums</application> refuses to run in any mode.
+ Start the cluster and shut it down cleanly to reset the state, then
+ retry; on a standby, let replication complete the transition first.
</para>
<para>
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 8c4ae59222d..95757d0424b 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -586,6 +586,21 @@ main(int argc, char *argv[])
ControlFile->state != DB_SHUTDOWNED_IN_RECOVERY)
pg_fatal("cluster must be shut down");
+ /*
+ * An inprogress state means an online transition was cut short. A
+ * standby stopped mid-transition carries either state; a cleanly shut
+ * down primary can still carry inprogress-off, which a fast shutdown
+ * during pg_disable_data_checksums() leaves behind, while inprogress-on
+ * is always resolved by the launcher's exit cleanup.
+ */
+ if (ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted");
+ pg_log_error_hint("Start and cleanly shut down the cluster once to reset the data checksum state, then retry. On a standby, let replication complete the transition first.");
+ exit(1);
+ }
+
if (ControlFile->data_checksum_version != PG_DATA_CHECKSUM_VERSION &&
mode == PG_MODE_CHECK)
pg_fatal("data checksums are not enabled in cluster");
@@ -673,6 +688,12 @@ main(int argc, char *argv[])
printf(_("Checksums enabled in cluster\n"));
else
printf(_("Checksums disabled in cluster\n"));
+
+ printf(_("This change applies to this data directory only.\n"));
+ if (ControlFile->state == DB_SHUTDOWNED_IN_RECOVERY)
+ printf(_("This node appears to be a standby; apply the same change to the primary and all other standbys.\n"));
+ else
+ printf(_("In a replication setup, apply the same change to every node while all are stopped, before restarting any of them.\n"));
}
return 0;
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
index 241a1282ae9..5efe11f993e 100644
--- a/src/test/modules/test_checksums/t/012_offline_standby.pl
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -149,7 +149,10 @@ $newnode->stop('immediate');
# Converge the cluster: enable offline on the standby too.
$standby->stop;
-$standby->checksum_enable_offline;
+command_checks_all(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ 0, [qr/appears to be a standby/],
+ [], 'standby-role notice on offline enable');
$standby->start;
test_checksum_state($standby, 'on');
$primary->wait_for_catchup($standby);
@@ -217,11 +220,11 @@ unlike(
);
# Scenario 4: a standby stopped while replaying an interrupted online
-# transition keeps the interrupted state in its own control file. A
-# primary is never caught this way, as its checksums launcher resolves
-# inprogress-on back to off from its exit cleanup. A standby has no
-# launcher; it carries forward whatever the last replayed record left
-# it in.
+# transition keeps the interrupted state in its own control file, and
+# pg_checksums must refuse to touch it. A primary is never caught this
+# way, as its checksums launcher resolves inprogress-on back to off from
+# its exit cleanup. A standby has no launcher; it carries forward
+# whatever the last replayed record left it in.
# Block an online enable on the primary at inprogress-on with a
# blocking temp table, same trick as in 004_offline.pl.
@@ -235,6 +238,20 @@ wait_for_checksum_state($standby, 'inprogress-on');
# Stop the standby cleanly; its restartpoint persists inprogress-on to
# its own control file, since nothing on a standby resolves it away.
$standby->stop;
+
+command_fails_like(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --enable refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --check refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--disable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --disable refuses a standby stopped mid-transition');
+
$standby->start;
wait_for_checksum_state($standby, 'inprogress-on');
@@ -257,7 +274,7 @@ wait_for_checksum_state($primary, 'on');
$primary->wait_for_catchup($standby);
wait_for_checksum_state($standby, 'on');
-is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
'10001', 'standby readable once the transition completes');
# Scenario 4 continued: the backup taken mid-transition must also
--
2.55.0
[application/octet-stream] v12-0001-Do-not-adopt-data-checksum-state-from-another-no.patch (133.8K, ../../CAN4CZFNG57F8DY5NF=y=bdbCju6+_JxrBnUVYg5Xz3LOMCrSug@mail.gmail.com/4-v12-0001-Do-not-adopt-data-checksum-state-from-another-no.patch)
download | inline diff:
From 869df3c3a71c287e2a98f0700f6ea655760b4130 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 14 Aug 2026 17:06:02 +0000
Subject: [PATCH v12 1/4] Do not adopt data checksum state from another node
during replay
Offline pg_checksums changes are local to one node, but replay adopted
the data checksum state carried by checkpoint records unconditionally.
After an offline enable on the primary, a standby whose pages were
never rewritten started verifying checksums it does not have; after an
offline change on a standby, the next replayed checkpoint silently
reverted it.
To fix, make the control file track this node's state alone, and have
replay cross-check the replayed state against it instead of adopting
it, warning once per divergent value and reporting when the states
agree again. pg_control gains a watermark, normally the end LSN of
the newest XLOG2_CHECKSUMS record the node has written or applied, so
that replay can skip transition records whose effect the control file
already contains, and a flag marking a state last written by
pg_checksums, which recovery must never overwrite with a replayed one.
Since the control file may only claim "on" once every page on disk
carries a checksum, persisting the state is tied to flushes:
XLOG2_CHECKSUMS replay persists every state but "on" immediately and
leaves "on" to the next restartpoint, checkpoints and restartpoints
only persist a state their flush ran under from beginning to end, and
transitions publish their state under the new
DataChecksumTransitionLock so that the states carried by WAL records
match their WAL order. The comments in xlog.c spell out the
individual rules.
Recovery from a base backup may be an exception to not adopting: its
control file was copied at an arbitrary moment, so the state carried
by the starting checkpoint is the one the WAL from there on was
written under. That holds for a backup taken from a primary, and only
when the copied state is not node local and its watermark is below the
starting checkpoint; a backup taken from a standby keeps the copied
state. pg_rewind keeps the target's own state and clamps the
watermark to the divergence point, since a watermark set on the
target's abandoned fork could numerically cover transition records the
source wrote after the divergence.
Bump PG_CONTROL_VERSION.
Document the offline procedure for replication setups, the lockstep
one: stop all nodes, run pg_checksums on each of them, and only then
restart. An offline change writes no WAL and has no ordering against
WAL a node has not replayed yet, so a node must not stop before
replaying all WAL of its upstream, or the change is overridden on
restart.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Author: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
---
doc/src/sgml/ref/pg_checksums.sgml | 87 ++-
doc/src/sgml/wal.sgml | 23 +
src/backend/access/transam/xlog.c | 570 ++++++++++++++++--
src/backend/postmaster/datachecksum_state.c | 26 +-
.../utils/activity/wait_event_names.txt | 1 +
src/bin/pg_checksums/pg_checksums.c | 10 +
src/bin/pg_controldata/pg_controldata.c | 4 +
src/bin/pg_resetwal/pg_resetwal.c | 7 +
src/bin/pg_rewind/pg_rewind.c | 35 +-
src/bin/pg_upgrade/controldata.c | 2 +-
src/include/catalog/pg_control.h | 29 +-
src/include/storage/lwlocklist.h | 1 +
src/test/modules/test_checksums/Makefile | 2 +-
src/test/modules/test_checksums/meson.build | 10 +
.../test_checksums/t/012_offline_standby.pl | 289 +++++++++
.../modules/test_checksums/t/013_rewind.pl | 201 ++++++
.../modules/test_checksums/t/014_lockstep.pl | 182 ++++++
.../t/015_standby_crash_after_disable.pl | 125 ++++
.../t/016_promote_enable_crash.pl | 134 ++++
.../test_checksums/t/017_restartpoint_race.pl | 138 +++++
.../t/018_enable_crash_windows.pl | 496 +++++++++++++++
.../t/019_standby_shutdown_catchup.pl | 126 ++++
.../t/020_cascade_divergence.pl | 115 ++++
.../t/021_rewind_divergent_transitions.pl | 193 ++++++
24 files changed, 2722 insertions(+), 84 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/012_offline_standby.pl
create mode 100644 src/test/modules/test_checksums/t/013_rewind.pl
create mode 100644 src/test/modules/test_checksums/t/014_lockstep.pl
create mode 100644 src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
create mode 100644 src/test/modules/test_checksums/t/016_promote_enable_crash.pl
create mode 100644 src/test/modules/test_checksums/t/017_restartpoint_race.pl
create mode 100644 src/test/modules/test_checksums/t/018_enable_crash_windows.pl
create mode 100644 src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
create mode 100644 src/test/modules/test_checksums/t/020_cascade_divergence.pl
create mode 100644 src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index ae66fad3f0f..aa6d30e436a 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -242,15 +242,84 @@ PostgreSQL documentation
data directory must not be started or else data loss may occur.
</para>
<para>
- When using a replication setup with tools which perform direct copies
- of relation file blocks (for example <xref linkend="app-pgrewind"/>),
- enabling or disabling checksums can lead to page corruptions in the
- shape of incorrect checksums if the operation is not done consistently
- across all nodes. When enabling or disabling checksums in a replication
- setup, it is thus recommended to stop all the clusters before switching
- them all consistently. Destroying all standbys, performing the operation
- on the primary and finally recreating the standbys from scratch is also
- safe.
+ Enabling or disabling checksums with
+ <application>pg_checksums</application> changes only the local data
+ directory; the new state is not replicated to any other node. In a
+ replication setup the same change must be applied to every node:
+ </para>
+
+ <procedure>
+ <step>
+ <title>Shut down all nodes</title>
+ <para>
+ All nodes participating in the replication must be stopped with a
+ clean shutdown; <application>pg_checksums</application> refuses to run
+ on a data directory left behind by an immediate shutdown. Before
+ stopping a standby, make sure it has replayed all WAL of the primary:
+ stop the primary first, read its <quote>Latest checkpoint
+ location</quote> with <xref linkend="app-pgcontroldata"/>, and check
+ that <function>pg_last_wal_replay_lsn()</function> on the standby has
+ advanced past it. Comparing the replay position with
+ <function>pg_last_wal_receive_lsn()</function> is not enough, as it
+ only shows that the WAL the standby received has been replayed.
+ </para>
+ </step>
+
+ <step>
+ <title>Enable or disable data checksums on each node</title>
+ <para>
+ Run <application>pg_checksums</application> on the data directory of
+ each node in the replication setup. Nodes can be processed in
+ parallel while they are shut down. Processing must complete
+ successfully on all nodes before continuing.
+ </para>
+ </step>
+
+ <step>
+ <title>Restart all nodes</title>
+ <para>
+ Start the nodes normally, verify that
+ <xref linkend="guc-data-checksums"/> matches on all of them, and
+ monitor the logs of the standbys for data checksum state mismatch
+ warnings.
+ </para>
+ </step>
+ </procedure>
+
+ <para>
+ Tools that copy relation file blocks directly between nodes, such as
+ <xref linkend="app-pgrewind"/>, likewise require all nodes to be in
+ the same data checksum state.
+ </para>
+ <para>
+ The replay requirement exists because an offline change is recorded
+ only in the control file and has no defined ordering against WAL the
+ node has not replayed yet; see
+ <xref linkend="checksums-offline-enable-disable"/>. A node stopped
+ before replaying an online checksum state change applies that change
+ when it is restarted, overriding the offline change, and the states of
+ the nodes silently diverge until a later checkpoint record triggers
+ the warning described below. Because of this it is best not to mix
+ the two mechanisms: change the state of a replication setup either
+ with the offline procedure above or with an online transition, and
+ make sure the previous change has reached every node before starting
+ the next one.
+ </para>
+ <para>
+ If the change is applied inconsistently, each node keeps its own state,
+ and a standby logs a warning when the state recorded in the replayed WAL
+ differs from its own. The same warning can appear transiently while a
+ standby catches up over WAL written before a consistent change; it stops
+ once a checkpoint record carrying the new state has been replayed.
+ </para>
+ <para>
+ A standby whose data directory was never checksummed must not have
+ checksums enabled by catching up this way. Converge the cluster by
+ running <application>pg_checksums</application> on it while stopped as
+ outlined above, by enabling checksums online, or by recreating it from
+ a base backup. Note
+ that an online enable only starts from a primary whose checksums are off,
+ so if they are already enabled there, disable them online first.
</para>
<para>
If <application>pg_checksums</application> is aborted or killed while
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index ec62d17fbcc..dbbfdae5e4f 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -317,6 +317,29 @@
verify checksums, on an offline cluster.
</para>
+ <para>
+ An offline change provides durability differently from an
+ <link linkend="checksums-online-enable-disable">online change</link>.
+ An online transition is WAL-logged: it is ordered against all other
+ WAL records, it is replayed after a crash, and it propagates to
+ standbys. An offline change is recorded only in the cluster's
+ control file: it writes no WAL, it is invisible to replication, and
+ it has no defined ordering against WAL the node has not replayed
+ yet. When a node later replays WAL that contains an online checksum
+ state change, that change takes effect on the node even if it was
+ written before the offline change was made.
+ </para>
+
+ <para>
+ An offline change only affects the data directory it is run on; the
+ new state does not propagate over replication. In a replication setup
+ the same change must be applied to all nodes while all of them are
+ stopped, as described in <xref linkend="app-pgchecksums"/>. A standby
+ whose state diverges logs a warning but keeps its local setting. The
+ mismatch persists until the states are brought together again, with
+ the offline procedure or with an online transition; do this promptly.
+ </para>
+
</sect2>
<sect2 id="checksums-online-enable-disable" xreflabel="Online Enabling of Checksums">
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index de4c96e135f..ac466305bad 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -556,14 +556,22 @@ typedef struct XLogCtlData
*/
XLogRecPtr lastFpwDisableRecPtr;
- /* last data_checksum_version we've seen */
+ /* current data checksum state of this node */
uint32 data_checksum_version;
+ /*
+ * Copies of control file fields with the same names, see pg_control.h for
+ * an in-depth description of these fields. Must be updated together with
+ * data_checksum_version under info_lck.
+ */
+ XLogRecPtr data_checksum_lsn;
+ bool data_checksum_is_local;
+
slock_t info_lck; /* locks shared variables shown above */
/*
* lastChecksumChangeRecPtr points to the end of the last XLOG2_CHECKSUMS
- * record inserted or replayed, i.e. the last change of
+ * record inserted or replayed which corresponds to the last change of
* data_checksum_version. InvalidXLogRecPtr if the state hasn't changed
* since the server started.
*/
@@ -690,6 +698,20 @@ static ChecksumStateType LocalDataChecksumState = 0;
*/
int data_checksums = 0;
+/*
+ * Whether replay of the next checkpoint-family record must adopt the data
+ * checksum state it carries. Set when recovery starts from a base backup,
+ * where the state at the redo point takes precedence over the control file
+ * copied with the backup later.
+ */
+static bool adoptChecksumStateFromNextCheckpoint = false;
+
+/*
+ * Sentinel for the data checksum mismatch warning tracking in
+ * CheckReplayedDataChecksumState(): no warning is outstanding.
+ */
+#define NO_WARNING_ISSUED PG_UINT32_MAX
+
/* For WALInsertLockAcquire/Release functions */
static int MyLockNo = 0;
static bool holdingAllLocks = false;
@@ -730,6 +752,8 @@ static void ValidateXLOGDirectoryStructure(void);
static void CleanupBackupHistory(void);
static void UpdateMinRecoveryPoint(XLogRecPtr lsn, bool force);
static bool PerformRecoveryXLogAction(void);
+static void CheckReplayedDataChecksumState(uint32 replayed_version);
+static void AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn);
static void InitControlFile(uint64 sysidentifier, uint32 data_checksum_version);
static void WriteControlFile(void);
static void ReadControlFile(void);
@@ -757,7 +781,7 @@ static void WALInsertLockAcquireExclusive(void);
static void WALInsertLockRelease(void);
static void WALInsertLockUpdateInsertingAt(XLogRecPtr insertingAt);
-static void XLogChecksums(uint32 new_type);
+static XLogRecPtr XLogChecksums(uint32 new_type);
/*
* Insert an XLOG record represented by an already-constructed chain of data
@@ -4292,6 +4316,8 @@ InitControlFile(uint64 sysidentifier, uint32 data_checksum_version)
ControlFile->track_commit_timestamp = track_commit_timestamp;
ControlFile->data_checksum_version = data_checksum_version;
ControlFile->data_checksum_version_init = data_checksum_version;
+ ControlFile->data_checksum_is_local = false;
+ ControlFile->data_checksum_lsn = InvalidXLogRecPtr;
/*
* Set the data_checksum_version value into XLogCtl, which is where all
@@ -4774,6 +4800,7 @@ void
SetDataChecksumsOnInProgress(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
/*
* The state transition is performed in a critical section with
@@ -4782,14 +4809,12 @@ SetDataChecksumsOnInProgress(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4827,6 +4852,8 @@ void
SetDataChecksumsOn(void)
{
uint64 barrier;
+ bool persist;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4847,23 +4874,11 @@ SetDataChecksumsOn(void)
SpinLockRelease(&XLogCtl->info_lck);
INJECTION_POINT("datachecksums-enable-checksums-delay", NULL);
+ INJECTION_POINT_LOAD("datachecksums-on-before-publish");
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_VERSION);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- SpinLockRelease(&XLogCtl->info_lck);
-
- /*
- * Update the controlfile before waiting since if we have an immediate
- * shutdown while waiting we want to come back up with checksums enabled.
- */
- LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- UpdateControlFile();
- LWLockRelease(ControlFileLock);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_VERSION);
barrier = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_CHECKSUM_ON);
@@ -4873,6 +4888,40 @@ SetDataChecksumsOn(void)
INJECTION_POINT("datachecksums-on-before-checkpoint", NULL);
RequestCheckpoint(CHECKPOINT_FORCE | CHECKPOINT_WAIT | CHECKPOINT_FAST);
+
+ INJECTION_POINT("datachecksums-on-after-checkpoint", NULL);
+
+ /*
+ * Persist "on" only now that the checkpoint has flushed the pages the
+ * transition rewrote. Crash recovery initializes verification from the
+ * control file but resumes from a checkpoint that can predate the
+ * transition, so an "on" persisted earlier would verify pages during
+ * replay whose rewrite never reached the disk: the pages found by the data
+ * checksums worker in shared buffers are not written back by its ring
+ * buffer.
+ *
+ * The checkpoint above normally persists the state itself; this covers the
+ * case where it started before the record written above and left the field
+ * alone. Crashing before this point is safe, as replay then re-establishes
+ * "on" from the full page images of the rewrite. Skip the write if the
+ * state moved on meanwhile, since whatever moved it persists its own.
+ * Compare the watermark rather than the state: a state comparison could
+ * not tell our transition from a later round trip back to "on".
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ persist = (XLogCtl->data_checksum_lsn == recptr);
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (persist)
+ {
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+ }
+
WaitForProcSignalBarrier(barrier);
}
@@ -4893,6 +4942,7 @@ void
SetDataChecksumsOff(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4918,14 +4968,12 @@ SetDataChecksumsOff(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4956,14 +5004,12 @@ SetDataChecksumsOff(void)
/* Ensure that we don't incur a checkpoint during disabling checksums */
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -5001,6 +5047,138 @@ SetLocalDataChecksumState(uint32 data_checksum_version)
data_checksums = data_checksum_version;
}
+/*
+ * CheckReplayedDataChecksumState
+ * Cross-check the data checksum state carried by a replayed checkpoint
+ * record against the state of this node.
+ *
+ * The state in the record belongs to whichever node wrote the WAL, and must
+ * not be adopted: an offline state change made with pg_checksums on one node
+ * of a replication set generates no WAL, and adopting would leak it into the
+ * other nodes through replay. Backup label recovery is the one exception,
+ * see AdoptReplayedDataChecksumState(). XLOG_CHECKPOINT_ONLINE needs no call
+ * here, as an online checkpoint's state already traveled in the preceding
+ * XLOG_CHECKPOINT_REDO record.
+ *
+ * Only archive recovery can see a lasting mismatch, as only there can the WAL
+ * and the control file come from different nodes or different times. In
+ * crash recovery a mismatch means replay resumed from a restartpoint
+ * predating an already-applied XLOG2_CHECKSUMS record, and replaying forward
+ * re-establishes the same state.
+ */
+static void
+CheckReplayedDataChecksumState(uint32 replayed_version)
+{
+ /*
+ * Warn once per remote value, so a lasting mismatch does not flood the
+ * log. Matching states re-arm the warning. Backend-local state is
+ * enough: replay only runs in the startup process, and restarting it
+ * re-arms as well.
+ */
+ static uint32 last_warned_version = NO_WARNING_ISSUED;
+ uint32 local_version;
+
+ if (!ArchiveRecoveryRequested)
+ return;
+
+ /*
+ * Re-replayed WAL below the consistency point was already cross-checked
+ * before minRecoveryPoint was last persisted, and the persisted state can
+ * legitimately be newer than what checkpoint records there carry:
+ * XLOG2_CHECKSUMS replay persists most states ahead of the restartpoint
+ * horizon. In particular the checkpoint record recovery restarts from is
+ * such a re-replay.
+ */
+ if (!reachedConsistency)
+ return;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ local_version = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (replayed_version == local_version)
+ {
+ /*
+ * Report convergence if this process warned before. Nothing else
+ * tells the operator that running pg_checksums on the other nodes, or
+ * a rebuild, took effect.
+ */
+ if (last_warned_version != NO_WARNING_ISSUED)
+ ereport(LOG,
+ errmsg("data checksum state \"%s\" of this node now agrees with the replayed WAL",
+ get_checksum_state_string(local_version)));
+
+ last_warned_version = NO_WARNING_ISSUED;
+ return;
+ }
+
+ /* the nodes legitimately differ while an online transition runs */
+ if (replayed_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ replayed_version == PG_DATA_CHECKSUM_INPROGRESS_OFF ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ return;
+
+ if (replayed_version == last_warned_version)
+ return;
+ last_warned_version = replayed_version;
+
+ ereport(WARNING,
+ errmsg("data checksum state \"%s\" of this node does not match the state \"%s\" in the replayed WAL",
+ get_checksum_state_string(local_version),
+ get_checksum_state_string(replayed_version)),
+ errdetail("The data checksum state was most likely changed with pg_checksums on another node."),
+ errhint("Apply the same change with pg_checksums on the primary and all standby servers, or rebuild this server from a base backup."));
+}
+
+/*
+ * AdoptReplayedDataChecksumState
+ * Adopt the data checksum state at the redo point of backup label
+ * recovery.
+ *
+ * The state is persisted immediately so that a crash before the first
+ * restartpoint does not resurrect the state copied with the backup; a crash
+ * at this point restarts from the same redo point, so the control file does
+ * not run ahead of the replay position. If the value is unchanged the
+ * control file already carries it, so both the barrier and the persist are
+ * skipped.
+ *
+ * lsn is the location of the checkpoint-family record the state was taken
+ * from and becomes the new watermark: the adopted state covers everything
+ * below the redo point, which replay never revisits. For the same reason
+ * the old watermark can stay when the value is unchanged: the records
+ * between the two positions are never replayed, so nothing depends on
+ * which one is recorded.
+ */
+static void
+AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn)
+{
+ bool changed = false;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (XLogCtl->data_checksum_version != new_version)
+ {
+ XLogCtl->data_checksum_version = new_version;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(new_version);
+ changed = true;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (!changed)
+ return;
+
+ EmitAndWaitDataChecksumsBarrier(new_version);
+
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = new_version;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+}
+
/* guc hook */
const char *
show_data_checksums(void)
@@ -5454,6 +5632,8 @@ XLOGShmemInit(void *arg)
/* Use the checksum info from control file */
XLogCtl->data_checksum_version = ControlFile->data_checksum_version;
+ XLogCtl->data_checksum_lsn = ControlFile->data_checksum_lsn;
+ XLogCtl->data_checksum_is_local = ControlFile->data_checksum_is_local;
SetLocalDataChecksumState(XLogCtl->data_checksum_version);
SpinLockInit(&XLogCtl->Insert.insertpos_lck);
@@ -6031,6 +6211,52 @@ StartupXLOG(void)
SetCommitTsLimit(checkPoint.oldestCommitTsXid,
checkPoint.newestCommitTsXid);
+ /*
+ * When recovery starts from a base backup, the control file was copied at
+ * an arbitrary moment and its data checksum state may differ from the
+ * state at the redo point, which is what the WAL from there on was
+ * written under. Adopt the state of the starting checkpoint: a shutdown
+ * checkpoint is not replayed, so take it from the record read above; the
+ * redo point of an online checkpoint is its CHECKPOINT_REDO record, so
+ * let the replay of that record adopt it. Check backupStartPoint in
+ * addition to the label: on a crash restart during backup recovery the
+ * label file is already renamed away, but the start point persists until
+ * the backup end record.
+ *
+ * Not for a base backup taken from a standby, though. Its starting
+ * checkpoint is the standby's last restartpoint, a record written by the
+ * upstream primary, whose state is not the one the copied files were
+ * written under. The copied control file is already correct: a standby
+ * persists its state only at restartpoint horizons and never claims
+ * more than what reached disk. Such backups are recognized by
+ * backupEndPoint together with backupEndRequired; backupEndPoint is only
+ * set for "BACKUP FROM: standby" labels and persists across a crash
+ * restart. pg_rewind writes a standby label as well, but no
+ * backupEndPoint, and its recovery keeps adopting: the control file it
+ * installs carries the target's own checksum state, which can lag the
+ * redo point of the last common checkpoint the same way a restartpoint
+ * horizon can.
+ *
+ * Never adopt over a state the control file's watermark or local flag
+ * marks as newer than the starting checkpoint. A pg_checksums change is
+ * local to the node and generates no WAL, so nothing in the replayed WAL
+ * could ever restore it once overwritten; and a watermark above the redo
+ * point means the control file already contains the effect of every
+ * transition record up to there.
+ */
+ if ((haveBackupLabel || XLogRecPtrIsValid(ControlFile->backupStartPoint)) &&
+ !(XLogRecPtrIsValid(ControlFile->backupEndPoint) &&
+ ControlFile->backupEndRequired) &&
+ !ControlFile->data_checksum_is_local &&
+ checkPoint.redo > ControlFile->data_checksum_lsn)
+ {
+ if (wasShutdown)
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ checkPoint.redo);
+ else
+ adoptChecksumStateFromNextCheckpoint = true;
+ }
+
/*
* Clear out any old relcache cache files. This is *necessary* if we do
* any WAL replay, since that would probably result in the cache files
@@ -6633,11 +6859,7 @@ StartupXLOG(void)
if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
ereport(WARNING,
@@ -6654,11 +6876,7 @@ StartupXLOG(void)
else if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
}
@@ -6683,6 +6901,8 @@ StartupXLOG(void)
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
XLogCtl->SharedRecoveryState = RECOVERY_STATE_DONE;
SpinLockRelease(&XLogCtl->info_lck);
@@ -6814,6 +7034,23 @@ static bool
PerformRecoveryXLogAction(void)
{
bool promoted = false;
+ bool flushForChecksums;
+ uint32 checksum_state;
+
+ /*
+ * The end-of-recovery record persists the data checksum state without
+ * flushing the buffer pool, but the control file may only claim "on" once
+ * every page on disk carries a checksum. If replay entered that state
+ * without a restartpoint following it, the pages rewritten by the
+ * transition are still only in the buffer pool, so take the full
+ * checkpoint below instead of the lightweight record.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ flushForChecksums = (checksum_state == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile->data_checksum_version != checksum_state);
/*
* Perform a checkpoint to update all our recovery activity to disk.
@@ -6829,7 +7066,7 @@ PerformRecoveryXLogAction(void)
* fully out of recovery mode and already accepting queries.
*/
if (ArchiveRecoveryRequested && IsUnderPostmaster &&
- PromoteIsTriggered())
+ PromoteIsTriggered() && !flushForChecksums)
{
promoted = true;
@@ -7436,6 +7673,7 @@ CreateCheckPoint(int flags)
uint32 freespace;
XLogRecPtr PriorRedoPtr;
XLogRecPtr last_important_lsn;
+ XLogRecPtr checksumLsn;
VirtualTransactionId *vxids;
int nvxids;
int oldXLogAllowed = 0;
@@ -7548,11 +7786,14 @@ CreateCheckPoint(int flags)
checkPoint.wal_level = wal_level;
/*
- * Get the current data_checksum_version value from xlogctl, valid at the
- * time of the checkpoint.
+ * Get the current data_checksum_version value from xlogctl. This is
+ * final only for a shutdown checkpoint, where no concurrent transition
+ * is possible; an online checkpoint resamples it together with the redo
+ * record below.
*/
SpinLockAcquire(&XLogCtl->info_lck);
checkPoint.dataChecksumState = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
if (shutdown)
@@ -7609,10 +7850,21 @@ CreateCheckPoint(int flags)
{
xl_checkpoint_redo redo_rec;
+ /*
+ * Sample the data checksum state and insert the redo record under
+ * DataChecksumTransitionLock, so that a concurrent transition cannot
+ * insert its XLOG2_CHECKSUMS record between the sampling and the
+ * insertion below. Without this, the redo record could follow the
+ * transition record in WAL while carrying the pre-transition state,
+ * and recovery resuming here would never learn about the transition.
+ * See XLogChecksums().
+ */
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
WALInsertLockAcquire();
redo_rec.wal_level = wal_level;
SpinLockAcquire(&XLogCtl->info_lck);
redo_rec.data_checksum_version = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
WALInsertLockRelease();
@@ -7620,6 +7872,14 @@ CreateCheckPoint(int flags)
XLogBeginInsert();
XLogRegisterData(&redo_rec, sizeof(xl_checkpoint_redo));
(void) XLogInsert(RM_XLOG_ID, XLOG_CHECKPOINT_REDO);
+ LWLockRelease(DataChecksumTransitionLock);
+
+ /*
+ * The checkpoint record must carry the same state as the redo record
+ * just inserted: the sample taken before redo determination can be
+ * stale by now, and the pair would otherwise disagree.
+ */
+ checkPoint.dataChecksumState = redo_rec.data_checksum_version;
/*
* XLogInsertRecord will have updated XLogCtl->Insert.RedoRecPtr in
@@ -7823,6 +8083,40 @@ CreateCheckPoint(int flags)
ControlFile->minRecoveryPoint = InvalidXLogRecPtr;
ControlFile->minRecoveryPointTLI = 0;
+ /*
+ * Persist the data checksum state this node runs under. Only the
+ * top-level field tracks this node; ControlFile->checkPointCopy above is
+ * a historical record used to resume replay.
+ *
+ * checkPoint.dataChecksumState was sampled under
+ * DataChecksumTransitionLock together with the redo record, so it is the
+ * state in effect at the redo point. If it was "on",
+ * the XLOG2_CHECKSUMS record announcing that precedes the redo point and
+ * every page the transition rewrote was dirtied before it, so
+ * CheckPointGuts() has just written all of them out. Recording the state
+ * here is what keeps a finished transition from being resolved as
+ * interrupted when this checkpoint is the one crash recovery resumes
+ * from: replay never sees the record announcing it.
+ *
+ * Persist it only if the state did not change while the flush was in
+ * progress. If it changed in between, the pages written out straddle two
+ * states, and the newer one could claim checksums that pages already on
+ * disk do not carry; leave the field to the next checkpoint then, the
+ * transition itself has already persisted every state that is safe
+ * without a flush.
+ * Compare the watermark rather than the state: record positions are
+ * unique, so a full round trip back to the sampled state cannot alias,
+ * while its flushed pages straddle the intermediate states all the same.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksumLsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checkPoint.dataChecksumState;
+ ControlFile->data_checksum_lsn = checksumLsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
/*
* Persist unloggedLSN value. It's reset on crash recovery, so this goes
* unused on non-shutdown checkpoints, but seems useful to store it always
@@ -7967,9 +8261,11 @@ CreateEndOfRecoveryRecord(void)
ControlFile->minRecoveryPoint = recptr;
ControlFile->minRecoveryPointTLI = xlrec.ThisTimeLineID;
- /* start with the latest checksum version (as of the end of recovery) */
+ /* persist the data checksum state this node ended recovery with */
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
@@ -8174,6 +8470,9 @@ CreateRestartPoint(int flags)
XLogRecPtr endptr;
XLogSegNo _logSegNo;
TimestampTz xtime;
+ uint32 checksum_state;
+ XLogRecPtr checksum_lsn;
+ bool checksum_is_local;
/* Concurrent checkpoint/restartpoint cannot happen */
Assert(!IsUnderPostmaster || MyBackendType == B_CHECKPOINTER);
@@ -8220,8 +8519,47 @@ CreateRestartPoint(int flags)
UpdateMinRecoveryPoint(InvalidXLogRecPtr, true);
if (flags & CHECKPOINT_IS_SHUTDOWN)
{
+ bool catchUpChecksums;
+
+ /*
+ * There is no new restartpoint to persist the data checksum state
+ * with, but a cleanly stopped node should not leave the control
+ * file behind the state replay reached: pg_checksums and
+ * pg_rewind read it, and an in-progress state there makes them
+ * refuse to run. Catching it up needs the same guarantee a
+ * restartpoint gives: that every page on disk carries a checksum,
+ * so flush the buffer pool first. Only a transition to
+ * "on" that no restartpoint followed can get here; the other
+ * states are already persisted by XLOG2_CHECKSUMS replay. Replay
+ * has ended by now, so the state cannot change under us.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ checksum_is_local = XLogCtl->data_checksum_is_local;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockAcquire(ControlFileLock, LW_SHARED);
+ catchUpChecksums =
+ (checksum_lsn != ControlFile->data_checksum_lsn &&
+ XLogRecPtrIsValid(lastCheckPointRecPtr));
+ LWLockRelease(ControlFileLock);
+
+ if (catchUpChecksums)
+ {
+ MemSet(&CheckpointStats, 0, sizeof(CheckpointStats));
+ CheckpointStats.ckpt_start_t = GetCurrentTimestamp();
+ CheckPointGuts(lastCheckPoint.redo, flags);
+ }
+
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
+ if (catchUpChecksums)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = checksum_is_local;
+ }
UpdateControlFile();
LWLockRelease(ControlFileLock);
}
@@ -8262,6 +8600,17 @@ CreateRestartPoint(int flags)
/* Update the process title */
update_checkpoint_display(flags, true, false);
+ /*
+ * Note the data checksum state the flush below starts under. Replay runs
+ * concurrently and can change the state while the flush is in progress,
+ * in which case the flush covers pages written under both states; see
+ * where the state is persisted further down.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
CheckPointGuts(lastCheckPoint.redo, flags);
/*
@@ -8321,8 +8670,26 @@ CreateRestartPoint(int flags)
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
}
- /* we shall start with the latest checksum version */
- ControlFile->data_checksum_version = lastCheckPoint.dataChecksumState;
+ /*
+ * Persist the data checksum state of this node. Not the state of the
+ * replayed checkpoint: that one belongs to the node that wrote it and
+ * may differ after an offline change on either side.
+ * ControlFile->checkPointCopy above keeps the replayed value on
+ * purpose, being a historical record used to resume replay rather
+ * than a tracker of node state.
+ *
+ * Persist only if the flush above ran under one state throughout; see
+ * CreateCheckPoint() for why, including why this compares the
+ * watermark and not the state.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksum_lsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
}
@@ -8763,9 +9130,22 @@ XLogReportParameters(void)
}
/*
- * Log the new state of checksums
+ * XLogChecksums
+ * Log and publish the new state of checksums
+ *
+ * Inserting the record and publishing the new state must be atomic with
+ * respect to a checkpoint sampling the state for its XLOG_CHECKPOINT_REDO
+ * record: without that, a checkpoint could read the old state after the
+ * record is already in WAL and insert a redo record that both follows the
+ * transition in WAL order and carries the pre-transition state. Recovery
+ * resuming from such a redo point would never replay the transition record
+ * and resolve the finished transition as interrupted.
+ * DataChecksumTransitionLock serializes the two; see CreateCheckPoint().
+ *
+ * Returns the end LSN of the inserted record, which the caller persists
+ * together with the new state as the data checksum watermark.
*/
-static void
+static XLogRecPtr
XLogChecksums(uint32 new_type)
{
xl_checksum_state xlrec;
@@ -8773,12 +9153,28 @@ XLogChecksums(uint32 new_type)
xlrec.new_checksum_state = new_type;
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
+
XLogBeginInsert();
XLogRegisterData((char *) &xlrec, sizeof(xl_checksum_state));
recptr = XLogInsert(RM_XLOG2_ID, XLOG2_CHECKSUMS);
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, recptr);
+
+ /* only loaded by SetDataChecksumsOn(), a no-op for the other callers */
+ INJECTION_POINT_CACHED("datachecksums-on-before-publish", NULL);
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ XLogCtl->data_checksum_version = new_type;
+ XLogCtl->data_checksum_lsn = recptr;
+ XLogCtl->data_checksum_is_local = false;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockRelease(DataChecksumTransitionLock);
+
XLogFlush(recptr);
+
+ return recptr;
}
/*
@@ -8966,11 +9362,19 @@ xlog_redo(XLogReaderState *record)
/* ControlFile->checkPointCopy always tracks the latest ckpt XID */
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->checkPointCopy.nextXid = checkPoint.nextXid;
- ControlFile->data_checksum_version = checkPoint.dataChecksumState;
UpdateControlFile();
LWLockRelease(ControlFileLock);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(checkPoint.dataChecksumState);
+
/*
* We should've already switched to the new TLI before replaying this
* record.
@@ -9206,19 +9610,17 @@ xlog_redo(XLogReaderState *record)
else if (info == XLOG_CHECKPOINT_REDO)
{
xl_checkpoint_redo redo_rec;
- bool new_state = false;
memcpy(&redo_rec, XLogRecGetData(record), sizeof(xl_checkpoint_redo));
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = redo_rec.data_checksum_version;
- SetLocalDataChecksumState(redo_rec.data_checksum_version);
- if (redo_rec.data_checksum_version != ControlFile->data_checksum_version)
- new_state = true;
- SpinLockRelease(&XLogCtl->info_lck);
-
- if (new_state)
- EmitAndWaitDataChecksumsBarrier(redo_rec.data_checksum_version);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(redo_rec.data_checksum_version,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(redo_rec.data_checksum_version);
}
else if (info == XLOG_LOGICAL_DECODING_STATUS_CHANGE)
{
@@ -9280,25 +9682,63 @@ xlog2_redo(XLogReaderState *record)
{
xl_checksum_state state;
XLogRecPtr lsn = record->EndRecPtr;
+ XLogRecPtr watermark;
memcpy(&state, XLogRecGetData(record), sizeof(xl_checksum_state));
+ SpinLockAcquire(&XLogCtl->info_lck);
+ watermark = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ /*
+ * Skip records this node has already applied. The control file
+ * carries the watermark, so this holds across restarts: recovery
+ * resuming below a record whose effect the control file already
+ * contains must not re-apply it, or it would revert a state change
+ * made with pg_checksums in between, which moves the state without
+ * writing any record of its own.
+ */
+ if (lsn <= watermark)
+ return;
+
/* advertise the location before the new state becomes visible */
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, lsn);
SpinLockAcquire(&XLogCtl->info_lck);
XLogCtl->data_checksum_version = state.new_checksum_state;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(state.new_checksum_state);
SpinLockRelease(&XLogCtl->info_lck);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = state.new_checksum_state;
+
+ /*
+ * Persist the new state, except when it is "on". Only "on" verifies
+ * checksums during reads, and between the last restartpoint and this
+ * record there may be pages on disk flushed under the old state; a
+ * crash-restart initializes verification from the control file and
+ * replay reads those pages back, so the control file may only say
+ * "on" once everything written under the transition has been flushed,
+ * as restartpoints and the end of recovery do.
+ * The opposite direction cannot wait for the restartpoint: once this
+ * record is replayed, evicted pages are written without checksums,
+ * and a control file still saying "on" would fail verification on
+ * exactly those pages after a crash.
+ */
+ if (state.new_checksum_state != PG_DATA_CHECKSUM_VERSION)
+ {
+ ControlFile->data_checksum_version = state.new_checksum_state;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ }
/*
* Update minRecoveryPoint to ensure that if recovery is aborted, we
* recover back up to this point before allowing hot standby again.
- * The new state is durable in pg_control while its location is only
- * tracked in shared memory; a standby becoming consistent below this
- * record would let base backups resume checksum verification with the
+ * The change location is only tracked in shared memory and is lost
+ * over a restart; a standby becoming consistent below this record
+ * would let base backups resume checksum verification with the
* location unknown. The local copies cannot be updated as long as
* crash recovery is happening and we expect all the WAL to be
* replayed.
diff --git a/src/backend/postmaster/datachecksum_state.c b/src/backend/postmaster/datachecksum_state.c
index f69258bc33d..099a6b4fe2e 100644
--- a/src/backend/postmaster/datachecksum_state.c
+++ b/src/backend/postmaster/datachecksum_state.c
@@ -94,9 +94,9 @@
*
* If processing is started in an online cluster then all backends are in Bd.
* If processing was halted by the cluster shutting down (due to a crash or
- * intentional restart), the controlfile state "inprogress-on" will be observed
- * on system startup and all backends will be placed in Bd. The controlfile
- * state will also be set to "off".
+ * intentional restart), the control file state "inprogress-on" will be
+ * observed on system startup and all backends will be placed in Bd. The
+ * control file state will also be set to "off".
*
* Backends transition Bd -> Bi via a procsignalbarrier which is emitted by the
* DataChecksumsWorkerLauncherMain. When all backends have acknowledged the
@@ -146,7 +146,25 @@
* stop writing data checksums as no backend is enforcing data checksum
* validation any longer.
*
- * 4. Future opportunities for optimizations
+ * 4. Interaction with offline data checksum changes
+ * -------------------------------------------------
+ * Enabling or disabling checksums offline with pg_checksums uses none of the
+ * machinery in this file, but the two mechanisms share the state kept in the
+ * control file, so their interaction is documented here.
+ *
+ * pg_checksums writes the new state to the control file and sets
+ * data_checksum_is_local, marking a state that no WAL record accounts for.
+ * Recovery then does not adopt the state carried by a replayed checkpoint
+ * record over it. The control file also carries a watermark, the WAL
+ * position through which data checksum transitions are covered. Replay skips
+ * transition records ending at or below the watermark, as their effect is
+ * already contained in the control file, and applies records above it as
+ * usual, whether they were written before or after an offline change. This
+ * is why an offline change in a replicated setup must be made on every node
+ * while all of them are stopped and caught up; see the pg_checksums
+ * documentation for the procedure.
+ *
+ * 5. Future opportunities for optimizations
* -----------------------------------------
* Below are some potential optimizations and improvements which were brought
* up during reviews of this feature, but which weren't implemented in the
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 256b3a3c02e..77fa543f9d8 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -371,6 +371,7 @@ WaitLSN "Waiting to read or update shared Wait-for-LSN state."
LogicalDecodingControl "Waiting to read or update logical decoding status information."
DataChecksumsWorker "Waiting for data checksums worker."
AioWorkerControl "Waiting to update AIO worker information."
+DataChecksumTransition "Waiting for a data checksum state transition to be written to WAL."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 3b3ae23f1a6..8c4ae59222d 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -648,6 +648,16 @@ main(int argc, char *argv[])
ControlFile->data_checksum_version =
(mode == PG_MODE_ENABLE) ? PG_DATA_CHECKSUM_VERSION : PG_DATA_CHECKSUM_OFF;
+ /*
+ * Mark the state as changed locally, without a WAL record. Recovery
+ * then does not let a replayed checkpoint overwrite it, as no record
+ * could restore the change afterwards. The watermark is left alone:
+ * XLOG2_CHECKSUMS records at or below it stay covered, while records
+ * above it, which this node has not applied yet, still take effect
+ * on replay no matter when they were written.
+ */
+ ControlFile->data_checksum_is_local = true;
+
if (do_sync)
{
pg_log_info("syncing data directory");
diff --git a/src/bin/pg_controldata/pg_controldata.c b/src/bin/pg_controldata/pg_controldata.c
index 6fc87ed114d..c363ce3dbb7 100644
--- a/src/bin/pg_controldata/pg_controldata.c
+++ b/src/bin/pg_controldata/pg_controldata.c
@@ -349,6 +349,10 @@ main(int argc, char *argv[])
(ControlFile->float8ByVal ? _("by value") : _("by reference")));
printf(_("Data page checksum version: %u\n"),
ControlFile->data_checksum_version);
+ printf(_("Data checksum watermark: %X/%08X\n"),
+ LSN_FORMAT_ARGS(ControlFile->data_checksum_lsn));
+ printf(_("Data checksum state is node-local: %s\n"),
+ (ControlFile->data_checksum_is_local ? _("yes") : _("no")));
printf(_("Default char data signedness: %s\n"),
(ControlFile->default_char_signedness ? _("signed") : _("unsigned")));
printf(_("Mock authentication nonce: %s\n"),
diff --git a/src/bin/pg_resetwal/pg_resetwal.c b/src/bin/pg_resetwal/pg_resetwal.c
index 1542a56ca4b..cdfb9898066 100644
--- a/src/bin/pg_resetwal/pg_resetwal.c
+++ b/src/bin/pg_resetwal/pg_resetwal.c
@@ -923,6 +923,13 @@ RewriteControlFile(void)
ControlFile.backupEndPoint = InvalidXLogRecPtr;
ControlFile.backupEndRequired = false;
+ /*
+ * The old WAL is gone and the new position may lie below the old
+ * watermark, which would make replay ignore future checksum transition
+ * records. The state itself is kept.
+ */
+ ControlFile.data_checksum_lsn = InvalidXLogRecPtr;
+
/*
* Force the defaults for max_* settings. The values don't really matter
* as long as wal_level='minimal'; the postmaster will reset these fields
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 2e86fd158d0..466f4223501 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -37,7 +37,8 @@ static void usage(const char *progname);
static void perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo);
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec);
static void createBackupLabel(XLogRecPtr startpoint, TimeLineID starttli,
XLogRecPtr checkpointloc);
@@ -531,7 +532,8 @@ main(int argc, char **argv)
* This is the point of no return. Once we start copying things, there is
* no turning back!
*/
- perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo);
+ perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo,
+ divergerec);
if (showprogress)
pg_log_info("syncing target data directory");
@@ -566,7 +568,8 @@ static void
perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo)
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec)
{
XLogRecPtr endrec;
TimeLineID endtli;
@@ -738,6 +741,32 @@ perform_rewind(filemap_t *filemap, rewind_source *source,
ControlFile_new.minRecoveryPoint = endrec;
ControlFile_new.minRecoveryPointTLI = endtli;
ControlFile_new.state = DB_IN_ARCHIVE_RECOVERY;
+
+ /*
+ * Keep the target's own data checksum state. Most of the data directory
+ * is still the target's: only blocks it changed since the divergence
+ * were copied from the source, so the source's state says nothing about
+ * the pages that stay. Replay from the last common checkpoint applies
+ * any WAL-logged transition the target has not seen (the watermark tells
+ * them apart), which converges the rewound server to the source's state
+ * whenever the WAL carries it.
+ */
+ ControlFile_new.data_checksum_version = ControlFile_target.data_checksum_version;
+ ControlFile_new.data_checksum_lsn = ControlFile_target.data_checksum_lsn;
+ ControlFile_new.data_checksum_is_local = ControlFile_target.data_checksum_is_local;
+
+ /*
+ * The watermark is only meaningful within the history the node replays.
+ * Records at or below the divergence point are common to both histories
+ * and stay covered, but a watermark above it was set by a transition
+ * record on the target's own abandoned fork: numerically it can cover
+ * transition records the source wrote after the divergence, and replay
+ * would skip them as already applied. Clamp it to the divergence point,
+ * so that every transition record on the source's history takes effect.
+ */
+ if (ControlFile_new.data_checksum_lsn > divergerec)
+ ControlFile_new.data_checksum_lsn = divergerec;
+
if (!dry_run)
update_controlfile(datadir_target, &ControlFile_new, do_sync);
}
diff --git a/src/bin/pg_upgrade/controldata.c b/src/bin/pg_upgrade/controldata.c
index b3bd4ccde83..de539479bdc 100644
--- a/src/bin/pg_upgrade/controldata.c
+++ b/src/bin/pg_upgrade/controldata.c
@@ -431,7 +431,7 @@ get_control_data(ClusterInfo *cluster)
cluster->controldata.date_is_int = strstr(p, "64-bit integers") != NULL;
got_date_is_int = true;
}
- else if ((p = strstr(bufin, "checksum")) != NULL)
+ else if ((p = strstr(bufin, "Data page checksum version:")) != NULL)
{
p = strchr(p, ':');
diff --git a/src/include/catalog/pg_control.h b/src/include/catalog/pg_control.h
index 7b5404460ec..5365f4e8288 100644
--- a/src/include/catalog/pg_control.h
+++ b/src/include/catalog/pg_control.h
@@ -22,7 +22,7 @@
/* Version identifier for this pg_control format */
-#define PG_CONTROL_VERSION 1903
+#define PG_CONTROL_VERSION 1904
/* Nonce key length, see below */
#define MOCK_AUTH_NONCE_LEN 32
@@ -237,6 +237,33 @@ typedef struct ControlFileData
/* Current data checksums state */
uint32 data_checksum_version;
+ /*
+ * WAL position through which data checksum transitions are covered.
+ * Replay ignores XLOG2_CHECKSUMS records ending at or below this point:
+ * their effect is already contained in data_checksum_version, or an
+ * offline pg_checksums change made after they were first applied
+ * supersedes them. Ordinarily this is the end of the newest such record
+ * this node has written or applied, but a tool may store any position
+ * that covers the same set of records. If the node has never written or
+ * applied such a record this field shall be set to InvalidXLogRecPtr.
+ *
+ * The comparison has no timeline context, so the value is only valid
+ * within the WAL history this node replays. A tool that moves the node
+ * to another history must clamp the watermark to the point where the
+ * histories fork, as pg_rewind does, or reset it, as pg_resetwal does.
+ */
+ XLogRecPtr data_checksum_lsn;
+
+ /*
+ * True when data_checksum_version was last set by pg_checksums in an
+ * offline operation rather than by an online, WAL-logged transition.
+ * Such a state is local to this node and not derived from WAL, so
+ * recovery must not replace it with a state taken from a checkpoint
+ * record; nothing in the WAL could restore the change once it is
+ * overwritten. Cleared by the next WAL-logged transition.
+ */
+ bool data_checksum_is_local;
+
/*
* True if the default signedness of char is "signed" on a platform where
* the cluster is initialized.
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index d7eb648bd27..8d858be9927 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -89,6 +89,7 @@ PG_LWLOCK(54, WaitLSN)
PG_LWLOCK(55, LogicalDecodingControl)
PG_LWLOCK(56, DataChecksumsWorker)
PG_LWLOCK(57, AioWorkerControl)
+PG_LWLOCK(58, DataChecksumTransition)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/test/modules/test_checksums/Makefile b/src/test/modules/test_checksums/Makefile
index 71455cd5577..80f54bf6d8c 100644
--- a/src/test/modules/test_checksums/Makefile
+++ b/src/test/modules/test_checksums/Makefile
@@ -9,7 +9,7 @@
#
#-------------------------------------------------------------------------
-EXTRA_INSTALL = src/test/modules/injection_points
+EXTRA_INSTALL = contrib/pg_buffercache src/test/modules/injection_points
export enable_injection_points
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index fb7129d796f..e5d38fafb7d 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -35,6 +35,16 @@ tests += {
't/009_fpi.pl',
't/010_backup_straddle.pl',
't/011_standby_straddle.pl',
+ 't/012_offline_standby.pl',
+ 't/013_rewind.pl',
+ 't/014_lockstep.pl',
+ 't/015_standby_crash_after_disable.pl',
+ 't/016_promote_enable_crash.pl',
+ 't/017_restartpoint_race.pl',
+ 't/018_enable_crash_windows.pl',
+ 't/019_standby_shutdown_catchup.pl',
+ 't/020_cascade_divergence.pl',
+ 't/021_rewind_divergent_transitions.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
new file mode 100644
index 00000000000..241a1282ae9
--- /dev/null
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -0,0 +1,289 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Offline checksum changes with pg_checksums are local to one node. A
+# standby must neither adopt the state of the primary from replayed
+# checkpoint records, nor lose its own offline change to them.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', 'autovacuum = off');
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Scenario 1: enable offline on the primary only. The standby must
+# stay off, warn about the mismatch, and remain readable.
+$standby->stop;
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+test_checksum_state($primary, 'on');
+test_checksum_state($standby, 'off');
+
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after offline enable on the primary');
+
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+my @warnings = $log =~ /(does not match the state)/g;
+is(scalar(@warnings), 1, 'mismatch warned once per remote value');
+
+# Matching states re-arm the warning: undo the divergence on the primary,
+# then diverge again to the same value, all without restarting the standby.
+$primary->stop;
+$primary->checksum_disable_offline;
+$logstart = -s $standby->logfile;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no warning while the states match again');
+like($log, qr/now agrees with the replayed WAL/,
+ 'convergence is reported once the states match again');
+
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+test_checksum_state($standby, 'off');
+
+# The local state survives both clean and immediate restarts.
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+# Still in scenario 1's divergence: take a base backup *from the
+# standby*. A backup taken on a standby uses the last restartpoint as
+# its starting checkpoint (do_pg_backup_start()), so the record at the
+# redo point was written by the upstream primary and carries the
+# primary's state. Recovery from such a backup must not adopt that
+# state: the files were copied from the standby, and the new node
+# would come up verifying checksums its files do not have.
+
+# Make sure the standby has written pages under its own "off" state, so
+# its files really do lack checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,50000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', "SELECT count(*) FROM t2;");
+
+$standby->backup('from_standby');
+my $newnode = PostgreSQL::Test::Cluster->new('newnode');
+$newnode->init_from_backup($standby, 'from_standby');
+
+# Stream from the primary, not from the standby the files were copied
+# from, so the new node replays the "on" primary's records.
+$newnode->enable_streaming($primary);
+$newnode->start;
+
+# The new node's files all came from a cluster running with checksums
+# off. Anything but "off" here means it adopted the primary's state
+# through the checkpoint record at the redo point.
+my ($rc, $stdout, $stderr) = $newnode->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'the node copied from the standby accepts connections')
+ or diag("stderr: $stderr");
+is($stdout, 'off', 'backup of an "off" standby comes up with checksums off');
+
+# A checkpoint record streamed from the "on" primary must not flip the
+# state either.
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($newnode);
+test_checksum_state($newnode, 'off');
+
+# And it must be able to read the pages the standby wrote without
+# checksums.
+(undef, $stdout, $stderr) =
+ $newnode->psql('postgres', "SELECT count(*) FROM t2;");
+is($stdout, '50000', 'pages copied from the standby are readable')
+ or diag("stderr: $stderr");
+
+my $newnode_log = PostgreSQL::Test::Utils::slurp_file($newnode->logfile);
+unlike(
+ $newnode_log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the node copied from the standby');
+
+$newnode->stop('immediate');
+
+# Converge the cluster: enable offline on the standby too.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+$primary->wait_for_catchup($standby);
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after converging');
+
+# Scenario 2: disable offline on the standby only. The replayed
+# checkpoint records of the still-enabled primary must not override it.
+$standby->stop;
+$standby->checksum_disable_offline;
+$standby->start;
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'off');
+
+# Restartpoints must persist the local state, not the replayed copy.
+$standby->safe_psql('postgres', "CHECKPOINT;");
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable with checksums disabled locally');
+
+# Scenario 3: crash-restart right after an online transition, before the
+# next restartpoint. Replay then resumes from an older restartpoint whose
+# checkpoint records still carry the pre-transition state. Those must
+# still match the state seeded from the control file, and the transition
+# itself must be re-established by re-replaying the XLOG2_CHECKSUMS record,
+# without a spurious mismatch warning along the way.
+
+# Converge first: bring the standby back to "on" offline.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+disable_data_checksums($primary, wait => 'off');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+# Crash-restart the standby immediately, before any restartpoint has had a
+# chance to persist the new state to its control file.
+$logstart = -s $standby->logfile;
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001',
+ 'standby readable after crash-restart across an online transition');
+
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no spurious mismatch warning after crash-restart across an online transition'
+);
+
+# Scenario 4: a standby stopped while replaying an interrupted online
+# transition keeps the interrupted state in its own control file. A
+# primary is never caught this way, as its checksums launcher resolves
+# inprogress-on back to off from its exit cleanup. A standby has no
+# launcher; it carries forward whatever the last replayed record left
+# it in.
+
+# Block an online enable on the primary at inprogress-on with a
+# blocking temp table, same trick as in 004_offline.pl.
+my $bsession = $primary->background_psql('postgres');
+$bsession->query_safe('CREATE TEMPORARY TABLE tt (a integer);');
+enable_data_checksums($primary, wait => 'inprogress-on');
+
+# The standby picks up the in-progress state from the XLOG2_CHECKSUMS record.
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# Stop the standby cleanly; its restartpoint persists inprogress-on to
+# its own control file, since nothing on a standby resolves it away.
+$standby->stop;
+$standby->start;
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# The primary still sits at inprogress-on, so a base backup taken now
+# copies a control file and a redo point that both carry the
+# in-progress state; replay of the XLOG2_CHECKSUMS records completes
+# the transition on the new standby.
+$primary->backup('inprogress_backup');
+my $standby2 = PostgreSQL::Test::Cluster->new('standby2');
+$standby2->init_from_backup($primary, 'inprogress_backup',
+ has_streaming => 1);
+$standby2->start;
+
+# Backup label recovery adopts the redo point's state immediately,
+# before any further WAL is replayed.
+wait_for_checksum_state($standby2, 'inprogress-on');
+
+$bsession->quit;
+wait_for_checksum_state($primary, 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable once the transition completes');
+
+# Scenario 4 continued: the backup taken mid-transition must also
+# complete the transition and stay readable.
+$primary->wait_for_catchup($standby2);
+wait_for_checksum_state($standby2, 'on');
+
+is($standby2->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'mid-transition backup standby readable after the transition');
+
+$standby2->stop;
+
+# Nothing along this standby's replay can legitimately disagree: it
+# started out at the same inprogress-on state as the primary.
+my $standby2_log = PostgreSQL::Test::Utils::slurp_file($standby2->logfile);
+unlike(
+ $standby2_log,
+ qr/does not match the state/,
+ 'no mismatch warning on a backup taken mid-transition');
+
+# No extra CHECKPOINT needed: the transition's completion checkpoint
+# wrote every page with a checksum, and the shutdown restartpoint
+# flushes the rest.
+command_ok([ 'pg_checksums', '--check', '-D', $standby2->data_dir ],
+ 'checksums valid on the mid-transition backup standby');
+
+$standby->stop;
+$primary->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/013_rewind.pl b/src/test/modules/test_checksums/t/013_rewind.pl
new file mode 100644
index 00000000000..a791e24317d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/013_rewind.pl
@@ -0,0 +1,201 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test pg_rewind across an online data checksum enable.
+#
+# A clean switchover leaves the shutdown checkpoint of the old primary
+# as the last common checkpoint between the two nodes. When data
+# checksums are enabled online on the new primary before the old one is
+# rewound, pg_rewind installs the control file of the new primary, which
+# already claims checksums are fully enabled, while replay begins at the
+# shutdown checkpoint whose record still carries the old state. Replay
+# of the WAL stretch from before the enable must run with checksums off,
+# as recorded in the checkpoint, else it would verify pages which never
+# had checksums written and fail recovery.
+#
+# The new primary requests a checkpoint right after promotion, and
+# replaying its CHECKPOINT_REDO record would repair the state before
+# any interesting WAL is reached. Hold the checkpointer on the standby
+# in a restartpoint over the promotion, like in the recovery test
+# 041_checkpoint_at_promote, so that the post-promotion writes end up
+# in WAL before the first checkpoint of the new timeline.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Old primary. full_page_writes is off so that the updates done on the
+# promoted node do not carry full page images, forcing replay on the
+# rewound node to read the pages from disk. wal_log_hints is required
+# by pg_rewind on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+full_page_writes = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+log_checkpoints = on
+]);
+$node_a->start;
+
+if (!$node_a->check_extension('injection_points'))
+{
+ plan skip_all => 'Extension injection_points not installed';
+}
+
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+$node_a->safe_psql('postgres', "CREATE EXTENSION injection_points;");
+
+# Set the hint bits on t before taking the backup, so that reads on the
+# promoted node do not emit full page images for its pages later.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->safe_psql('postgres', "SELECT count(*) FROM t;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_a, 'off');
+test_checksum_state($node_b, 'off');
+
+# Hold the next restartpoint on the standby.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point', 'wait');");
+
+# Give the restartpoint a checkpoint record to work on, then start it
+# in a background session; it will block on the injection point with
+# the checkpointer busy until released.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+my $bg_psql = $node_b->background_psql('postgres', on_error_stop => 0);
+$bg_psql->query_until(
+ qr/starting_restartpoint/, q(
+ \echo starting_restartpoint
+ CHECKPOINT;
+));
+$node_b->wait_for_event('checkpointer', 'create-restart-point');
+
+# Clean switchover: the shutdown checkpoint of A streams to B and
+# becomes the last common checkpoint.
+$node_a->stop('fast');
+
+my ($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+$stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+my $shutdown_ckpt = $1;
+$node_b->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_ckpt'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+
+my $logstart = -s $node_b->logfile;
+$node_b->promote;
+
+# Accidental restart of the old primary, diverging its timeline.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->stop('fast');
+
+# Updates on the new primary before its first checkpoint; replay of
+# these on the rewound node has to read the pages from disk.
+$node_b->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 25 = 0;");
+ok( !$node_b->log_contains("checkpoint complete", $logstart),
+ "no checkpoint on the new timeline before the updates");
+
+# Release the checkpointer; the queued post-promotion checkpoint runs
+# after the updates.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+$bg_psql->quit;
+
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+# pg_rewind refuses to run with full_page_writes disabled on the
+# source; the updates it was disabled for are already in WAL.
+$node_b->safe_psql('postgres', "ALTER SYSTEM SET full_page_writes = on;");
+$node_b->safe_psql('postgres', "SELECT pg_reload_conf();");
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the new primary');
+
+# Replay on the rewound node must start at the shutdown checkpoint of
+# the switchover, with the control file of the new primary.
+my $backup_label = slurp_file($node_a->data_dir . '/backup_label');
+$backup_label =~ /^CHECKPOINT LOCATION: ([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from backup_label";
+is($1, $shutdown_ckpt, 'replay starts at the switchover checkpoint');
+
+($stdout, $stderr) = run_command(
+ [
+ 'pg_waldump',
+ '-p' => $node_a->data_dir . '/pg_wal',
+ '-t' => 1,
+ '-s' => $shutdown_ckpt,
+ '-n' => 1,
+ ]);
+like($stdout, qr/CHECKPOINT_SHUTDOWN/,
+ 'last common checkpoint is a shutdown checkpoint');
+
+# pg_rewind keeps the target's own checksum state in the control file it
+# writes; the online enable reaches the rewound node through WAL replay
+# below, not through the copied control file.
+($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+like(
+ $stdout,
+ qr/^Data page checksum version:\s*0$/m,
+ 'rewound node keeps its own checksum state in the control file');
+
+# Start the rewound node as a standby of the new primary. Replay runs
+# through the pre-enable WAL stretch and the online enable.
+#
+# The rewind replaced the configuration files with those of the new
+# primary, so put the port back.
+my $connstr = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_b->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/014_lockstep.pl b/src/test/modules/test_checksums/t/014_lockstep.pl
new file mode 100644
index 00000000000..713f37fb22c
--- /dev/null
+++ b/src/test/modules/test_checksums/t/014_lockstep.pl
@@ -0,0 +1,182 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# The lockstep procedure for offline checksum changes in a replication
+# setup: stop all nodes, run pg_checksums on all of them, restart.
+# Replay of checkpoint records written before the change must not
+# revert the state of the standby.
+#
+# The same pair then runs the procedure in the disable direction, and
+# finally across a divergence: after a failover both nodes get their
+# checksums enabled offline, while the last common checkpoint still
+# carries "off". pg_rewind must keep the offline "on" on the rewound
+# node, since no record in the replayed WAL could ever restore it.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# wal_log_hints keeps the pair eligible for pg_rewind without checksums.
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Part 1: the lockstep procedure in the enable direction. Stop the
+# standby first: the WAL written after this point is replayed only
+# after the offline switch, and every checkpoint record in it still
+# carries the old state.
+$standby->stop;
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 0;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 1;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+# The lockstep procedure.
+$primary->checksum_enable_offline;
+$standby->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to off.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "off" in the replayed WAL/,
+ 0);
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep enable');
+
+# Crash the standby and replay the same stretch again.
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'on');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after crash restart');
+
+# Once a post-switch checkpoint has been replayed the states match and
+# no warning may be logged.
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no mismatch warning once the states match');
+
+# Every page the standby wrote in this window must carry a checksum.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$standby->stop;
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'checksums valid on the standby');
+
+# Part 2: the lockstep procedure in the other direction. The standby
+# is already stopped; give it a pre-switch WAL stretch to replay, with
+# every checkpoint record in it still carrying "on" -- an explicit
+# checkpoint plus the shutdown checkpoint written by stop().
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 2;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+$primary->checksum_disable_offline;
+$standby->checksum_disable_offline;
+
+my $disable_logstart = -s $standby->logfile;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to on.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $disable_logstart);
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep disable');
+
+# Part 3: failover, divergence and pg_rewind across offline enables.
+# The roles swap from here on: the standby becomes the new primary and
+# the old primary is rewound to follow it, but the variables keep the
+# names they had above.
+#
+# Failover: promote the standby, then diverge the old primary.
+$standby->promote;
+$standby->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$primary->stop;
+
+# The lockstep procedure across the divergence: both nodes get their
+# checksums enabled offline.
+$primary->checksum_enable_offline;
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+
+# The states match, so the rewind proceeds without complaint.
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $primary->data_dir,
+ '--source-server' => $standby->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled offline on both nodes');
+
+# The rewind replaced the configuration files with those of the source,
+# so put the port back. The copied file also carries the source's own
+# stale primary_conninfo; enable_streaming() below appends ours, which
+# wins as the later entry.
+$primary->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$primary->port]}
+]);
+$primary->enable_streaming($standby);
+
+my $rewind_logstart = -s $primary->logfile;
+$primary->start;
+$standby->wait_for_catchup($primary);
+
+is($primary->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+
+# The offline enable survives: replay from the common checkpoint must
+# not resurrect the pre-divergence "off".
+test_checksum_state($primary, 'on');
+
+$log =
+ PostgreSQL::Test::Utils::slurp_file($primary->logfile, $rewind_logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no divergence reported between the rewound server and its source');
+
+$primary->stop;
+$standby->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
new file mode 100644
index 00000000000..315b27b93d4
--- /dev/null
+++ b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
@@ -0,0 +1,125 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that a standby crashing after a replayed online disable, but before
+# its next restartpoint, restarts cleanly.
+#
+# Pages dirtied before the XLOG2_CHECKSUMS record and evicted after it are
+# written out under the new "off" state, without checksums. The control
+# file must follow the record immediately in this direction: a crash-restart
+# initializes checksum verification from the control file, and one still
+# saying "on" would fail verification on exactly those pages while replaying
+# records older than the state change.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$primary->start;
+
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,1000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+
+# Small buffer pool so that replaying a bulk load evicts the pages we care
+# about, and no restartpoints so the control file stays where it is.
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 1MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_delay = 10000
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'on');
+
+# No full page images, so that replay has to read the pages of "t" back from
+# disk instead of overwriting them from the WAL.
+$primary->append_conf('postgresql.conf', 'full_page_writes = off');
+$primary->reload;
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" on the standby, *before* the state change. They are
+# not flushed: the standby has no restartpoint from here on.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+# Online disable. The standby replays it and switches to "off", but its
+# control file is not updated.
+$primary->safe_psql('postgres', 'SELECT pg_disable_data_checksums();');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+my ($ctl_before) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($ctl_state) = $ctl_before =~ /Data page checksum version:\s+(\d+)/;
+note(
+ "standby control file data_checksum_version after the disable: "
+ . "$ctl_state (0 = off, 1 = on)");
+
+# Force the standby to evict the dirty pages of "t" now that it is "off":
+# they are written back without checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE filler AS SELECT generate_series(1,300000) AS a;");
+$primary->wait_for_catchup($standby);
+
+# Crash the standby before it gets a chance to run a restartpoint.
+$standby->stop('immediate');
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'standby restarts after crashing between a checksum state '
+ . 'change and the next restartpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'the pages written while "off" are readable')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("standby log tail:\n" . join("\n", @lines[ -12 .. -1 ]))
+ if @lines >= 12;
+ fail('the pages written while "off" are readable');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/016_promote_enable_crash.pl b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
new file mode 100644
index 00000000000..97f84da5685
--- /dev/null
+++ b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a standby promoted right after replaying an online enable, before any
+# restartpoint has flushed the rewritten pages.
+#
+# The end-of-recovery record persists the checksum state without flushing the
+# buffer pool, while the control file still points at a restartpoint older than
+# the transition. Persisting "on" there would make a crash before the
+# post-promotion checkpoint verify checksums over the pages that were flushed
+# while the state was still "off"; promotion must take the full checkpoint
+# instead.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Dirty the pages of "t" without full page images, then push them out to disk
+# on the standby while checksums are still off.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Online enable on the primary; the standby replays it into its own buffers,
+# where the rewritten pages stay dirty (no restartpoint, no bgwriter).
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+note("standby control file before the promotion: $before");
+
+# Promote. The end-of-recovery record persists the live state without any
+# buffer flush; the checkpoint requested afterwards is not immediate.
+$standby->promote;
+$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery();')
+ or die 'timed out waiting for the promotion';
+$standby->stop('immediate');
+
+($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($ckpt) = $ctl =~ /Latest checkpoint location:\s+(\S+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note(
+ "standby control file after the promotion: $after, checkpoint $ckpt, redo $redo"
+);
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'promoted node restarts after crashing before its first checkpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/017_restartpoint_race.pl b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
new file mode 100644
index 00000000000..65aa834f366
--- /dev/null
+++ b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
@@ -0,0 +1,138 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a restartpoint whose flush races the replay of an online enable.
+#
+# Replay keeps running while CheckPointGuts() writes out the buffer pool, so
+# the state at the end of the flush can be newer than the one the flush ran
+# under. Persisting it would claim checksums for pages the flush wrote out
+# before the transition, against a redo pointer that predates it.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# C1: the checkpoint the pending restartpoint will target.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" without full page images and push them to disk on
+# the standby while it is still "off".
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Start a restartpoint for C1 and hold it right after CheckPointGuts().
+$standby->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point','wait');");
+my $bg = $standby->background_psql('postgres');
+$bg->query_until(qr//, "\\echo restartpoint\nCHECKPOINT;\n");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint injection point';
+
+# The startup process keeps replaying while the checkpointer is held: the
+# whole online enable lands, and the rewritten pages stay dirty.
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+# Let the restartpoint finish; it persists the state it samples now.
+$standby->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint to finish';
+$standby->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+
+$standby->stop('immediate');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note("standby control file after the restartpoint: $after, redo $redo");
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started, 'standby restarts after a restartpoint that raced the enable');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/018_enable_crash_windows.pl b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
new file mode 100644
index 00000000000..17f96ff6272
--- /dev/null
+++ b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
@@ -0,0 +1,496 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Crash and checkpoint windows inside an online data checksum enable on
+# a single node. Each scenario starts from "off" on the same cluster,
+# runs the enable into a held injection point, and checks what a crash
+# or a concurrent checkpoint in that window may and may not do to the
+# control file and the pages on disk.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+use IPC::Run;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Bring the node back to "off" with no leftover table between
+# scenarios. The crash restarts already resolve an interrupted enable
+# to "off", so only disable when something is left to disable. Those
+# restarts also clear any attached injection points. The checkpoint
+# bounds the WAL the next scenario's crash recovery has to replay.
+sub reset_scenario
+{
+ my ($node) = @_;
+
+ BAIL_OUT('node is not running after the previous scenario')
+ unless $node->is_alive;
+
+ my $state = $node->safe_psql('postgres', 'SHOW data_checksums;');
+ disable_data_checksums($node, wait => 'off') if $state ne 'off';
+ $node->safe_psql('postgres', 'DROP TABLE IF EXISTS t;');
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+ test_checksum_state($node, 'off');
+}
+
+my $node = PostgreSQL::Test::Cluster->new('crash_windows');
+$node->init(no_data_checksums => 1);
+
+# Scenario 1 needs the pages of "t" to reach disk without a full page
+# image, hence full_page_writes and wal_log_hints off. Scenario 2
+# needs them to stay dirty in shared buffers until a checkpoint, hence
+# no bgwriter and a large shared_buffers. wal_level is pinned because
+# init() writes "minimal" for a non-streaming node. The rest merely
+# tolerate the settings.
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+wal_level = replica
+max_wal_size = 10GB
+full_page_writes = off
+wal_log_hints = off
+shared_buffers = 512MB
+bgwriter_lru_maxpages = 0
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$node->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+
+test_checksum_state($node, 'off');
+
+# Scenario 1: a primary crashes inside SetDataChecksumsOn(), just before the
+# forced checkpoint that flushes the rewritten pages, with crash recovery
+# resuming from a checkpoint older than the transition. The control file
+# must still say "inprogress-on" there: replay does not adopt the state of
+# the checkpoint record it resumes from, so an "on" written before the flush
+# would stay in effect while replay reads pages whose rewrite never reached
+# disk. Here the pages went out through the rewriting worker's ring buffer,
+# so only the control file state discriminates; see the next scenario for
+# the case where they did not.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Establish the checkpoint that crash recovery will resume from.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty the pages of "t" without full page images, then push them out
+ # to disk while checksums are still off.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+ my $evicted = $node->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+ note("evict_relation on t: $evicted, relpath $relpath");
+
+ # Stop the enable right after the control file has been updated to
+ # "on" but before the checkpoint that flushes the rewritten pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+ $node->safe_psql('postgres', 'SELECT pg_enable_data_checksums();');
+
+ $node->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'datachecksums-on-before-checkpoint';"
+ ) or die 'timed out waiting for the injection point';
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ $node->stop('immediate');
+
+ # Show the on-disk checksum field of the first page of "t".
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 2: the same crash as above, but with the rewritten pages resident
+# in shared buffers instead of gone out through the ring. The rewriting
+# worker reads through a BAS_VACUUM ring, which writes pages back as the
+# ring recycles, but a page already resident in shared buffers is not read
+# through the ring: ReadBufferExtended() hands back the existing buffer, and
+# it stays dirty until a checkpoint. The control file may therefore not
+# say "on" before that checkpoint has run, or crash recovery would resume
+# from a checkpoint older than the transition with verification already
+# enabled, and replay records that read those still-unchecksummed pages.
+# full_page_writes is off so that the records do not simply overwrite the
+# pages with a full page image.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,100000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Establish the checkpoint that crash recovery will resume from, with the
+ # pages of "t" written out while checksums are still off.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty those pages again without emitting full page images, and leave
+ # them in shared buffers. Replay of these records has to read the
+ # pages from disk.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+
+ my $dirty_before = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_before, '>', 0,
+ 'pages of t are resident and dirty before enabling checksums');
+
+ # Hold the enabling right before the checkpoint that flushes the rewritten
+ # pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ # The rewritten pages must still be sitting dirty in shared buffers, or
+ # the window this test is about does not exist.
+ my $dirty_after = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_after, '>', 0,
+ 'rewritten pages of t are still dirty before the checkpoint');
+
+ $node->stop('immediate');
+
+ # The on-disk copy is the one written before the transition, without a
+ # checksum.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ is($pd_checksum, 0, 'block 0 of t on disk carries no checksum');
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 3: a checkpoint that runs between the XLOG2_CHECKSUMS("on")
+# record and the control file write at the end of SetDataChecksumsOn() must
+# not make a crash throw the completed transition away. SetDataChecksumsOn()
+# writes the record, flips shared memory to "on", emits the barrier and only
+# then requests the checkpoint that flushes the rewritten pages; the control
+# file is written after that checkpoint returns. A crash in that window is
+# harmless only as long as recovery still starts before the record. Any
+# checkpoint completing in the window moves the redo point past the record,
+# so recovery would never see it, would come up with the control file's
+# "inprogress-on" and StartupXLOG() would demote that to "off", even though
+# every page on disk carries a checksum by then. CreateCheckPoint()
+# therefore persists the state the checkpoint ran under, the same way
+# CreateRestartPoint() does. Here the window is made deterministic by
+# holding the launcher at the datachecksums-on-before-checkpoint injection
+# point and checkpointing from another session.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Hold the launcher after the record, the shared memory flip and the
+ # barrier, but before the checkpoint SetDataChecksumsOn() requests
+ # itself.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ # Every backend already sees "on" and writes checksums.
+ test_checksum_state($node, 'on');
+
+ # ... while the control file still says "inprogress-on".
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($before, '3', 'control file says "inprogress-on" inside the window');
+
+ # A concurrent checkpoint. It flushes the rewritten pages and moves
+ # the redo point past the XLOG2_CHECKSUMS("on") record, so it has to
+ # record the "on" state in the control file as well.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($after, '1',
+ 'the concurrent checkpoint records "on" in the control file');
+
+ $node->stop('immediate');
+
+ # The checkpoint flushed the rewritten pages, so they carry a checksum on
+ # disk: the transition really did complete.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ isnt($pd_checksum, 0, 'block 0 of t on disk carries a checksum');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition is complete on disk, so the cluster has to come back
+ # "on".
+ test_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+}
+reset_scenario($node);
+
+# Scenario 4: a crash between the checkpoint an online enable requests and
+# the control file write that follows it must not lose the transition. The
+# last steps of SetDataChecksumsOn() are
+#
+# WAL record -> shmem -> barrier -> checkpoint -> persist
+#
+# The checkpoint is what makes "on" safe to persist, but it also moves the
+# redo point above the XLOG2_CHECKSUMS record that carries the new state.
+# Crash recovery started from that checkpoint therefore never replays the
+# record, so without the state the checkpoint itself persists, a crash in
+# the remaining window would bring the cluster back at "inprogress-on",
+# which StartupXLOG() resolves to "off", discarding a transition whose
+# pages are all on disk with a checksum. The previous scenario exercises
+# the same window through a checkpoint requested by another session; this
+# one closes it from the other side: the transition's own checkpoint is the
+# one that moves the redo point, so persisting the state from
+# CreateCheckPoint() is what has to save it, not the write below the
+# injection point.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Hold the launcher after the checkpoint that licenses "on" has
+ # completed and before the state reaches the control file.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-after-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-after-checkpoint');
+
+ # The transition is complete as far as the running cluster is concerned.
+ test_checksum_state($node, 'on');
+
+ # The checkpoint has already recorded it, so the pending write below the
+ # injection point has nothing left to do.
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($version) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($version, '1',
+ 'the requested checkpoint recorded "on" in the control file');
+
+ # Crash before the control file write that follows the checkpoint.
+ $node->stop('immediate');
+
+ $node->start;
+
+ # Every page on disk carries a checksum and the checkpoint that
+ # flushed them completed, so the cluster has to come back verifying
+ # them.
+ test_checksum_state($node, 'on');
+
+ is($node->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '10000', 'relation readable after the crash');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ # The data directory must be one the offline tools accept.
+ $node->stop;
+ $node->command_ok([ 'pg_checksums', '--check', '-D', $node->data_dir ],
+ 'pg_checksums accepts the data directory');
+
+ # Leave the node running for the reset.
+ $node->start;
+}
+reset_scenario($node);
+
+# Scenario 5: a checkpoint racing SetDataChecksumsOn() between the insertion
+# of the XLOG2_CHECKSUMS("on") record and the shared memory update must not
+# insert an XLOG_CHECKPOINT_REDO record that follows the transition in WAL
+# order while still carrying "inprogress-on". Recovery resuming from such a
+# redo point never replays the preceding transition record, comes up in
+# "inprogress-on" and resolves the finished transition as interrupted.
+# XLogChecksums() closes the window by inserting the record and publishing
+# the new state under DataChecksumTransitionLock, which CreateCheckPoint()
+# takes around sampling the state and inserting the redo record. Here the
+# launcher is held between the two steps at the
+# datachecksums-on-before-publish injection point, and a concurrent
+# CHECKPOINT has to block on the lock instead of completing inside the
+# window.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # The datachecksums-on-before-publish point fires inside a critical
+ # section, where the wait machinery must not allocate. Waiting once
+ # at the datachecksums-enable-checksums-delay point, which the
+ # launcher runs outside the critical section, initializes it; see
+ # 050_redo_segment_missing.pl for the same recipe around
+ # create-checkpoint-run.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-enable-checksums-delay','wait');"
+ );
+
+ # Hold the launcher after the XLOG2_CHECKSUMS("on") record is in WAL but
+ # before the new state is published in shared memory.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-publish','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-enable-checksums-delay');
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-enable-checksums-delay');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-enable-checksums-delay');");
+
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-publish');
+
+ # The record is in WAL, the published state is still the old one.
+ test_checksum_state($node, 'inprogress-on');
+
+ # A concurrent checkpoint. It must block on DataChecksumTransitionLock
+ # before inserting its redo record rather than complete inside the window.
+ my $checkpointer = IPC::Run::start(
+ [ 'psql', '-XAtq', '-d', $node->connstr('postgres'), '-c', 'CHECKPOINT;' ],
+ '<' => '/dev/null',
+ '>' => '/dev/null',
+ '2>' => '/dev/null',
+ IPC::Run::timer($PostgreSQL::Test::Utils::timeout_default));
+
+ ok( $node->poll_query_until(
+ 'postgres',
+ "SELECT wait_event = 'DataChecksumTransition' "
+ . "FROM pg_stat_activity WHERE backend_type = 'checkpointer';"),
+ 'concurrent checkpoint blocks on DataChecksumTransitionLock');
+
+ # Release the launcher; the checkpoint then samples the published "on" and
+ # its redo record follows the transition record.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-publish');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-publish');");
+
+ $checkpointer->finish;
+
+ # Crash while the launcher's own checkpoint may still be in flight.
+ # Recovery resumes from the concurrent checkpoint's redo point, which
+ # now lies above the transition record and carries "on".
+ $node->stop('immediate');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition completed, so the cluster has to come back "on".
+ wait_for_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ $node->stop;
+}
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
new file mode 100644
index 00000000000..aaa35f8f7dd
--- /dev/null
+++ b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
@@ -0,0 +1,126 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A standby stopped after replaying an online enable, but before the
+# checkpoint record that follows it on the primary.
+#
+# The shutdown restartpoint has no new checkpoint record to work from
+# and is skipped, so the control file would keep the in-progress state
+# the transition passed through, even though replay left the node at
+# "on" and rewrote every page. pg_checksums reads that field and
+# refuses to run on an in-progress state, so the shutdown has to flush
+# and catch it up instead.
+#
+# The caught-up control file then carries the watermark of the "on"
+# record, which the second half of this test relies on: an offline
+# disable made while the standby is down must survive the re-replay of
+# that record on the next startup, or the change would be silently
+# reverted.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+wal_keep_size = 1GB
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Anchor the standby's last restartpoint here, so that nothing replayed
+# from now on gives the shutdown restartpoint a newer checkpoint record
+# to use.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Hold the enable right after the state change record, before the
+# checkpoint it requests once the transition is complete.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+enable_data_checksums($primary);
+$primary->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+wait_for_checksum_state($primary, 'on');
+
+# The state change record is already flushed and streams on its own;
+# no checkpoint record follows while the launcher is held.
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+$standby->stop('fast');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($state, '1',
+ 'shutdown catches the control file up with the replayed state');
+
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'pg_checksums verifies the stopped standby');
+
+# Disable checksums offline while the standby is down. The restart
+# below resumes replay under the last restartpoint, re-reading the "on"
+# record; the watermark persisted with the catch-up above marks it as
+# already applied, so the offline change must survive.
+$standby->checksum_disable_offline;
+
+# Let the enable finish on the primary.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-checkpoint');");
+$primary->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-checkpoint');");
+
+# The launcher exits only after the checkpoint it requests completes,
+# so a checkpoint record now follows the transition in WAL.
+$primary->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_catalog.pg_stat_activity "
+ . "WHERE backend_type = 'datachecksums launcher';");
+
+my $log_offset = -s $standby->logfile;
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+
+# The nodes legitimately diverged, which replayed checkpoints report.
+$standby->wait_for_log(
+ qr/data checksum state "off" of this node does not match the state "on" in the replayed WAL/,
+ $log_offset);
+
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/020_cascade_divergence.pl b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
new file mode 100644
index 00000000000..b4b474fa583
--- /dev/null
+++ b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
@@ -0,0 +1,115 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# An offline data checksum state change made on one node of a cascading setup
+# is reported all the way down the chain.
+#
+# CheckReplayedDataChecksumState() is reached from xlog_redo() when a
+# checkpoint record (XLOG_CHECKPOINT_SHUTDOWN or XLOG_CHECKPOINT_REDO) is
+# replayed after consistency has been reached. Those records are written by
+# the root primary only and are relayed verbatim by every intermediate
+# standby, so a cascaded standby that was changed offline learns about the
+# divergence too. An intermediate standby's own restartpoints are not
+# WAL-logged and cannot surface it.
+#
+# Note that this makes detection checkpoint-driven: an idle or read-only
+# primary surfaces the divergence no sooner than its next checkpoint, so up to
+# checkpoint_timeout may pass before the operator sees anything. Closing that
+# window would require comparing the state when a walreceiver connects.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# checkpoint_timeout is deliberately long so that the only checkpoint record in
+# this test is the one requested explicitly below.
+my $conf = qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+);
+
+my $primary = PostgreSQL::Test::Cluster->new('cascade_primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', $conf);
+$primary->start;
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,1000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('cascade_standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf('postgresql.conf', $conf);
+$standby->start;
+
+$standby->backup('backup');
+my $cascade = PostgreSQL::Test::Cluster->new('cascade_cascade');
+$cascade->init_from_backup($standby, 'backup', has_streaming => 1);
+$cascade->append_conf('postgresql.conf', $conf);
+$cascade->start;
+
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+test_checksum_state($cascade, 'off');
+
+# Change the cascaded standby offline, and only it.
+$cascade->stop;
+$cascade->command_ok([ 'pg_checksums', '--enable', '-D', $cascade->data_dir ],
+ 'pg_checksums enables checksums on the cascaded standby only');
+
+my $logstart = -s $cascade->logfile;
+$cascade->start;
+
+test_checksum_state($cascade, 'on');
+
+# Ordinary WAL traffic carries no state, so the cascaded standby keeps
+# streaming from a chain whose state is "off" without noticing.
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (1);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+is($cascade->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '1001', 'cascaded standby keeps streaming from a divergent chain');
+
+# Restartpoints on the intermediate standby are not WAL-logged, so this does
+# not surface anything on the cascaded standby either.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (2);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+my $log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+unlike(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'an intermediate restartpoint does not surface the divergence');
+
+# A checkpoint on the root primary does: the record is relayed down the whole
+# chain, so both the direct and the cascaded standby compare it against their
+# own state. wait_for_catchup() below waits for replay, not just receipt, so
+# the record has been through xlog_redo() on both nodes once it returns.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+$log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+like(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'the cascaded standby reports the divergence on the primary checkpoint');
+
+$cascade->stop;
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
new file mode 100644
index 00000000000..a6c8a1091ee
--- /dev/null
+++ b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
@@ -0,0 +1,193 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind across online data checksum transitions on both sides of a
+# divergence, with the two nodes trading roles between the scenarios.
+#
+# Scenario 1: after a switchover the new primary enables checksums
+# online, while the old primary restarts on its old timeline, advances
+# its WAL beyond the enable records and runs an online enable/disable
+# cycle of its own. The old primary ends "off" with a checksum
+# watermark numerically above every checksum record the new primary has
+# written. pg_rewind clamps the watermark it keeps to the divergence
+# point; without the clamp, replay on the rewound node would skip the
+# source's enable as already applied and stay "off" under an "on"
+# primary.
+#
+# Scenario 2: checksums are disabled again, and after another
+# switchover both nodes enable them online independently, so the
+# divergence checkpoint carries "off" while both control files say
+# "on". The rewind is allowed, the target keeps its own "on" state,
+# and replay re-walks the source's enable onto the already enabled
+# node.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+sub controldata_watermark
+{
+ my ($node) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $node->data_dir ]);
+ $stdout =~ /^Data checksum watermark:\s*([0-9A-F]+)\/([0-9A-F]+)$/m
+ or die "watermark missing from pg_controldata output";
+ return (hex($1) << 32) + hex($2);
+}
+
+# Wait until the standby has replayed the shutdown checkpoint of the
+# stopped primary, so that a subsequent promotion diverges after it and
+# the shutdown checkpoint becomes the last common checkpoint.
+sub wait_for_shutdown_checkpoint_replay
+{
+ my ($primary, $standby) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $primary->data_dir ]);
+ $stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+ my $shutdown_checkpoint = $1;
+
+ $standby->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_checkpoint'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+}
+
+# Old primary, checksums off. wal_log_hints is required by pg_rewind
+# on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+# Clean switchover to B; enable checksums online on it.
+$node_a->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_a, $node_b);
+$node_b->promote;
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+$node_b->stop('fast');
+my $watermark_b = controldata_watermark($node_b);
+$node_b->start;
+
+# Accidental restart of the old primary on the old timeline. Advance
+# its WAL beyond the enable watermark of B, then run an online enable
+# and disable cycle: the node ends "off" with a watermark above every
+# checksum record B has written.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t_pad AS SELECT generate_series(1,200000) AS a;");
+enable_data_checksums($node_a, wait => 'on');
+disable_data_checksums($node_a, wait => 'off');
+test_checksum_state($node_a, 'off');
+$node_a->stop('fast');
+
+my $watermark_a = controldata_watermark($node_a);
+die "test broken: target watermark not above the source's enable"
+ unless $watermark_a > $watermark_b;
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the promoted node');
+
+# Start the rewound node as a standby of B. Replay from the last
+# common checkpoint runs through B's online enable, which the target
+# never saw, so the rewound node must converge to "on".
+my $connstr_b = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr_b application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+# Scenario 2, reusing the pair with the roles reversed. Disable
+# checksums online so the next divergence point carries "off", and let
+# A replay the change.
+disable_data_checksums($node_b, wait => 'off');
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'off');
+
+# Clean switchover back to A; enable checksums online on it.
+$node_b->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_b, $node_a);
+$node_a->promote;
+enable_data_checksums($node_a, wait => 'on');
+test_checksum_state($node_a, 'on');
+my $source_enable_watermark = controldata_watermark($node_a);
+
+# The old primary restarts on its old timeline and enables checksums
+# online independently: both control files say "on", the divergence
+# checkpoint says "off".
+$node_b->start;
+$node_b->safe_psql('postgres', "INSERT INTO t_div VALUES (2);");
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+$node_b->stop('fast');
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_b->data_dir,
+ '--source-server' => $node_a->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled on both nodes');
+
+my $connstr_a = $node_a->connstr;
+$node_b->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_b->port]}
+primary_conninfo = '$connstr_a application_name=@{[$node_b->name]}'
+]);
+$node_b->set_standby_mode;
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_b, 'on');
+
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_b->stop('fast');
+is(controldata_watermark($node_b), $source_enable_watermark,
+ 'rewound node replayed the source checksum transition');
+command_ok([ 'pg_checksums', '--check', '-D', $node_b->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the twice-rewound node');
+
+done_testing();
--
2.55.0
[application/octet-stream] v12-0004-pg_combinebackup-Refuse-mixed-data-checksum-stat.patch (9.5K, ../../CAN4CZFNG57F8DY5NF=y=bdbCju6+_JxrBnUVYg5Xz3LOMCrSug@mail.gmail.com/5-v12-0004-pg_combinebackup-Refuse-mixed-data-checksum-stat.patch)
download | inline diff:
From 865ffc7aa437ef054a3cdff0afbd1244f894c0d0 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Tue, 25 Aug 2026 15:06:58 +0000
Subject: [PATCH v12 4/4] pg_combinebackup: Refuse mixed data checksum states
in a backup chain
check_control_files() detected a chain whose backups were taken under
different data checksum states, warned, and proceeded. The output
directory keeps the last backup's control file, so a full backup taken
with checksums off combined with an incremental taken after an offline
enable produces a cluster whose control file says "on" while most of
its blocks carry no checksums; it starts, and then every connection
dies on the first unchecksummed catalog page. An offline enable
between two backups of a chain is all it takes, since it rewrites
every page without logging anything, so the incremental backup does
not re-ship the pages.
Turn the warning into an error, matching what pg_rewind does for the
equivalent combinations. The check stays asymmetric on purpose: when
the last backup was taken with checksums off, stale checksums from an
earlier backup are never verified, and that chain remains usable.
---
doc/src/sgml/ref/pg_combinebackup.sgml | 19 +--
src/bin/pg_combinebackup/pg_combinebackup.c | 14 +-
src/test/modules/test_checksums/meson.build | 1 +
.../t/024_combinebackup_mixed.pl | 134 ++++++++++++++++++
4 files changed, 155 insertions(+), 13 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
diff --git a/doc/src/sgml/ref/pg_combinebackup.sgml b/doc/src/sgml/ref/pg_combinebackup.sgml
index 9a6d201e0b8..f5d7e177b36 100644
--- a/doc/src/sgml/ref/pg_combinebackup.sgml
+++ b/doc/src/sgml/ref/pg_combinebackup.sgml
@@ -306,18 +306,19 @@ PostgreSQL documentation
<para>
<literal>pg_combinebackup</literal> does not recompute page checksums when
- writing the output directory. Therefore, if any of the backups used for
- reconstruction were taken with checksums disabled, but the final backup was
- taken with checksums enabled, the resulting directory may contain pages
- with invalid checksums.
+ writing the output directory. It therefore refuses a chain in which some
+ of the backups used for reconstruction were taken with checksums disabled
+ but the final backup was taken with checksums enabled: the resulting
+ directory would contain pages with invalid checksums, and the cluster
+ would fail checksum verification as soon as it reads them. Take a new
+ full backup after enabling data checksums with
+ <xref linkend="app-pgchecksums"/>.
</para>
<para>
- To avoid this problem, taking a new full backup after changing the checksum
- state of the cluster using <xref linkend="app-pgchecksums"/> is
- recommended. Otherwise, you can disable and then optionally reenable
- checksums on the directory produced by <literal>pg_combinebackup</literal>
- in order to correct the problem.
+ The reverse case is accepted: when the final backup was taken with
+ checksums disabled, stale checksums copied from an earlier backup are
+ never verified.
</para>
</refsect1>
diff --git a/src/bin/pg_combinebackup/pg_combinebackup.c b/src/bin/pg_combinebackup/pg_combinebackup.c
index 86e2ee37c40..254a27b125b 100644
--- a/src/bin/pg_combinebackup/pg_combinebackup.c
+++ b/src/bin/pg_combinebackup/pg_combinebackup.c
@@ -670,13 +670,19 @@ check_control_files(int n_backups, char **backup_dirs)
pg_log_debug("system identifier is %" PRIu64, system_identifier);
/*
- * Warn the user if not all backups are in the same state with regards to
- * checksums.
+ * Reject a chain whose backups were not all in the same state with
+ * regards to checksums. The loop above only flags that when the last
+ * backup has checksums enabled: the blocks taken from the older backups
+ * have no checksums, and pg_combinebackup does not recompute them, so the
+ * combined cluster would fail verification as soon as it reads them. The
+ * other direction is harmless, since stale checksums are never verified
+ * while checksums are disabled.
*/
if (data_checksum_mismatch)
{
- pg_log_warning("only some backups have checksums enabled");
- pg_log_warning_hint("Disable, and optionally reenable, checksums on the output directory to avoid failures.");
+ pg_log_error("only some backups have checksums enabled");
+ pg_log_error_hint("Take a new full backup after changing the data checksum state with pg_checksums.");
+ exit(1);
}
return system_identifier;
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index 7d07c757052..7eccd5156b9 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -47,6 +47,7 @@ tests += {
't/021_rewind_divergent_transitions.pl',
't/022_rewind_state.pl',
't/023_rewind_standby_target.pl',
+ 't/024_combinebackup_mixed.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
new file mode 100644
index 00000000000..ce9dee9413b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
@@ -0,0 +1,134 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_combinebackup refuses a chain whose backups were taken under
+# different data checksum states when the final state has them enabled.
+#
+# The output directory keeps the last backup's control file, so a full
+# backup taken with checksums off plus an incremental taken after an
+# offline enable would produce a directory whose control file says "on"
+# while most of its blocks have no checksums. An offline enable between
+# the two backups is enough to get there: it rewrites every page but logs
+# nothing, so the incremental backup does not re-ship the pages.
+#
+# The reverse order stays allowed: with the final backup taken with
+# checksums off, stale checksums from an earlier backup are never verified.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+my $node = PostgreSQL::Test::Cluster->new('node');
+$node->init(
+ no_data_checksums => 1,
+ has_archiving => 1,
+ allows_streaming => 1);
+$node->append_conf('postgresql.conf', 'summarize_wal = on');
+$node->append_conf('postgresql.conf', 'autovacuum = off');
+$node->start;
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+test_checksum_state($node, 'off');
+
+$node->backup('full');
+
+# Offline enable: rewrites every page, logs nothing.
+$node->stop;
+system_or_bail('pg_checksums', '--enable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'on');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full/backup_manifest',
+ ],
+ 'incremental backup with checksums on');
+
+# Combining the chain must be refused: the result would say "on" while the
+# blocks inherited from the full backup have no checksums.
+command_fails_like(
+ [
+ 'pg_combinebackup',
+ $node->backup_dir . '/full',
+ $node->backup_dir . '/incr',
+ '--output' => $node->backup_dir . '/combined',
+ ],
+ qr/only some backups have checksums enabled/,
+ 'pg_combinebackup refuses a chain crossing an offline enable');
+
+# The other direction: full backup with checksums on, offline disable, then
+# an incremental. The last backup wins, so the combined cluster comes up off.
+$node->backup('full2');
+
+$node->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr2',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full2/backup_manifest',
+ ],
+ 'incremental backup with checksums off');
+
+my $restored = PostgreSQL::Test::Cluster->new('restored');
+$restored->init_from_backup(
+ $node, 'incr2',
+ combine_with_prior => ['full2'],
+ has_restoring => 1,
+ standby => 0);
+
+$restored->start;
+
+my ($rc, $stdout, $stderr) = $restored->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'combined backup accepts connections') or diag("stderr: $stderr");
+is($stdout, 'off', 'combined backup follows the last backup state');
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t3;');
+is($rc, 0, 'blocks from the incremental backup are readable')
+ or diag("stderr: $stderr");
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t1;');
+is($rc, 0, 'blocks inherited from the full backup are readable')
+ or diag("stderr: $stderr");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($restored->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures in the combined backup');
+
+$restored->stop('immediate');
+$node->stop;
+
+done_testing();
--
2.55.0
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-03 11:54 Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
parent: Zsolt Parragi <zsolt.parragi@percona.com>
0 siblings, 1 reply; 43+ messages in thread
From: Bertrand Drouvot @ 2026-09-03 11:54 UTC (permalink / raw)
To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; pgsql-hackers@lists.postgresql.org
Hi,
On Thu, Sep 03, 2026 at 12:06:42PM +0100, Zsolt Parragi wrote:
> Thanks!
>
> v12 addresses these, otherwise it is unchanged to compared 11.
Thanks! v12 LGTM.
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-03 23:08 Daniel Gustafsson <daniel@yesql.se>
parent: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 1 reply; 43+ messages in thread
From: Daniel Gustafsson @ 2026-09-03 23:08 UTC (permalink / raw)
To: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-hackers@lists.postgresql.org
> On 3 Sep 2026, at 13:54, Bertrand Drouvot <bertranddrouvot.pg@gmail.com> wrote:
> On Thu, Sep 03, 2026 at 12:06:42PM +0100, Zsolt Parragi wrote:
>> v12 addresses these, otherwise it is unchanged to compared 11.
>
> Thanks! v12 LGTM.
Thanks for review. I've attached a v13 where I've moved most of the new tests
under PG_TEST_EXTRA to keep test times down. I placed most tests under
'checksum' and 18, 21 and 23 under 'checksum_extended', but the exact split may
be tweaked further. Since the origin of this open item is missing test
coverage, I prefer to add all these tests even though they aren't executed
during normal testruns. There are at least one BF animal running the full
suite which ensures timely execution of the tests.
This concludes the only open item left (thus far). Being able to error
standbys out of mismatched clusters would be nice, and is a potential
development area for 20, but it's not a showstopper if we never add it IMHO.
My current plan is to commit this to master only either tomorrow or Monday
after staring at it a little bit more, to a) give it exposure in the buildfarm
before an eventual backpatching; b) allow time for the revert discussion. If
we decide to revert I prefer to avoid more v19 churn.
On the latter topic. Since this thread is about closing the final open item
with the deadline for a new beta looming, it seems like a good place to bring
up the discussion of reverting. I am admittedly not sure about what the best
course of action would be, and some input is greatly appreciated. AFAICT these
are the main reasons for a revert:
* Architectural concerns
* Existing unfixed bugs (open items)
* Forwards incompatible post feature-freeze fixes of stopcap nature
* Potential bugs we don't know about
* RMT mandating a revert
I'll try to address each below, apart from the latter one.
Re-reading the recent threads, I don't see any reports of architectural
concerns which can be acted on. Pointers to threads would be appreciated If
I've missed any.
There are no further open issues (after this one). The argument for not
applying this patchset and closing the item would be that it's too complicated
and invasive at this point in the cycle (requires a pg_control change for
example). I have a lot of sympathy for this. The 0001 diffstat might seem
terrifying due to the new tests, but there is also a nontrivial amount of code
added to xlog.c. 0002 and onwards are addressing existing bugs in offline
checksums, though they will look a tad different if online checksums is
reverted.
The number of postcommit fixes was highlighted as one area of concern. Since
number of commits is a poor measurement of anything I've tried to quantify it
by compiling all the postcommit fixes and categorized them (see attached .txt).
Of the 29 thus far there have been 12 bugfixes, 14 spelling/typo fixes, source
code improvements and test stabilization. 3 commits are improvements or
optimizations. Whether or not those numbers are more interesting than 29, or
if either makes a case for reverting, I don't know. What I do know from
looking them over is that none of them were solved any different than what they
would have been if they were only in master.
Another case for a revert is if there is a general unspecified uneasiness over
the feature or its readiness among the maintainers and/or the RMT. And that's
totally fair. It's by far the most complicated and ambitious thing I've ever
built for Postgres. I have no interest in causing stress or worry in my fellow
postgres hackers, especially now, so if there is concensus on bad gut-feeling
then that's also a fair case for reverting. However, if we decide to revert
over an unspecified reason, I think we need to couple that with defining what
needs to change/happen over the v20 cycle to put that uneasiness to rest.
Simply leaving it in master without actionable technical concerns raised
against the code or the architecture for a year won't change anything.
Once/if this patchset lands, left on my TODO is to look at potentially dead
code around setting delayChkptFlags which Tomas Vondra identified, and keep
trying to break the code with the help of various LLM models and their ability
to see gaps in testing. I have posted a revert commit (to the original thread)
to give anyone interested a chance to see what it would look like, and what I
propose leaving behind in v19.
Any thoughts/concerns?
--
Daniel Gustafsson
Bugs
----
* 20260406: d771b0a907e: Handle checksumworker startup wait race
Handle the race condition that data checksums are fully enabled and the worker
finishes before the launcher starts waiting for it.
* 20260430: 25b922ec582: Fix invalid checksum state transition in checkpoints
There was a race condition around updating the controlfile leading to a backend
starting at exactly the wrong time reading the previous state. A related issue
was if multiple checksum state changes happened between a backend initializing
procsignals and finishing XLOG startup. Somewhat hard to hit since it requires
a backend magnitudes slower than other background workers in the same cluster.
* 20260430: 8fb8ded8895: Handle data_checksum state changes during launcher_exit
Silly bug, the state transition for reverting from inprogress-on to off via
inprogress-off in one error path was missing the state machince allowed state
table.
* 20260430: b120358c612: Prevent pg_enable/disable_data_checksums() on standby
The SQL functions for disabling and enabling checksums could be executed on a
hot standby.
* 20260506: 9a39056c418: Apply data-checksum worker throttling parameters
The API for for cost parameters had changed between when I wrote the code
(years ago) and now and I had missed that, so the cost settings were not
properly applied.
* 20260529: 5fee7cab1b8: Fix checksum state transition during promotion
When a primary crashed during checksum enabling, the standby didn't emit the
procsignalbarrier during promotion to revert the cluster state to off.
* 20260804: 01805b7d16b: Do not reuse rd_smgr in fork loop when enabling data checksums
The worker incorrectly accessed rd_smgr which can be set to NULL at a relcache
invalidation.
* 20260817: 3a18526e8d6: Make data checksums launcher cancel its worker at SIGINT
The launcher waited for the worker to complete before terminating at SIGINT
which cause pointless work and can take some time. Fix to make it faster by
having the launcher terminate and worker at SIGINT and then exit itself.
* 20260430: 1df361e3d82: Improve database detection logic in datachecksumsworker
* 20260804: 343d98c3601: Don't skip invalid databases when enabling data checksums
Commit 1df361e3d82 attempted to improve the logic for handling invalid
databases, later fixed in 343d98c3601 to handle databases which have had a DROP
DATABASE operation crash.
* 20260828: e469e4784ea: Handle invalid and dropped databases during checksum enable
Detect invalid databases earlier, and also detect when a database is dropped
while checksums are being enabled in that very database, previously this case
resulted in a process error with the state rolled back to off.
* 20260818: 0907112d388: basebackup: do not verify checksums on pages from before enabling
A base backup which started before a checksum state transition, or when
checksums are disabled and re-enabled while the backup is running, would
produce false negatives.
Improvements and Optimizations
------------------------------
* 20260430: bf25e5571b3: Improve handling of concurrent checksum requests
Made the logic for handling concurrent invocations of enabling or disabling
checksums not require creating a new launcher for detection, which reduce
cluster process usage.
* 20260506: 2018bd61679: Skip WAL for unlogged main fork during online checksum enable
Pages for unlogged relations were not exempt from processing which caused
needless WAL traffic. All pages were still properly checksummed, it "just"
generated too much WAL traffic.
* 20260818: 397f0fd06ed: Add data_page_checksum_version to pg_control_checkpoint
pg_controldata shows the checksum version from the checkpoint but system view
pg_control_checkpoint omitted it.
Minor Source Code changes
-------------------------
* 20260406: b3a37ffbc5b: Use PG_DATA_CHECKSUM_OFF instead of hardcoded value
Replace 0 with the PG_DATA_CHECKSUM_OFF label in the code. Readability
improvement over using the hardcoded zero which has been used since forever.
* 20260529: 0ca1b301059: Use correct datatype for PID
Use pid_t instead of int.
* 20260817: abac86c7a27: Reorder function prototypes to match definition order
Clearly not required, but there was an ask to do it in HEAD and backpatching
seemed a cheap insurance against backpatching conflicts.
Pre-existing bugs
-----------------
* 20260818: aaf8b9989f7: Record initial state of data checksums in controlfile
Fixes a longstanding bug in offline checksums where pg_control_init is defined
to list the initial cluster state, but pg_checksums just overwrites and
effectively made it show the current state.
Test Stabilization
-------------------
* 20260406: 07009121c23: Test stabilization for online checksums
* 20260901: 5313fc415d9: Wait for checksum state transition in test
Documentation, Error messages and Code comments etc
---------------------------------------------------
* 20260408: b364828f825: doc: Fix data_checksums data type
* 20260430: 381d19da153: Typo and spelling fixups for online checksums
* 20260529: cd857dec0e0: Improve comments in online checksums code
* 20260529: 5ab239c9a90: Constistent naming for datacheckusms processes
* 20260605: 4ae3e98c02c: doc: Mention online checksum enabling in pg_checksums docs
* 20260605: e5e1f6dc795: Reword activity message to avoid truncation
* 20260618: 8d22f523245: Fix comments on data checksum cost settings
* 20260716: e3a27cad462: doc: Fix link text for data checksums
* 20260801: 602f19c84ca: doc: Fix glossary entry for data checksums workers
Attachments:
[application/octet-stream] v13-0001-Do-not-adopt-data-checksum-state-from-another-no.patch (137.1K, ../../97139167-3B7B-4B00-B623-2DD5C0656DDB@yesql.se/2-v13-0001-Do-not-adopt-data-checksum-state-from-another-no.patch)
download | inline diff:
From 35e9ffe0495c60d648b967bed1249267a0d79e4b Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 3 Sep 2026 23:29:17 +0200
Subject: [PATCH v13 1/4] Do not adopt data checksum state from another node
during replay
Offline pg_checksums changes are local to one node, but replay adopted
the data checksum state carried by checkpoint records unconditionally.
After an offline enable on the primary, a standby whose pages were
never rewritten started verifying checksums it does not have; after an
offline change on a standby, the next replayed checkpoint silently
reverted it.
To fix, make the control file track this node's state alone, and have
replay cross-check the replayed state against it instead of adopting
it, warning once per divergent value and reporting when the states
agree again. pg_control gains a watermark, normally the end LSN of
the newest XLOG2_CHECKSUMS record the node has written or applied, so
that replay can skip transition records whose effect the control file
already contains, and a flag marking a state last written by
pg_checksums, which recovery must never overwrite with a replayed one.
Since the control file may only claim "on" once every page on disk
carries a checksum, persisting the state is tied to flushes:
XLOG2_CHECKSUMS replay persists every state but "on" immediately and
leaves "on" to the next restartpoint, checkpoints and restartpoints
only persist a state their flush ran under from beginning to end, and
transitions publish their state under the new
DataChecksumTransitionLock so that the states carried by WAL records
match their WAL order. The comments in xlog.c spell out the
individual rules.
Recovery from a base backup may be an exception to not adopting: its
control file was copied at an arbitrary moment, so the state carried
by the starting checkpoint is the one the WAL from there on was
written under. That holds for a backup taken from a primary, and only
when the copied state is not node local and its watermark is below the
starting checkpoint; a backup taken from a standby keeps the copied
state. pg_rewind keeps the target's own state and clamps the
watermark to the divergence point, since a watermark set on the
target's abandoned fork could numerically cover transition records the
source wrote after the divergence.
Bump PG_CONTROL_VERSION.
Document the offline procedure for replication setups, the lockstep
one: stop all nodes, run pg_checksums on each of them, and only then
restart. An offline change writes no WAL and has no ordering against
WAL a node has not replayed yet, so a node must not stop before
replaying all WAL of its upstream, or the change is overridden on
restart.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Author: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reported-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Discussion: https://postgr.es/m/anwm6UPxoVS41QA2@bdtpg
---
doc/src/sgml/ref/pg_checksums.sgml | 87 ++-
doc/src/sgml/wal.sgml | 23 +
src/backend/access/transam/xlog.c | 571 ++++++++++++++++--
src/backend/postmaster/datachecksum_state.c | 26 +-
.../utils/activity/wait_event_names.txt | 1 +
src/bin/pg_checksums/pg_checksums.c | 10 +
src/bin/pg_controldata/pg_controldata.c | 4 +
src/bin/pg_resetwal/pg_resetwal.c | 7 +
src/bin/pg_rewind/pg_rewind.c | 35 +-
src/bin/pg_upgrade/controldata.c | 2 +-
src/include/catalog/pg_control.h | 29 +-
src/include/storage/lwlocklist.h | 1 +
src/test/modules/test_checksums/Makefile | 2 +-
src/test/modules/test_checksums/meson.build | 10 +
.../test_checksums/t/012_offline_standby.pl | 304 ++++++++++
.../modules/test_checksums/t/013_rewind.pl | 201 ++++++
.../modules/test_checksums/t/014_lockstep.pl | 182 ++++++
.../t/015_standby_crash_after_disable.pl | 137 +++++
.../t/016_promote_enable_crash.pl | 146 +++++
.../test_checksums/t/017_restartpoint_race.pl | 151 +++++
.../t/018_enable_crash_windows.pl | 508 ++++++++++++++++
.../t/019_standby_shutdown_catchup.pl | 138 +++++
.../t/020_cascade_divergence.pl | 127 ++++
.../t/021_rewind_divergent_transitions.pl | 205 +++++++
24 files changed, 2823 insertions(+), 84 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/012_offline_standby.pl
create mode 100644 src/test/modules/test_checksums/t/013_rewind.pl
create mode 100644 src/test/modules/test_checksums/t/014_lockstep.pl
create mode 100644 src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
create mode 100644 src/test/modules/test_checksums/t/016_promote_enable_crash.pl
create mode 100644 src/test/modules/test_checksums/t/017_restartpoint_race.pl
create mode 100644 src/test/modules/test_checksums/t/018_enable_crash_windows.pl
create mode 100644 src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
create mode 100644 src/test/modules/test_checksums/t/020_cascade_divergence.pl
create mode 100644 src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index ae66fad3f0f..2d2057a8c19 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -242,15 +242,79 @@ PostgreSQL documentation
data directory must not be started or else data loss may occur.
</para>
<para>
- When using a replication setup with tools which perform direct copies
- of relation file blocks (for example <xref linkend="app-pgrewind"/>),
- enabling or disabling checksums can lead to page corruptions in the
- shape of incorrect checksums if the operation is not done consistently
- across all nodes. When enabling or disabling checksums in a replication
- setup, it is thus recommended to stop all the clusters before switching
- them all consistently. Destroying all standbys, performing the operation
- on the primary and finally recreating the standbys from scratch is also
- safe.
+ Enabling or disabling checksums with
+ <application>pg_checksums</application> changes only the local data
+ directory; the new state is not replicated to any other node. In a
+ replication setup the same change must be applied to every node:
+ </para>
+
+ <procedure>
+ <step id="shut-down-nodes">
+ <title>Shut down all nodes</title>
+ <para>
+ All nodes participating in the replication must be stopped with a
+ clean shutdown; <application>pg_checksums</application> refuses to run
+ on a data directory left behind by an immediate shutdown. Before
+ stopping a standby, make sure it has replayed all WAL of the primary:
+ stop the primary first, read its <quote>Latest checkpoint
+ location</quote> with <xref linkend="app-pgcontroldata"/>, and check
+ that <function>pg_last_wal_replay_lsn()</function> on the standby has
+ advanced past it. Comparing the replay position with
+ <function>pg_last_wal_receive_lsn()</function> is not enough, as it
+ only shows that the WAL the standby received has been replayed.
+ </para>
+ </step>
+
+ <step>
+ <title>Enable or disable data checksums on each node</title>
+ <para>
+ Run <application>pg_checksums</application> on the data directory of
+ each node in the replication setup. Nodes can be processed in
+ parallel while they are shut down. Processing must complete
+ successfully on all nodes before continuing.
+ </para>
+ </step>
+
+ <step>
+ <title>Restart all nodes</title>
+ <para>
+ Start the nodes normally, verify that
+ <xref linkend="guc-data-checksums"/> matches on all of them, and
+ monitor the logs of the standbys for data checksum state mismatch
+ warnings.
+ </para>
+ </step>
+ </procedure>
+
+ <para>
+ The replay requirement in <xref linkend="shut-down-nodes"/>exists
+ because an offline change is recorded only in the control file and
+ has no defined ordering against WAL the node has not replayed yet; see
+ <xref linkend="checksums-offline-enable-disable"/>. A node stopped
+ before replaying an online checksum state change applies that change
+ when it is restarted, overriding the offline change, and the states of
+ the nodes silently diverge until a later checkpoint record triggers
+ the warning described below. Because of this it is best not to mix
+ the two mechanisms: change the state of a replication setup either
+ with the offline procedure above or with an online transition, and
+ make sure the previous change has reached every node before starting
+ the next one.
+ </para>
+ <para>
+ If the change is applied inconsistently, each node keeps its own state,
+ and a standby logs a warning when the state recorded in the replayed WAL
+ differs from its own. The same warning can appear transiently while a
+ standby catches up over WAL written before a consistent change; it stops
+ once a checkpoint record carrying the new state has been replayed.
+ </para>
+ <para>
+ A standby whose data directory was never checksummed must not have
+ checksums enabled by catching up this way. Converge the cluster by
+ running <application>pg_checksums</application> on it while stopped as
+ outlined above, by enabling checksums online, or by recreating it from
+ a base backup. Note
+ that an online enable only starts from a primary whose checksums are off,
+ so if they are already enabled there, disable them online first.
</para>
<para>
If <application>pg_checksums</application> is aborted or killed while
@@ -258,6 +322,11 @@ PostgreSQL documentation
remains unchanged, and <application>pg_checksums</application> can be
re-run to perform the same operation.
</para>
+ <para>
+ Tools that copy relation file blocks directly between nodes, such as
+ <xref linkend="app-pgrewind"/>, require all nodes to be in
+ the same data checksum state, else there is risk for data corruption.
+ </para>
<para>
The target cluster must have the same major version as
<application>pg_checksums</application>.
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index ec62d17fbcc..dbbfdae5e4f 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -317,6 +317,29 @@
verify checksums, on an offline cluster.
</para>
+ <para>
+ An offline change provides durability differently from an
+ <link linkend="checksums-online-enable-disable">online change</link>.
+ An online transition is WAL-logged: it is ordered against all other
+ WAL records, it is replayed after a crash, and it propagates to
+ standbys. An offline change is recorded only in the cluster's
+ control file: it writes no WAL, it is invisible to replication, and
+ it has no defined ordering against WAL the node has not replayed
+ yet. When a node later replays WAL that contains an online checksum
+ state change, that change takes effect on the node even if it was
+ written before the offline change was made.
+ </para>
+
+ <para>
+ An offline change only affects the data directory it is run on; the
+ new state does not propagate over replication. In a replication setup
+ the same change must be applied to all nodes while all of them are
+ stopped, as described in <xref linkend="app-pgchecksums"/>. A standby
+ whose state diverges logs a warning but keeps its local setting. The
+ mismatch persists until the states are brought together again, with
+ the offline procedure or with an online transition; do this promptly.
+ </para>
+
</sect2>
<sect2 id="checksums-online-enable-disable" xreflabel="Online Enabling of Checksums">
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index 2e3f177100b..190043720d5 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -556,14 +556,22 @@ typedef struct XLogCtlData
*/
XLogRecPtr lastFpwDisableRecPtr;
- /* last data_checksum_version we've seen */
+ /* current data checksum state of this node */
uint32 data_checksum_version;
+ /*
+ * Copies of control file fields with the same names, see pg_control.h for
+ * an in-depth description of these fields. Must be updated together with
+ * data_checksum_version under info_lck.
+ */
+ XLogRecPtr data_checksum_lsn;
+ bool data_checksum_is_local;
+
slock_t info_lck; /* locks shared variables shown above */
/*
* lastChecksumChangeRecPtr points to the end of the last XLOG2_CHECKSUMS
- * record inserted or replayed, i.e. the last change of
+ * record inserted or replayed which corresponds to the last change of
* data_checksum_version. InvalidXLogRecPtr if the state hasn't changed
* since the server started.
*/
@@ -690,6 +698,20 @@ static ChecksumStateType LocalDataChecksumState = 0;
*/
int data_checksums = 0;
+/*
+ * Whether replay of the next checkpoint-family record must adopt the data
+ * checksum state it carries. Set when recovery starts from a base backup,
+ * where the state at the redo point takes precedence over the control file
+ * copied with the backup later.
+ */
+static bool adoptChecksumStateFromNextCheckpoint = false;
+
+/*
+ * Sentinel for the data checksum mismatch warning tracking in
+ * CheckReplayedDataChecksumState(): no warning is outstanding.
+ */
+#define NO_WARNING_ISSUED PG_UINT32_MAX
+
/* For WALInsertLockAcquire/Release functions */
static int MyLockNo = 0;
static bool holdingAllLocks = false;
@@ -730,6 +752,8 @@ static void ValidateXLOGDirectoryStructure(void);
static void CleanupBackupHistory(void);
static void UpdateMinRecoveryPoint(XLogRecPtr lsn, bool force);
static bool PerformRecoveryXLogAction(void);
+static void CheckReplayedDataChecksumState(uint32 replayed_version);
+static void AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn);
static void InitControlFile(uint64 sysidentifier, uint32 data_checksum_version);
static void WriteControlFile(void);
static void ReadControlFile(void);
@@ -757,7 +781,7 @@ static void WALInsertLockAcquireExclusive(void);
static void WALInsertLockRelease(void);
static void WALInsertLockUpdateInsertingAt(XLogRecPtr insertingAt);
-static void XLogChecksums(uint32 new_type);
+static XLogRecPtr XLogChecksums(uint32 new_type);
/*
* Insert an XLOG record represented by an already-constructed chain of data
@@ -4292,6 +4316,8 @@ InitControlFile(uint64 sysidentifier, uint32 data_checksum_version)
ControlFile->track_commit_timestamp = track_commit_timestamp;
ControlFile->data_checksum_version = data_checksum_version;
ControlFile->data_checksum_version_init = data_checksum_version;
+ ControlFile->data_checksum_is_local = false;
+ ControlFile->data_checksum_lsn = InvalidXLogRecPtr;
/*
* Set the data_checksum_version value into XLogCtl, which is where all
@@ -4774,6 +4800,7 @@ void
SetDataChecksumsOnInProgress(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
/*
* The state transition is performed in a critical section with
@@ -4782,14 +4809,12 @@ SetDataChecksumsOnInProgress(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4827,6 +4852,8 @@ void
SetDataChecksumsOn(void)
{
uint64 barrier;
+ bool persist;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4847,23 +4874,11 @@ SetDataChecksumsOn(void)
SpinLockRelease(&XLogCtl->info_lck);
INJECTION_POINT("datachecksums-enable-checksums-delay", NULL);
+ INJECTION_POINT_LOAD("datachecksums-on-before-publish");
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_VERSION);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- SpinLockRelease(&XLogCtl->info_lck);
-
- /*
- * Update the controlfile before waiting since if we have an immediate
- * shutdown while waiting we want to come back up with checksums enabled.
- */
- LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- UpdateControlFile();
- LWLockRelease(ControlFileLock);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_VERSION);
barrier = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_CHECKSUM_ON);
@@ -4873,6 +4888,41 @@ SetDataChecksumsOn(void)
INJECTION_POINT("datachecksums-on-before-checkpoint", NULL);
RequestCheckpoint(CHECKPOINT_FORCE | CHECKPOINT_WAIT | CHECKPOINT_FAST);
+
+ INJECTION_POINT("datachecksums-on-after-checkpoint", NULL);
+
+ /*
+ * Persist "on" only now that the checkpoint has flushed the pages the
+ * transition rewrote. Crash recovery initializes verification from the
+ * control file but resumes from a checkpoint that can predate the
+ * transition, so an "on" persisted earlier would verify pages during
+ * replay whose rewrite never reached the disk: the pages found by the
+ * data checksums worker in shared buffers are not written back by its
+ * ring buffer.
+ *
+ * The checkpoint above normally persists the state itself; this covers
+ * the case where it started before the record written above and left the
+ * field alone. Crashing before this point is safe, as replay then
+ * re-establishes "on" from the full page images of the rewrite. Skip the
+ * write if the state moved on meanwhile, since whatever moved it persists
+ * its own. Compare the watermark rather than the state: a state
+ * comparison could not tell our transition from a later round trip back
+ * to "on".
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ persist = (XLogCtl->data_checksum_lsn == recptr);
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (persist)
+ {
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+ }
+
WaitForProcSignalBarrier(barrier);
}
@@ -4893,6 +4943,7 @@ void
SetDataChecksumsOff(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4918,14 +4969,12 @@ SetDataChecksumsOff(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4956,14 +5005,12 @@ SetDataChecksumsOff(void)
/* Ensure that we don't incur a checkpoint during disabling checksums */
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -5001,6 +5048,138 @@ SetLocalDataChecksumState(uint32 data_checksum_version)
data_checksums = data_checksum_version;
}
+/*
+ * CheckReplayedDataChecksumState
+ * Cross-check the data checksum state carried by a replayed checkpoint
+ * record against the state of this node.
+ *
+ * The state in the record belongs to whichever node wrote the WAL, and must
+ * not be adopted: an offline state change made with pg_checksums on one node
+ * of a replication set generates no WAL, and adopting would leak it into the
+ * other nodes through replay. Backup label recovery is the one exception,
+ * see AdoptReplayedDataChecksumState(). XLOG_CHECKPOINT_ONLINE needs no call
+ * here, as an online checkpoint's state already traveled in the preceding
+ * XLOG_CHECKPOINT_REDO record.
+ *
+ * Only archive recovery can see a lasting mismatch, as only there can the WAL
+ * and the control file come from different nodes or different times. In
+ * crash recovery a mismatch means replay resumed from a restartpoint
+ * predating an already-applied XLOG2_CHECKSUMS record, and replaying forward
+ * re-establishes the same state.
+ */
+static void
+CheckReplayedDataChecksumState(uint32 replayed_version)
+{
+ /*
+ * Warn once per remote value, so a lasting mismatch does not flood the
+ * log. Matching states re-arm the warning. Backend-local state is
+ * enough: replay only runs in the startup process, and restarting it
+ * re-arms as well.
+ */
+ static uint32 last_warned_version = NO_WARNING_ISSUED;
+ uint32 local_version;
+
+ if (!ArchiveRecoveryRequested)
+ return;
+
+ /*
+ * Re-replayed WAL below the consistency point was already cross-checked
+ * before minRecoveryPoint was last persisted, and the persisted state can
+ * legitimately be newer than what checkpoint records there carry:
+ * XLOG2_CHECKSUMS replay persists most states ahead of the restartpoint
+ * horizon. In particular the checkpoint record recovery restarts from is
+ * such a re-replay.
+ */
+ if (!reachedConsistency)
+ return;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ local_version = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (replayed_version == local_version)
+ {
+ /*
+ * Report convergence if this process warned before. Nothing else
+ * tells the operator that running pg_checksums on the other nodes, or
+ * a rebuild, took effect.
+ */
+ if (last_warned_version != NO_WARNING_ISSUED)
+ ereport(LOG,
+ errmsg("data checksum state \"%s\" of this node now agrees with the replayed WAL",
+ get_checksum_state_string(local_version)));
+
+ last_warned_version = NO_WARNING_ISSUED;
+ return;
+ }
+
+ /* the nodes legitimately differ while an online transition runs */
+ if (replayed_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ replayed_version == PG_DATA_CHECKSUM_INPROGRESS_OFF ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ return;
+
+ if (replayed_version == last_warned_version)
+ return;
+ last_warned_version = replayed_version;
+
+ ereport(WARNING,
+ errmsg("data checksum state \"%s\" of this node does not match the state \"%s\" in the replayed WAL",
+ get_checksum_state_string(local_version),
+ get_checksum_state_string(replayed_version)),
+ errdetail("The data checksum state may have been changed with pg_checksums on another node."),
+ errhint("Apply the same change with pg_checksums on the primary and all standby servers, or rebuild this server from a base backup."));
+}
+
+/*
+ * AdoptReplayedDataChecksumState
+ * Adopt the data checksum state at the redo point of backup label
+ * recovery.
+ *
+ * The state is persisted immediately so that a crash before the first
+ * restartpoint does not resurrect the state copied with the backup; a crash
+ * at this point restarts from the same redo point, so the control file does
+ * not run ahead of the replay position. If the value is unchanged the
+ * control file already carries it, so both the barrier and the persist are
+ * skipped.
+ *
+ * lsn is the location of the checkpoint-family record the state was taken
+ * from and becomes the new watermark: the adopted state covers everything
+ * below the redo point, which replay never revisits. For the same reason
+ * the old watermark can stay when the value is unchanged: the records
+ * between the two positions are never replayed, so nothing depends on
+ * which one is recorded.
+ */
+static void
+AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn)
+{
+ bool changed = false;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (XLogCtl->data_checksum_version != new_version)
+ {
+ XLogCtl->data_checksum_version = new_version;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(new_version);
+ changed = true;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (!changed)
+ return;
+
+ EmitAndWaitDataChecksumsBarrier(new_version);
+
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = new_version;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+}
+
/* guc hook */
const char *
show_data_checksums(void)
@@ -5454,6 +5633,8 @@ XLOGShmemInit(void *arg)
/* Use the checksum info from control file */
XLogCtl->data_checksum_version = ControlFile->data_checksum_version;
+ XLogCtl->data_checksum_lsn = ControlFile->data_checksum_lsn;
+ XLogCtl->data_checksum_is_local = ControlFile->data_checksum_is_local;
SetLocalDataChecksumState(XLogCtl->data_checksum_version);
SpinLockInit(&XLogCtl->Insert.insertpos_lck);
@@ -6031,6 +6212,51 @@ StartupXLOG(void)
SetCommitTsLimit(checkPoint.oldestCommitTsXid,
checkPoint.newestCommitTsXid);
+ /*
+ * When recovery starts from a base backup, the control file was copied at
+ * an arbitrary moment and its data checksum state may differ from the
+ * state at the redo point, which is what the WAL from there on was
+ * written under. Adopt the state of the starting checkpoint: a shutdown
+ * checkpoint is not replayed, so take it from the record read above; the
+ * redo point of an online checkpoint is its CHECKPOINT_REDO record, so
+ * let the replay of that record adopt it. Check backupStartPoint in
+ * addition to the label: on a crash restart during backup recovery the
+ * label file is already renamed away, but the start point persists until
+ * the backup end record.
+ *
+ * Not for a base backup taken from a standby, though. Its starting
+ * checkpoint is the standby's last restartpoint, a record written by the
+ * upstream primary, whose state is not the one the copied files were
+ * written under. The copied control file is already correct: a standby
+ * persists its state only at restartpoint horizons and never claims more
+ * than what reached disk. Such backups are recognized by backupEndPoint
+ * together with backupEndRequired; backupEndPoint is only set for "BACKUP
+ * FROM: standby" labels and persists across a crash restart. pg_rewind
+ * writes a standby label as well, but no backupEndPoint, and its recovery
+ * keeps adopting: the control file it installs carries the target's own
+ * checksum state, which can lag the redo point of the last common
+ * checkpoint the same way a restartpoint horizon can.
+ *
+ * Never adopt over a state the control file's watermark or local flag
+ * marks as newer than the starting checkpoint. A pg_checksums change is
+ * local to the node and generates no WAL, so nothing in the replayed WAL
+ * could ever restore it once overwritten; and a watermark above the redo
+ * point means the control file already contains the effect of every
+ * transition record up to there.
+ */
+ if ((haveBackupLabel || XLogRecPtrIsValid(ControlFile->backupStartPoint)) &&
+ !(XLogRecPtrIsValid(ControlFile->backupEndPoint) &&
+ ControlFile->backupEndRequired) &&
+ !ControlFile->data_checksum_is_local &&
+ checkPoint.redo > ControlFile->data_checksum_lsn)
+ {
+ if (wasShutdown)
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ checkPoint.redo);
+ else
+ adoptChecksumStateFromNextCheckpoint = true;
+ }
+
/*
* Clear out any old relcache cache files. This is *necessary* if we do
* any WAL replay, since that would probably result in the cache files
@@ -6634,11 +6860,7 @@ StartupXLOG(void)
if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
ereport(WARNING,
@@ -6655,11 +6877,7 @@ StartupXLOG(void)
else if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
}
@@ -6684,6 +6902,8 @@ StartupXLOG(void)
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
XLogCtl->SharedRecoveryState = RECOVERY_STATE_DONE;
SpinLockRelease(&XLogCtl->info_lck);
@@ -6815,6 +7035,23 @@ static bool
PerformRecoveryXLogAction(void)
{
bool promoted = false;
+ bool flushForChecksums;
+ uint32 checksum_state;
+
+ /*
+ * The end-of-recovery record persists the data checksum state without
+ * flushing the buffer pool, but the control file may only claim "on" once
+ * every page on disk carries a checksum. If replay entered that state
+ * without a restartpoint following it, the pages rewritten by the
+ * transition are still only in the buffer pool, so take the full
+ * checkpoint below instead of the lightweight record.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ flushForChecksums = (checksum_state == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile->data_checksum_version != checksum_state);
/*
* Perform a checkpoint to update all our recovery activity to disk.
@@ -6830,7 +7067,7 @@ PerformRecoveryXLogAction(void)
* fully out of recovery mode and already accepting queries.
*/
if (ArchiveRecoveryRequested && IsUnderPostmaster &&
- PromoteIsTriggered())
+ PromoteIsTriggered() && !flushForChecksums)
{
promoted = true;
@@ -7437,6 +7674,7 @@ CreateCheckPoint(int flags)
uint32 freespace;
XLogRecPtr PriorRedoPtr;
XLogRecPtr last_important_lsn;
+ XLogRecPtr checksumLsn;
VirtualTransactionId *vxids;
int nvxids;
int oldXLogAllowed = 0;
@@ -7549,11 +7787,14 @@ CreateCheckPoint(int flags)
checkPoint.wal_level = wal_level;
/*
- * Get the current data_checksum_version value from xlogctl, valid at the
- * time of the checkpoint.
+ * Get the current data_checksum_version value from xlogctl. This is
+ * final only for a shutdown checkpoint, where no concurrent transition is
+ * possible; an online checkpoint resamples it together with the redo
+ * record below.
*/
SpinLockAcquire(&XLogCtl->info_lck);
checkPoint.dataChecksumState = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
if (shutdown)
@@ -7610,10 +7851,21 @@ CreateCheckPoint(int flags)
{
xl_checkpoint_redo redo_rec;
+ /*
+ * Sample the data checksum state and insert the redo record under
+ * DataChecksumTransitionLock, so that a concurrent transition cannot
+ * insert its XLOG2_CHECKSUMS record between the sampling and the
+ * insertion below. Without this, the redo record could follow the
+ * transition record in WAL while carrying the pre-transition state,
+ * and recovery resuming here would never learn about the transition.
+ * See XLogChecksums().
+ */
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
WALInsertLockAcquire();
redo_rec.wal_level = wal_level;
SpinLockAcquire(&XLogCtl->info_lck);
redo_rec.data_checksum_version = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
WALInsertLockRelease();
@@ -7621,6 +7873,14 @@ CreateCheckPoint(int flags)
XLogBeginInsert();
XLogRegisterData(&redo_rec, sizeof(xl_checkpoint_redo));
(void) XLogInsert(RM_XLOG_ID, XLOG_CHECKPOINT_REDO);
+ LWLockRelease(DataChecksumTransitionLock);
+
+ /*
+ * The checkpoint record must carry the same state as the redo record
+ * just inserted: the sample taken before redo determination can be
+ * stale by now, and the pair would otherwise disagree.
+ */
+ checkPoint.dataChecksumState = redo_rec.data_checksum_version;
/*
* XLogInsertRecord will have updated XLogCtl->Insert.RedoRecPtr in
@@ -7824,6 +8084,41 @@ CreateCheckPoint(int flags)
ControlFile->minRecoveryPoint = InvalidXLogRecPtr;
ControlFile->minRecoveryPointTLI = 0;
+ /*
+ * Persist the data checksum state this node runs under into the control
+ * file. Only the top-level field tracks this node, checkPointCopy is a
+ * historical record used to resume replay.
+ *
+ * checkPoint.dataChecksumState was sampled while holding the
+ * DataChecksumTransitionLock together with the redo record, so it is the
+ * state in effect at the redo point. If it was "on", the XLOG2_CHECKSUMS
+ * record announcing that precedes the redo point and every page the
+ * transition rewrote was dirtied before it, so CheckPointGuts() has just
+ * written all of them out. Recording the state here is what keeps a
+ * finished transition from being resolved as interrupted when this
+ * checkpoint is the one crash recovery resumes from: replay never sees
+ * the record announcing it.
+ *
+ * Persist it only if the state did not change while the flush was in
+ * progress. If it changed in between, the pages written out straddle two
+ * states, and the newer one could claim checksums that pages already on
+ * disk do not carry; leave the field to the next checkpoint then, the
+ * transition itself has already persisted every state that is safe
+ * without a flush.
+ *
+ * Compare the watermark rather than the state: record positions are
+ * unique, so a full round trip back to the sampled state cannot alias,
+ * while its flushed pages straddle the intermediate states all the same.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksumLsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checkPoint.dataChecksumState;
+ ControlFile->data_checksum_lsn = checksumLsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
/*
* Persist unloggedLSN value. It's reset on crash recovery, so this goes
* unused on non-shutdown checkpoints, but seems useful to store it always
@@ -7968,9 +8263,11 @@ CreateEndOfRecoveryRecord(void)
ControlFile->minRecoveryPoint = recptr;
ControlFile->minRecoveryPointTLI = xlrec.ThisTimeLineID;
- /* start with the latest checksum version (as of the end of recovery) */
+ /* persist the data checksum state this node ended recovery with */
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
@@ -8175,6 +8472,9 @@ CreateRestartPoint(int flags)
XLogRecPtr endptr;
XLogSegNo _logSegNo;
TimestampTz xtime;
+ uint32 checksum_state;
+ XLogRecPtr checksum_lsn;
+ bool checksum_is_local;
/* Concurrent checkpoint/restartpoint cannot happen */
Assert(!IsUnderPostmaster || MyBackendType == B_CHECKPOINTER);
@@ -8221,8 +8521,47 @@ CreateRestartPoint(int flags)
UpdateMinRecoveryPoint(InvalidXLogRecPtr, true);
if (flags & CHECKPOINT_IS_SHUTDOWN)
{
+ bool catchUpChecksums;
+
+ /*
+ * There is no new restartpoint to persist the data checksum state
+ * with, but a cleanly stopped node should not leave the control
+ * file behind the state replay reached: pg_checksums and
+ * pg_rewind read it, and an in-progress state there makes them
+ * refuse to run. Catching it up needs the same guarantee a
+ * restartpoint gives: that every page on disk carries a checksum,
+ * so flush the buffer pool first. Only a transition to "on" that
+ * no restartpoint followed can get here; the other states are
+ * already persisted by XLOG2_CHECKSUMS replay. Replay has ended
+ * by now, so the state cannot change under us.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ checksum_is_local = XLogCtl->data_checksum_is_local;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockAcquire(ControlFileLock, LW_SHARED);
+ catchUpChecksums =
+ (checksum_lsn != ControlFile->data_checksum_lsn &&
+ XLogRecPtrIsValid(lastCheckPointRecPtr));
+ LWLockRelease(ControlFileLock);
+
+ if (catchUpChecksums)
+ {
+ MemSet(&CheckpointStats, 0, sizeof(CheckpointStats));
+ CheckpointStats.ckpt_start_t = GetCurrentTimestamp();
+ CheckPointGuts(lastCheckPoint.redo, flags);
+ }
+
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
+ if (catchUpChecksums)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = checksum_is_local;
+ }
UpdateControlFile();
LWLockRelease(ControlFileLock);
}
@@ -8263,6 +8602,17 @@ CreateRestartPoint(int flags)
/* Update the process title */
update_checkpoint_display(flags, true, false);
+ /*
+ * Note the data checksum state the flush below starts under. Replay runs
+ * concurrently and can change the state while the flush is in progress,
+ * in which case the flush covers pages written under both states; see
+ * where the state is persisted further down.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
CheckPointGuts(lastCheckPoint.redo, flags);
/*
@@ -8322,8 +8672,26 @@ CreateRestartPoint(int flags)
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
}
- /* we shall start with the latest checksum version */
- ControlFile->data_checksum_version = lastCheckPoint.dataChecksumState;
+ /*
+ * Persist the data checksum state of this node. Not the state of the
+ * replayed checkpoint: that one belongs to the node that wrote it and
+ * may differ after an offline change on either side.
+ * ControlFile->checkPointCopy above keeps the replayed value on
+ * purpose, being a historical record used to resume replay rather
+ * than a tracker of node state.
+ *
+ * Persist only if the flush above ran under one state throughout; see
+ * CreateCheckPoint() for why, including why this compares the
+ * watermark and not the state.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksum_lsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
}
@@ -8764,9 +9132,22 @@ XLogReportParameters(void)
}
/*
- * Log the new state of checksums
+ * XLogChecksums
+ * Log and publish the new state of checksums
+ *
+ * Inserting the record and publishing the new state must be atomic with
+ * respect to a checkpoint sampling the state for its XLOG_CHECKPOINT_REDO
+ * record: without that, a checkpoint could read the old state after the
+ * record is already in WAL and insert a redo record that both follows the
+ * transition in WAL order and carries the pre-transition state. Recovery
+ * resuming from such a redo point would never replay the transition record
+ * and resolve the finished transition as interrupted.
+ * DataChecksumTransitionLock serializes the two; see CreateCheckPoint().
+ *
+ * Returns the end LSN of the inserted record, which the caller persists
+ * together with the new state as the data checksum watermark.
*/
-static void
+static XLogRecPtr
XLogChecksums(uint32 new_type)
{
xl_checksum_state xlrec;
@@ -8774,12 +9155,28 @@ XLogChecksums(uint32 new_type)
xlrec.new_checksum_state = new_type;
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
+
XLogBeginInsert();
XLogRegisterData((char *) &xlrec, sizeof(xl_checksum_state));
recptr = XLogInsert(RM_XLOG2_ID, XLOG2_CHECKSUMS);
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, recptr);
+
+ /* only loaded by SetDataChecksumsOn(), a no-op for the other callers */
+ INJECTION_POINT_CACHED("datachecksums-on-before-publish", NULL);
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ XLogCtl->data_checksum_version = new_type;
+ XLogCtl->data_checksum_lsn = recptr;
+ XLogCtl->data_checksum_is_local = false;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockRelease(DataChecksumTransitionLock);
+
XLogFlush(recptr);
+
+ return recptr;
}
/*
@@ -8967,11 +9364,19 @@ xlog_redo(XLogReaderState *record)
/* ControlFile->checkPointCopy always tracks the latest ckpt XID */
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->checkPointCopy.nextXid = checkPoint.nextXid;
- ControlFile->data_checksum_version = checkPoint.dataChecksumState;
UpdateControlFile();
LWLockRelease(ControlFileLock);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(checkPoint.dataChecksumState);
+
/*
* We should've already switched to the new TLI before replaying this
* record.
@@ -9207,19 +9612,17 @@ xlog_redo(XLogReaderState *record)
else if (info == XLOG_CHECKPOINT_REDO)
{
xl_checkpoint_redo redo_rec;
- bool new_state = false;
memcpy(&redo_rec, XLogRecGetData(record), sizeof(xl_checkpoint_redo));
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = redo_rec.data_checksum_version;
- SetLocalDataChecksumState(redo_rec.data_checksum_version);
- if (redo_rec.data_checksum_version != ControlFile->data_checksum_version)
- new_state = true;
- SpinLockRelease(&XLogCtl->info_lck);
-
- if (new_state)
- EmitAndWaitDataChecksumsBarrier(redo_rec.data_checksum_version);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(redo_rec.data_checksum_version,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(redo_rec.data_checksum_version);
}
else if (info == XLOG_LOGICAL_DECODING_STATUS_CHANGE)
{
@@ -9281,25 +9684,63 @@ xlog2_redo(XLogReaderState *record)
{
xl_checksum_state state;
XLogRecPtr lsn = record->EndRecPtr;
+ XLogRecPtr watermark;
memcpy(&state, XLogRecGetData(record), sizeof(xl_checksum_state));
+ SpinLockAcquire(&XLogCtl->info_lck);
+ watermark = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ /*
+ * Skip records this node has already applied. The control file
+ * carries the watermark, so this holds across restarts: recovery
+ * resuming below a record whose effect the control file already
+ * contains must not re-apply it, or it would revert a state change
+ * made with pg_checksums in between, which moves the state without
+ * writing any record of its own.
+ */
+ if (lsn <= watermark)
+ return;
+
/* advertise the location before the new state becomes visible */
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, lsn);
SpinLockAcquire(&XLogCtl->info_lck);
XLogCtl->data_checksum_version = state.new_checksum_state;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(state.new_checksum_state);
SpinLockRelease(&XLogCtl->info_lck);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = state.new_checksum_state;
+
+ /*
+ * Persist the new state, except when it is "on". Only "on" verifies
+ * checksums during reads, and between the last restartpoint and this
+ * record there may be pages on disk flushed under the old state; a
+ * crash-restart initializes verification from the control file and
+ * replay reads those pages back, so the control file may only say
+ * "on" once everything written under the transition has been flushed,
+ * as restartpoints and the end of recovery do. The opposite direction
+ * cannot wait for the restartpoint: once this record is replayed,
+ * evicted pages are written without checksums, and a control file
+ * still saying "on" would fail verification on exactly those pages
+ * after a crash.
+ */
+ if (state.new_checksum_state != PG_DATA_CHECKSUM_VERSION)
+ {
+ ControlFile->data_checksum_version = state.new_checksum_state;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ }
/*
* Update minRecoveryPoint to ensure that if recovery is aborted, we
* recover back up to this point before allowing hot standby again.
- * The new state is durable in pg_control while its location is only
- * tracked in shared memory; a standby becoming consistent below this
- * record would let base backups resume checksum verification with the
+ * The change location is only tracked in shared memory and is lost
+ * over a restart; a standby becoming consistent below this record
+ * would let base backups resume checksum verification with the
* location unknown. The local copies cannot be updated as long as
* crash recovery is happening and we expect all the WAL to be
* replayed.
diff --git a/src/backend/postmaster/datachecksum_state.c b/src/backend/postmaster/datachecksum_state.c
index f69258bc33d..099a6b4fe2e 100644
--- a/src/backend/postmaster/datachecksum_state.c
+++ b/src/backend/postmaster/datachecksum_state.c
@@ -94,9 +94,9 @@
*
* If processing is started in an online cluster then all backends are in Bd.
* If processing was halted by the cluster shutting down (due to a crash or
- * intentional restart), the controlfile state "inprogress-on" will be observed
- * on system startup and all backends will be placed in Bd. The controlfile
- * state will also be set to "off".
+ * intentional restart), the control file state "inprogress-on" will be
+ * observed on system startup and all backends will be placed in Bd. The
+ * control file state will also be set to "off".
*
* Backends transition Bd -> Bi via a procsignalbarrier which is emitted by the
* DataChecksumsWorkerLauncherMain. When all backends have acknowledged the
@@ -146,7 +146,25 @@
* stop writing data checksums as no backend is enforcing data checksum
* validation any longer.
*
- * 4. Future opportunities for optimizations
+ * 4. Interaction with offline data checksum changes
+ * -------------------------------------------------
+ * Enabling or disabling checksums offline with pg_checksums uses none of the
+ * machinery in this file, but the two mechanisms share the state kept in the
+ * control file, so their interaction is documented here.
+ *
+ * pg_checksums writes the new state to the control file and sets
+ * data_checksum_is_local, marking a state that no WAL record accounts for.
+ * Recovery then does not adopt the state carried by a replayed checkpoint
+ * record over it. The control file also carries a watermark, the WAL
+ * position through which data checksum transitions are covered. Replay skips
+ * transition records ending at or below the watermark, as their effect is
+ * already contained in the control file, and applies records above it as
+ * usual, whether they were written before or after an offline change. This
+ * is why an offline change in a replicated setup must be made on every node
+ * while all of them are stopped and caught up; see the pg_checksums
+ * documentation for the procedure.
+ *
+ * 5. Future opportunities for optimizations
* -----------------------------------------
* Below are some potential optimizations and improvements which were brought
* up during reviews of this feature, but which weren't implemented in the
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 0a70cffa081..3d366fd1114 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -371,6 +371,7 @@ WaitLSN "Waiting to read or update shared Wait-for-LSN state."
LogicalDecodingControl "Waiting to read or update logical decoding status information."
DataChecksumsWorker "Waiting for data checksums worker."
AioWorkerControl "Waiting to update AIO worker information."
+DataChecksumTransition "Waiting for a data checksum state transition to be written to WAL."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 3b3ae23f1a6..5d6ea318784 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -648,6 +648,16 @@ main(int argc, char *argv[])
ControlFile->data_checksum_version =
(mode == PG_MODE_ENABLE) ? PG_DATA_CHECKSUM_VERSION : PG_DATA_CHECKSUM_OFF;
+ /*
+ * Mark the state as changed locally, without a WAL record. Recovery
+ * then does not let a replayed checkpoint overwrite it, as no record
+ * could restore the change afterwards. The watermark is left alone:
+ * XLOG2_CHECKSUMS records at or below it stay covered, while records
+ * above it, which this node has not applied yet, still take effect on
+ * replay no matter when they were written.
+ */
+ ControlFile->data_checksum_is_local = true;
+
if (do_sync)
{
pg_log_info("syncing data directory");
diff --git a/src/bin/pg_controldata/pg_controldata.c b/src/bin/pg_controldata/pg_controldata.c
index 6fc87ed114d..c363ce3dbb7 100644
--- a/src/bin/pg_controldata/pg_controldata.c
+++ b/src/bin/pg_controldata/pg_controldata.c
@@ -349,6 +349,10 @@ main(int argc, char *argv[])
(ControlFile->float8ByVal ? _("by value") : _("by reference")));
printf(_("Data page checksum version: %u\n"),
ControlFile->data_checksum_version);
+ printf(_("Data checksum watermark: %X/%08X\n"),
+ LSN_FORMAT_ARGS(ControlFile->data_checksum_lsn));
+ printf(_("Data checksum state is node-local: %s\n"),
+ (ControlFile->data_checksum_is_local ? _("yes") : _("no")));
printf(_("Default char data signedness: %s\n"),
(ControlFile->default_char_signedness ? _("signed") : _("unsigned")));
printf(_("Mock authentication nonce: %s\n"),
diff --git a/src/bin/pg_resetwal/pg_resetwal.c b/src/bin/pg_resetwal/pg_resetwal.c
index 1542a56ca4b..cdfb9898066 100644
--- a/src/bin/pg_resetwal/pg_resetwal.c
+++ b/src/bin/pg_resetwal/pg_resetwal.c
@@ -923,6 +923,13 @@ RewriteControlFile(void)
ControlFile.backupEndPoint = InvalidXLogRecPtr;
ControlFile.backupEndRequired = false;
+ /*
+ * The old WAL is gone and the new position may lie below the old
+ * watermark, which would make replay ignore future checksum transition
+ * records. The state itself is kept.
+ */
+ ControlFile.data_checksum_lsn = InvalidXLogRecPtr;
+
/*
* Force the defaults for max_* settings. The values don't really matter
* as long as wal_level='minimal'; the postmaster will reset these fields
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 2e86fd158d0..0e4c2df3f4b 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -37,7 +37,8 @@ static void usage(const char *progname);
static void perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo);
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec);
static void createBackupLabel(XLogRecPtr startpoint, TimeLineID starttli,
XLogRecPtr checkpointloc);
@@ -531,7 +532,8 @@ main(int argc, char **argv)
* This is the point of no return. Once we start copying things, there is
* no turning back!
*/
- perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo);
+ perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo,
+ divergerec);
if (showprogress)
pg_log_info("syncing target data directory");
@@ -566,7 +568,8 @@ static void
perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo)
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec)
{
XLogRecPtr endrec;
TimeLineID endtli;
@@ -738,6 +741,32 @@ perform_rewind(filemap_t *filemap, rewind_source *source,
ControlFile_new.minRecoveryPoint = endrec;
ControlFile_new.minRecoveryPointTLI = endtli;
ControlFile_new.state = DB_IN_ARCHIVE_RECOVERY;
+
+ /*
+ * Keep the target's own data checksum state. Most of the data directory
+ * is still the target's: only blocks it changed since the divergence were
+ * copied from the source, so the source's state says nothing about the
+ * pages that stay. Replay from the last common checkpoint applies any
+ * WAL-logged transition the target has not seen (the watermark tells them
+ * apart), which converges the rewound server to the source's state
+ * whenever the WAL carries it.
+ */
+ ControlFile_new.data_checksum_version = ControlFile_target.data_checksum_version;
+ ControlFile_new.data_checksum_lsn = ControlFile_target.data_checksum_lsn;
+ ControlFile_new.data_checksum_is_local = ControlFile_target.data_checksum_is_local;
+
+ /*
+ * The watermark is only meaningful within the history the node replays.
+ * Records at or below the divergence point are common to both histories
+ * and stay covered, but a watermark above it was set by a transition
+ * record on the target's own abandoned fork: numerically it can cover
+ * transition records the source wrote after the divergence, and replay
+ * would skip them as already applied. Clamp it to the divergence point,
+ * so that every transition record on the source's history takes effect.
+ */
+ if (ControlFile_new.data_checksum_lsn > divergerec)
+ ControlFile_new.data_checksum_lsn = divergerec;
+
if (!dry_run)
update_controlfile(datadir_target, &ControlFile_new, do_sync);
}
diff --git a/src/bin/pg_upgrade/controldata.c b/src/bin/pg_upgrade/controldata.c
index 7543a988045..7259538de39 100644
--- a/src/bin/pg_upgrade/controldata.c
+++ b/src/bin/pg_upgrade/controldata.c
@@ -431,7 +431,7 @@ get_control_data(ClusterInfo *cluster)
cluster->controldata.date_is_int = strstr(p, "64-bit integers") != NULL;
got_date_is_int = true;
}
- else if ((p = strstr(bufin, "checksum")) != NULL)
+ else if ((p = strstr(bufin, "Data page checksum version:")) != NULL)
{
p = strchr(p, ':');
diff --git a/src/include/catalog/pg_control.h b/src/include/catalog/pg_control.h
index 7b5404460ec..3a55228d180 100644
--- a/src/include/catalog/pg_control.h
+++ b/src/include/catalog/pg_control.h
@@ -22,7 +22,7 @@
/* Version identifier for this pg_control format */
-#define PG_CONTROL_VERSION 1903
+#define PG_CONTROL_VERSION 1904
/* Nonce key length, see below */
#define MOCK_AUTH_NONCE_LEN 32
@@ -237,6 +237,33 @@ typedef struct ControlFileData
/* Current data checksums state */
uint32 data_checksum_version;
+ /*
+ * WAL position through which data checksum transitions are covered.
+ * Replay ignores XLOG2_CHECKSUMS records ending at or below this point:
+ * their effect is already contained in data_checksum_version, or an
+ * offline pg_checksums change made after they were first applied
+ * supersedes them. Ordinarily this is the end of the newest such record
+ * this node has written or applied, but a tool may store any position
+ * that covers the same set of records. If the node has never written or
+ * applied such a record this field shall be set to InvalidXLogRecPtr.
+ *
+ * The comparison has no timeline context, so the value is only valid
+ * within the WAL history this node replays. A tool that moves the node
+ * to another history must clamp the watermark to the point where the
+ * histories fork, as pg_rewind does, or reset it, as pg_resetwal does.
+ */
+ XLogRecPtr data_checksum_lsn;
+
+ /*
+ * True when data_checksum_version was last set by pg_checksums in an
+ * offline operation rather than by an online, WAL-logged transition. Such
+ * a state is local to this node and not derived from WAL, so recovery
+ * must not replace it with a state taken from a checkpoint record;
+ * nothing in the WAL could restore the change once it is overwritten.
+ * Cleared by the next WAL-logged transition.
+ */
+ bool data_checksum_is_local;
+
/*
* True if the default signedness of char is "signed" on a platform where
* the cluster is initialized.
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index d7eb648bd27..8d858be9927 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -89,6 +89,7 @@ PG_LWLOCK(54, WaitLSN)
PG_LWLOCK(55, LogicalDecodingControl)
PG_LWLOCK(56, DataChecksumsWorker)
PG_LWLOCK(57, AioWorkerControl)
+PG_LWLOCK(58, DataChecksumTransition)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/test/modules/test_checksums/Makefile b/src/test/modules/test_checksums/Makefile
index 71455cd5577..80f54bf6d8c 100644
--- a/src/test/modules/test_checksums/Makefile
+++ b/src/test/modules/test_checksums/Makefile
@@ -9,7 +9,7 @@
#
#-------------------------------------------------------------------------
-EXTRA_INSTALL = src/test/modules/injection_points
+EXTRA_INSTALL = contrib/pg_buffercache src/test/modules/injection_points
export enable_injection_points
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index fb7129d796f..e5d38fafb7d 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -35,6 +35,16 @@ tests += {
't/009_fpi.pl',
't/010_backup_straddle.pl',
't/011_standby_straddle.pl',
+ 't/012_offline_standby.pl',
+ 't/013_rewind.pl',
+ 't/014_lockstep.pl',
+ 't/015_standby_crash_after_disable.pl',
+ 't/016_promote_enable_crash.pl',
+ 't/017_restartpoint_race.pl',
+ 't/018_enable_crash_windows.pl',
+ 't/019_standby_shutdown_catchup.pl',
+ 't/020_cascade_divergence.pl',
+ 't/021_rewind_divergent_transitions.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
new file mode 100644
index 00000000000..c0ec374c8be
--- /dev/null
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -0,0 +1,304 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Offline checksum changes with pg_checksums are local to one node. A
+# standby must neither adopt the state of the primary from replayed
+# checkpoint records, nor lose its own offline change to them.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Scenario 1: enable offline on the primary only. The standby must
+# stay off, warn about the mismatch, and remain readable.
+$standby->stop;
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+test_checksum_state($primary, 'on');
+test_checksum_state($standby, 'off');
+
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after offline enable on the primary');
+
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+my @warnings = $log =~ /(does not match the state)/g;
+is(scalar(@warnings), 1, 'mismatch warned once per remote value');
+
+# Matching states re-arm the warning: undo the divergence on the primary,
+# then diverge again to the same value, all without restarting the standby.
+$primary->stop;
+$primary->checksum_disable_offline;
+$logstart = -s $standby->logfile;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no warning while the states match again');
+like($log, qr/now agrees with the replayed WAL/,
+ 'convergence is reported once the states match again');
+
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+test_checksum_state($standby, 'off');
+
+# The local state survives both clean and immediate restarts.
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+# Still in scenario 1's divergence: take a base backup *from the
+# standby*. A backup taken on a standby uses the last restartpoint as
+# its starting checkpoint (do_pg_backup_start()), so the record at the
+# redo point was written by the upstream primary and carries the
+# primary's state. Recovery from such a backup must not adopt that
+# state: the files were copied from the standby, and the new node
+# would come up verifying checksums its files do not have.
+
+# Make sure the standby has written pages under its own "off" state, so
+# its files really do lack checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,50000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', "SELECT count(*) FROM t2;");
+
+$standby->backup('from_standby');
+my $newnode = PostgreSQL::Test::Cluster->new('newnode');
+$newnode->init_from_backup($standby, 'from_standby');
+
+# Stream from the primary, not from the standby the files were copied
+# from, so the new node replays the "on" primary's records.
+$newnode->enable_streaming($primary);
+$newnode->start;
+
+# The new node's files all came from a cluster running with checksums
+# off. Anything but "off" here means it adopted the primary's state
+# through the checkpoint record at the redo point.
+my ($rc, $stdout, $stderr) = $newnode->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'the node copied from the standby accepts connections')
+ or diag("stderr: $stderr");
+is($stdout, 'off', 'backup of an "off" standby comes up with checksums off');
+
+# A checkpoint record streamed from the "on" primary must not flip the
+# state either.
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($newnode);
+test_checksum_state($newnode, 'off');
+
+# And it must be able to read the pages the standby wrote without
+# checksums.
+(undef, $stdout, $stderr) =
+ $newnode->psql('postgres', "SELECT count(*) FROM t2;");
+is($stdout, '50000', 'pages copied from the standby are readable')
+ or diag("stderr: $stderr");
+
+my $newnode_log = PostgreSQL::Test::Utils::slurp_file($newnode->logfile);
+unlike(
+ $newnode_log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the node copied from the standby');
+
+$newnode->stop('immediate');
+
+# Converge the cluster: enable offline on the standby too.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+$primary->wait_for_catchup($standby);
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after converging');
+
+# Scenario 2: disable offline on the standby only. The replayed
+# checkpoint records of the still-enabled primary must not override it.
+$standby->stop;
+$standby->checksum_disable_offline;
+$standby->start;
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'off');
+
+# Restartpoints must persist the local state, not the replayed copy.
+$standby->safe_psql('postgres', "CHECKPOINT;");
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable with checksums disabled locally');
+
+# Scenario 3: crash-restart right after an online transition, before the
+# next restartpoint. Replay then resumes from an older restartpoint whose
+# checkpoint records still carry the pre-transition state. Those must
+# still match the state seeded from the control file, and the transition
+# itself must be re-established by re-replaying the XLOG2_CHECKSUMS record,
+# without a spurious mismatch warning along the way.
+
+# Converge first: bring the standby back to "on" offline.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+disable_data_checksums($primary, wait => 'off');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+# Crash-restart the standby immediately, before any restartpoint has had a
+# chance to persist the new state to its control file.
+$logstart = -s $standby->logfile;
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001',
+ 'standby readable after crash-restart across an online transition');
+
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no spurious mismatch warning after crash-restart across an online transition'
+);
+
+# Scenario 4: a standby stopped while replaying an interrupted online
+# transition keeps the interrupted state in its own control file. A
+# primary is never caught this way, as its checksums launcher resolves
+# inprogress-on back to off from its exit cleanup. A standby has no
+# launcher; it carries forward whatever the last replayed record left
+# it in.
+
+# Block an online enable on the primary at inprogress-on with a
+# blocking temp table, same trick as in 004_offline.pl.
+my $bsession = $primary->background_psql('postgres');
+$bsession->query_safe('CREATE TEMPORARY TABLE tt (a integer);');
+enable_data_checksums($primary, wait => 'inprogress-on');
+
+# The standby picks up the in-progress state from the XLOG2_CHECKSUMS record.
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# Stop the standby cleanly; its restartpoint persists inprogress-on to
+# its own control file, since nothing on a standby resolves it away.
+$standby->stop;
+$standby->start;
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# The primary still sits at inprogress-on, so a base backup taken now
+# copies a control file and a redo point that both carry the
+# in-progress state; replay of the XLOG2_CHECKSUMS records completes
+# the transition on the new standby.
+$primary->backup('inprogress_backup');
+my $standby2 = PostgreSQL::Test::Cluster->new('standby2');
+$standby2->init_from_backup($primary, 'inprogress_backup',
+ has_streaming => 1);
+$standby2->start;
+
+# Backup label recovery adopts the redo point's state immediately,
+# before any further WAL is replayed.
+wait_for_checksum_state($standby2, 'inprogress-on');
+
+$bsession->quit;
+wait_for_checksum_state($primary, 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable once the transition completes');
+
+# Scenario 4 continued: the backup taken mid-transition must also
+# complete the transition and stay readable.
+$primary->wait_for_catchup($standby2);
+wait_for_checksum_state($standby2, 'on');
+
+is($standby2->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'mid-transition backup standby readable after the transition');
+
+$standby2->stop;
+
+# Nothing along this standby's replay can legitimately disagree: it
+# started out at the same inprogress-on state as the primary.
+my $standby2_log = PostgreSQL::Test::Utils::slurp_file($standby2->logfile);
+unlike(
+ $standby2_log,
+ qr/does not match the state/,
+ 'no mismatch warning on a backup taken mid-transition');
+
+# No extra CHECKPOINT needed: the transition's completion checkpoint
+# wrote every page with a checksum, and the shutdown restartpoint
+# flushes the rest.
+command_ok([ 'pg_checksums', '--check', '-D', $standby2->data_dir ],
+ 'checksums valid on the mid-transition backup standby');
+
+$standby->stop;
+$primary->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/013_rewind.pl b/src/test/modules/test_checksums/t/013_rewind.pl
new file mode 100644
index 00000000000..a791e24317d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/013_rewind.pl
@@ -0,0 +1,201 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test pg_rewind across an online data checksum enable.
+#
+# A clean switchover leaves the shutdown checkpoint of the old primary
+# as the last common checkpoint between the two nodes. When data
+# checksums are enabled online on the new primary before the old one is
+# rewound, pg_rewind installs the control file of the new primary, which
+# already claims checksums are fully enabled, while replay begins at the
+# shutdown checkpoint whose record still carries the old state. Replay
+# of the WAL stretch from before the enable must run with checksums off,
+# as recorded in the checkpoint, else it would verify pages which never
+# had checksums written and fail recovery.
+#
+# The new primary requests a checkpoint right after promotion, and
+# replaying its CHECKPOINT_REDO record would repair the state before
+# any interesting WAL is reached. Hold the checkpointer on the standby
+# in a restartpoint over the promotion, like in the recovery test
+# 041_checkpoint_at_promote, so that the post-promotion writes end up
+# in WAL before the first checkpoint of the new timeline.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Old primary. full_page_writes is off so that the updates done on the
+# promoted node do not carry full page images, forcing replay on the
+# rewound node to read the pages from disk. wal_log_hints is required
+# by pg_rewind on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+full_page_writes = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+log_checkpoints = on
+]);
+$node_a->start;
+
+if (!$node_a->check_extension('injection_points'))
+{
+ plan skip_all => 'Extension injection_points not installed';
+}
+
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+$node_a->safe_psql('postgres', "CREATE EXTENSION injection_points;");
+
+# Set the hint bits on t before taking the backup, so that reads on the
+# promoted node do not emit full page images for its pages later.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->safe_psql('postgres', "SELECT count(*) FROM t;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_a, 'off');
+test_checksum_state($node_b, 'off');
+
+# Hold the next restartpoint on the standby.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point', 'wait');");
+
+# Give the restartpoint a checkpoint record to work on, then start it
+# in a background session; it will block on the injection point with
+# the checkpointer busy until released.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+my $bg_psql = $node_b->background_psql('postgres', on_error_stop => 0);
+$bg_psql->query_until(
+ qr/starting_restartpoint/, q(
+ \echo starting_restartpoint
+ CHECKPOINT;
+));
+$node_b->wait_for_event('checkpointer', 'create-restart-point');
+
+# Clean switchover: the shutdown checkpoint of A streams to B and
+# becomes the last common checkpoint.
+$node_a->stop('fast');
+
+my ($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+$stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+my $shutdown_ckpt = $1;
+$node_b->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_ckpt'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+
+my $logstart = -s $node_b->logfile;
+$node_b->promote;
+
+# Accidental restart of the old primary, diverging its timeline.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->stop('fast');
+
+# Updates on the new primary before its first checkpoint; replay of
+# these on the rewound node has to read the pages from disk.
+$node_b->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 25 = 0;");
+ok( !$node_b->log_contains("checkpoint complete", $logstart),
+ "no checkpoint on the new timeline before the updates");
+
+# Release the checkpointer; the queued post-promotion checkpoint runs
+# after the updates.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+$bg_psql->quit;
+
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+# pg_rewind refuses to run with full_page_writes disabled on the
+# source; the updates it was disabled for are already in WAL.
+$node_b->safe_psql('postgres', "ALTER SYSTEM SET full_page_writes = on;");
+$node_b->safe_psql('postgres', "SELECT pg_reload_conf();");
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the new primary');
+
+# Replay on the rewound node must start at the shutdown checkpoint of
+# the switchover, with the control file of the new primary.
+my $backup_label = slurp_file($node_a->data_dir . '/backup_label');
+$backup_label =~ /^CHECKPOINT LOCATION: ([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from backup_label";
+is($1, $shutdown_ckpt, 'replay starts at the switchover checkpoint');
+
+($stdout, $stderr) = run_command(
+ [
+ 'pg_waldump',
+ '-p' => $node_a->data_dir . '/pg_wal',
+ '-t' => 1,
+ '-s' => $shutdown_ckpt,
+ '-n' => 1,
+ ]);
+like($stdout, qr/CHECKPOINT_SHUTDOWN/,
+ 'last common checkpoint is a shutdown checkpoint');
+
+# pg_rewind keeps the target's own checksum state in the control file it
+# writes; the online enable reaches the rewound node through WAL replay
+# below, not through the copied control file.
+($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+like(
+ $stdout,
+ qr/^Data page checksum version:\s*0$/m,
+ 'rewound node keeps its own checksum state in the control file');
+
+# Start the rewound node as a standby of the new primary. Replay runs
+# through the pre-enable WAL stretch and the online enable.
+#
+# The rewind replaced the configuration files with those of the new
+# primary, so put the port back.
+my $connstr = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_b->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/014_lockstep.pl b/src/test/modules/test_checksums/t/014_lockstep.pl
new file mode 100644
index 00000000000..713f37fb22c
--- /dev/null
+++ b/src/test/modules/test_checksums/t/014_lockstep.pl
@@ -0,0 +1,182 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# The lockstep procedure for offline checksum changes in a replication
+# setup: stop all nodes, run pg_checksums on all of them, restart.
+# Replay of checkpoint records written before the change must not
+# revert the state of the standby.
+#
+# The same pair then runs the procedure in the disable direction, and
+# finally across a divergence: after a failover both nodes get their
+# checksums enabled offline, while the last common checkpoint still
+# carries "off". pg_rewind must keep the offline "on" on the rewound
+# node, since no record in the replayed WAL could ever restore it.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# wal_log_hints keeps the pair eligible for pg_rewind without checksums.
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Part 1: the lockstep procedure in the enable direction. Stop the
+# standby first: the WAL written after this point is replayed only
+# after the offline switch, and every checkpoint record in it still
+# carries the old state.
+$standby->stop;
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 0;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 1;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+# The lockstep procedure.
+$primary->checksum_enable_offline;
+$standby->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to off.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "off" in the replayed WAL/,
+ 0);
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep enable');
+
+# Crash the standby and replay the same stretch again.
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'on');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after crash restart');
+
+# Once a post-switch checkpoint has been replayed the states match and
+# no warning may be logged.
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no mismatch warning once the states match');
+
+# Every page the standby wrote in this window must carry a checksum.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$standby->stop;
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'checksums valid on the standby');
+
+# Part 2: the lockstep procedure in the other direction. The standby
+# is already stopped; give it a pre-switch WAL stretch to replay, with
+# every checkpoint record in it still carrying "on" -- an explicit
+# checkpoint plus the shutdown checkpoint written by stop().
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 2;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+$primary->checksum_disable_offline;
+$standby->checksum_disable_offline;
+
+my $disable_logstart = -s $standby->logfile;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to on.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $disable_logstart);
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep disable');
+
+# Part 3: failover, divergence and pg_rewind across offline enables.
+# The roles swap from here on: the standby becomes the new primary and
+# the old primary is rewound to follow it, but the variables keep the
+# names they had above.
+#
+# Failover: promote the standby, then diverge the old primary.
+$standby->promote;
+$standby->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$primary->stop;
+
+# The lockstep procedure across the divergence: both nodes get their
+# checksums enabled offline.
+$primary->checksum_enable_offline;
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+
+# The states match, so the rewind proceeds without complaint.
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $primary->data_dir,
+ '--source-server' => $standby->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled offline on both nodes');
+
+# The rewind replaced the configuration files with those of the source,
+# so put the port back. The copied file also carries the source's own
+# stale primary_conninfo; enable_streaming() below appends ours, which
+# wins as the later entry.
+$primary->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$primary->port]}
+]);
+$primary->enable_streaming($standby);
+
+my $rewind_logstart = -s $primary->logfile;
+$primary->start;
+$standby->wait_for_catchup($primary);
+
+is($primary->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+
+# The offline enable survives: replay from the common checkpoint must
+# not resurrect the pre-divergence "off".
+test_checksum_state($primary, 'on');
+
+$log =
+ PostgreSQL::Test::Utils::slurp_file($primary->logfile, $rewind_logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no divergence reported between the rewound server and its source');
+
+$primary->stop;
+$standby->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
new file mode 100644
index 00000000000..b9a79174546
--- /dev/null
+++ b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
@@ -0,0 +1,137 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that a standby crashing after a replayed online disable, but before
+# its next restartpoint, restarts cleanly.
+#
+# Pages dirtied before the XLOG2_CHECKSUMS record and evicted after it are
+# written out under the new "off" state, without checksums. The control
+# file must follow the record immediately in this direction: a crash-restart
+# initializes checksum verification from the control file, and one still
+# saying "on" would fail verification on exactly those pages while replaying
+# records older than the state change.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$primary->start;
+
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,1000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+
+# Small buffer pool so that replaying a bulk load evicts the pages we care
+# about, and no restartpoints so the control file stays where it is.
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 1MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_delay = 10000
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'on');
+
+# No full page images, so that replay has to read the pages of "t" back from
+# disk instead of overwriting them from the WAL.
+$primary->append_conf('postgresql.conf', 'full_page_writes = off');
+$primary->reload;
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" on the standby, *before* the state change. They are
+# not flushed: the standby has no restartpoint from here on.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+# Online disable. The standby replays it and switches to "off", but its
+# control file is not updated.
+$primary->safe_psql('postgres', 'SELECT pg_disable_data_checksums();');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+my ($ctl_before) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($ctl_state) = $ctl_before =~ /Data page checksum version:\s+(\d+)/;
+note(
+ "standby control file data_checksum_version after the disable: "
+ . "$ctl_state (0 = off, 1 = on)");
+
+# Force the standby to evict the dirty pages of "t" now that it is "off":
+# they are written back without checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE filler AS SELECT generate_series(1,300000) AS a;");
+$primary->wait_for_catchup($standby);
+
+# Crash the standby before it gets a chance to run a restartpoint.
+$standby->stop('immediate');
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'standby restarts after crashing between a checksum state '
+ . 'change and the next restartpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'the pages written while "off" are readable')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("standby log tail:\n" . join("\n", @lines[ -12 .. -1 ]))
+ if @lines >= 12;
+ fail('the pages written while "off" are readable');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/016_promote_enable_crash.pl b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
new file mode 100644
index 00000000000..13c983136ef
--- /dev/null
+++ b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
@@ -0,0 +1,146 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a standby promoted right after replaying an online enable, before any
+# restartpoint has flushed the rewritten pages.
+#
+# The end-of-recovery record persists the checksum state without flushing the
+# buffer pool, while the control file still points at a restartpoint older than
+# the transition. Persisting "on" there would make a crash before the
+# post-promotion checkpoint verify checksums over the pages that were flushed
+# while the state was still "off"; promotion must take the full checkpoint
+# instead.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Dirty the pages of "t" without full page images, then push them out to disk
+# on the standby while checksums are still off.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Online enable on the primary; the standby replays it into its own buffers,
+# where the rewritten pages stay dirty (no restartpoint, no bgwriter).
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+note("standby control file before the promotion: $before");
+
+# Promote. The end-of-recovery record persists the live state without any
+# buffer flush; the checkpoint requested afterwards is not immediate.
+$standby->promote;
+$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery();')
+ or die 'timed out waiting for the promotion';
+$standby->stop('immediate');
+
+($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($ckpt) = $ctl =~ /Latest checkpoint location:\s+(\S+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note(
+ "standby control file after the promotion: $after, checkpoint $ckpt, redo $redo"
+);
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'promoted node restarts after crashing before its first checkpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/017_restartpoint_race.pl b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
new file mode 100644
index 00000000000..10919c9741c
--- /dev/null
+++ b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
@@ -0,0 +1,151 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a restartpoint whose flush races the replay of an online enable.
+#
+# Replay keeps running while CheckPointGuts() writes out the buffer pool, so
+# the state at the end of the flush can be newer than the one the flush ran
+# under. Persisting it would claim checksums for pages the flush wrote out
+# before the transition, against a redo pointer that predates it.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# C1: the checkpoint the pending restartpoint will target.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" without full page images and push them to disk on
+# the standby while it is still "off".
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Start a restartpoint for C1 and hold it right after CheckPointGuts().
+$standby->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point','wait');");
+my $bg = $standby->background_psql('postgres');
+$bg->query_until(qr//, "\\echo restartpoint\nCHECKPOINT;\n");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint injection point';
+
+# The startup process keeps replaying while the checkpointer is held: the
+# whole online enable lands, and the rewritten pages stay dirty.
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+# Let the restartpoint finish; it persists the state it samples now.
+$standby->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint to finish';
+$standby->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+
+$standby->stop('immediate');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note("standby control file after the restartpoint: $after, redo $redo");
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started, 'standby restarts after a restartpoint that raced the enable');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/018_enable_crash_windows.pl b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
new file mode 100644
index 00000000000..4dac590e819
--- /dev/null
+++ b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
@@ -0,0 +1,508 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Crash and checkpoint windows inside an online data checksum enable on
+# a single node. Each scenario starts from "off" on the same cluster,
+# runs the enable into a held injection point, and checks what a crash
+# or a concurrent checkpoint in that window may and may not do to the
+# control file and the pages on disk.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+use IPC::Run;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum_extended' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum_extended\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Bring the node back to "off" with no leftover table between
+# scenarios. The crash restarts already resolve an interrupted enable
+# to "off", so only disable when something is left to disable. Those
+# restarts also clear any attached injection points. The checkpoint
+# bounds the WAL the next scenario's crash recovery has to replay.
+sub reset_scenario
+{
+ my ($node) = @_;
+
+ BAIL_OUT('node is not running after the previous scenario')
+ unless $node->is_alive;
+
+ my $state = $node->safe_psql('postgres', 'SHOW data_checksums;');
+ disable_data_checksums($node, wait => 'off') if $state ne 'off';
+ $node->safe_psql('postgres', 'DROP TABLE IF EXISTS t;');
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+ test_checksum_state($node, 'off');
+}
+
+my $node = PostgreSQL::Test::Cluster->new('crash_windows');
+$node->init(no_data_checksums => 1);
+
+# Scenario 1 needs the pages of "t" to reach disk without a full page
+# image, hence full_page_writes and wal_log_hints off. Scenario 2
+# needs them to stay dirty in shared buffers until a checkpoint, hence
+# no bgwriter and a large shared_buffers. wal_level is pinned because
+# init() writes "minimal" for a non-streaming node. The rest merely
+# tolerate the settings.
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+wal_level = replica
+max_wal_size = 10GB
+full_page_writes = off
+wal_log_hints = off
+shared_buffers = 512MB
+bgwriter_lru_maxpages = 0
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$node->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+
+test_checksum_state($node, 'off');
+
+# Scenario 1: a primary crashes inside SetDataChecksumsOn(), just before the
+# forced checkpoint that flushes the rewritten pages, with crash recovery
+# resuming from a checkpoint older than the transition. The control file
+# must still say "inprogress-on" there: replay does not adopt the state of
+# the checkpoint record it resumes from, so an "on" written before the flush
+# would stay in effect while replay reads pages whose rewrite never reached
+# disk. Here the pages went out through the rewriting worker's ring buffer,
+# so only the control file state discriminates; see the next scenario for
+# the case where they did not.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Establish the checkpoint that crash recovery will resume from.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty the pages of "t" without full page images, then push them out
+ # to disk while checksums are still off.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+ my $evicted = $node->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+ note("evict_relation on t: $evicted, relpath $relpath");
+
+ # Stop the enable right after the control file has been updated to
+ # "on" but before the checkpoint that flushes the rewritten pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+ $node->safe_psql('postgres', 'SELECT pg_enable_data_checksums();');
+
+ $node->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'datachecksums-on-before-checkpoint';"
+ ) or die 'timed out waiting for the injection point';
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ $node->stop('immediate');
+
+ # Show the on-disk checksum field of the first page of "t".
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 2: the same crash as above, but with the rewritten pages resident
+# in shared buffers instead of gone out through the ring. The rewriting
+# worker reads through a BAS_VACUUM ring, which writes pages back as the
+# ring recycles, but a page already resident in shared buffers is not read
+# through the ring: ReadBufferExtended() hands back the existing buffer, and
+# it stays dirty until a checkpoint. The control file may therefore not
+# say "on" before that checkpoint has run, or crash recovery would resume
+# from a checkpoint older than the transition with verification already
+# enabled, and replay records that read those still-unchecksummed pages.
+# full_page_writes is off so that the records do not simply overwrite the
+# pages with a full page image.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,100000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Establish the checkpoint that crash recovery will resume from, with the
+ # pages of "t" written out while checksums are still off.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty those pages again without emitting full page images, and leave
+ # them in shared buffers. Replay of these records has to read the
+ # pages from disk.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+
+ my $dirty_before = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_before, '>', 0,
+ 'pages of t are resident and dirty before enabling checksums');
+
+ # Hold the enabling right before the checkpoint that flushes the rewritten
+ # pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ # The rewritten pages must still be sitting dirty in shared buffers, or
+ # the window this test is about does not exist.
+ my $dirty_after = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_after, '>', 0,
+ 'rewritten pages of t are still dirty before the checkpoint');
+
+ $node->stop('immediate');
+
+ # The on-disk copy is the one written before the transition, without a
+ # checksum.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ is($pd_checksum, 0, 'block 0 of t on disk carries no checksum');
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 3: a checkpoint that runs between the XLOG2_CHECKSUMS("on")
+# record and the control file write at the end of SetDataChecksumsOn() must
+# not make a crash throw the completed transition away. SetDataChecksumsOn()
+# writes the record, flips shared memory to "on", emits the barrier and only
+# then requests the checkpoint that flushes the rewritten pages; the control
+# file is written after that checkpoint returns. A crash in that window is
+# harmless only as long as recovery still starts before the record. Any
+# checkpoint completing in the window moves the redo point past the record,
+# so recovery would never see it, would come up with the control file's
+# "inprogress-on" and StartupXLOG() would demote that to "off", even though
+# every page on disk carries a checksum by then. CreateCheckPoint()
+# therefore persists the state the checkpoint ran under, the same way
+# CreateRestartPoint() does. Here the window is made deterministic by
+# holding the launcher at the datachecksums-on-before-checkpoint injection
+# point and checkpointing from another session.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Hold the launcher after the record, the shared memory flip and the
+ # barrier, but before the checkpoint SetDataChecksumsOn() requests
+ # itself.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ # Every backend already sees "on" and writes checksums.
+ test_checksum_state($node, 'on');
+
+ # ... while the control file still says "inprogress-on".
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($before, '3', 'control file says "inprogress-on" inside the window');
+
+ # A concurrent checkpoint. It flushes the rewritten pages and moves
+ # the redo point past the XLOG2_CHECKSUMS("on") record, so it has to
+ # record the "on" state in the control file as well.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($after, '1',
+ 'the concurrent checkpoint records "on" in the control file');
+
+ $node->stop('immediate');
+
+ # The checkpoint flushed the rewritten pages, so they carry a checksum on
+ # disk: the transition really did complete.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ isnt($pd_checksum, 0, 'block 0 of t on disk carries a checksum');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition is complete on disk, so the cluster has to come back
+ # "on".
+ test_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+}
+reset_scenario($node);
+
+# Scenario 4: a crash between the checkpoint an online enable requests and
+# the control file write that follows it must not lose the transition. The
+# last steps of SetDataChecksumsOn() are
+#
+# WAL record -> shmem -> barrier -> checkpoint -> persist
+#
+# The checkpoint is what makes "on" safe to persist, but it also moves the
+# redo point above the XLOG2_CHECKSUMS record that carries the new state.
+# Crash recovery started from that checkpoint therefore never replays the
+# record, so without the state the checkpoint itself persists, a crash in
+# the remaining window would bring the cluster back at "inprogress-on",
+# which StartupXLOG() resolves to "off", discarding a transition whose
+# pages are all on disk with a checksum. The previous scenario exercises
+# the same window through a checkpoint requested by another session; this
+# one closes it from the other side: the transition's own checkpoint is the
+# one that moves the redo point, so persisting the state from
+# CreateCheckPoint() is what has to save it, not the write below the
+# injection point.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Hold the launcher after the checkpoint that licenses "on" has
+ # completed and before the state reaches the control file.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-after-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-after-checkpoint');
+
+ # The transition is complete as far as the running cluster is concerned.
+ test_checksum_state($node, 'on');
+
+ # The checkpoint has already recorded it, so the pending write below the
+ # injection point has nothing left to do.
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($version) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($version, '1',
+ 'the requested checkpoint recorded "on" in the control file');
+
+ # Crash before the control file write that follows the checkpoint.
+ $node->stop('immediate');
+
+ $node->start;
+
+ # Every page on disk carries a checksum and the checkpoint that
+ # flushed them completed, so the cluster has to come back verifying
+ # them.
+ test_checksum_state($node, 'on');
+
+ is($node->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '10000', 'relation readable after the crash');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ # The data directory must be one the offline tools accept.
+ $node->stop;
+ $node->command_ok([ 'pg_checksums', '--check', '-D', $node->data_dir ],
+ 'pg_checksums accepts the data directory');
+
+ # Leave the node running for the reset.
+ $node->start;
+}
+reset_scenario($node);
+
+# Scenario 5: a checkpoint racing SetDataChecksumsOn() between the insertion
+# of the XLOG2_CHECKSUMS("on") record and the shared memory update must not
+# insert an XLOG_CHECKPOINT_REDO record that follows the transition in WAL
+# order while still carrying "inprogress-on". Recovery resuming from such a
+# redo point never replays the preceding transition record, comes up in
+# "inprogress-on" and resolves the finished transition as interrupted.
+# XLogChecksums() closes the window by inserting the record and publishing
+# the new state under DataChecksumTransitionLock, which CreateCheckPoint()
+# takes around sampling the state and inserting the redo record. Here the
+# launcher is held between the two steps at the
+# datachecksums-on-before-publish injection point, and a concurrent
+# CHECKPOINT has to block on the lock instead of completing inside the
+# window.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # The datachecksums-on-before-publish point fires inside a critical
+ # section, where the wait machinery must not allocate. Waiting once
+ # at the datachecksums-enable-checksums-delay point, which the
+ # launcher runs outside the critical section, initializes it; see
+ # 050_redo_segment_missing.pl for the same recipe around
+ # create-checkpoint-run.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-enable-checksums-delay','wait');"
+ );
+
+ # Hold the launcher after the XLOG2_CHECKSUMS("on") record is in WAL but
+ # before the new state is published in shared memory.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-publish','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-enable-checksums-delay');
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-enable-checksums-delay');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-enable-checksums-delay');");
+
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-publish');
+
+ # The record is in WAL, the published state is still the old one.
+ test_checksum_state($node, 'inprogress-on');
+
+ # A concurrent checkpoint. It must block on DataChecksumTransitionLock
+ # before inserting its redo record rather than complete inside the window.
+ my $checkpointer = IPC::Run::start(
+ [ 'psql', '-XAtq', '-d', $node->connstr('postgres'), '-c', 'CHECKPOINT;' ],
+ '<' => '/dev/null',
+ '>' => '/dev/null',
+ '2>' => '/dev/null',
+ IPC::Run::timer($PostgreSQL::Test::Utils::timeout_default));
+
+ ok( $node->poll_query_until(
+ 'postgres',
+ "SELECT wait_event = 'DataChecksumTransition' "
+ . "FROM pg_stat_activity WHERE backend_type = 'checkpointer';"),
+ 'concurrent checkpoint blocks on DataChecksumTransitionLock');
+
+ # Release the launcher; the checkpoint then samples the published "on" and
+ # its redo record follows the transition record.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-publish');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-publish');");
+
+ $checkpointer->finish;
+
+ # Crash while the launcher's own checkpoint may still be in flight.
+ # Recovery resumes from the concurrent checkpoint's redo point, which
+ # now lies above the transition record and carries "on".
+ $node->stop('immediate');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition completed, so the cluster has to come back "on".
+ wait_for_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ $node->stop;
+}
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
new file mode 100644
index 00000000000..abcc234ab22
--- /dev/null
+++ b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
@@ -0,0 +1,138 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A standby stopped after replaying an online enable, but before the
+# checkpoint record that follows it on the primary.
+#
+# The shutdown restartpoint has no new checkpoint record to work from
+# and is skipped, so the control file would keep the in-progress state
+# the transition passed through, even though replay left the node at
+# "on" and rewrote every page. pg_checksums reads that field and
+# refuses to run on an in-progress state, so the shutdown has to flush
+# and catch it up instead.
+#
+# The caught-up control file then carries the watermark of the "on"
+# record, which the second half of this test relies on: an offline
+# disable made while the standby is down must survive the re-replay of
+# that record on the next startup, or the change would be silently
+# reverted.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+wal_keep_size = 1GB
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Anchor the standby's last restartpoint here, so that nothing replayed
+# from now on gives the shutdown restartpoint a newer checkpoint record
+# to use.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Hold the enable right after the state change record, before the
+# checkpoint it requests once the transition is complete.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+enable_data_checksums($primary);
+$primary->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+wait_for_checksum_state($primary, 'on');
+
+# The state change record is already flushed and streams on its own;
+# no checkpoint record follows while the launcher is held.
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+$standby->stop('fast');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($state, '1',
+ 'shutdown catches the control file up with the replayed state');
+
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'pg_checksums verifies the stopped standby');
+
+# Disable checksums offline while the standby is down. The restart
+# below resumes replay under the last restartpoint, re-reading the "on"
+# record; the watermark persisted with the catch-up above marks it as
+# already applied, so the offline change must survive.
+$standby->checksum_disable_offline;
+
+# Let the enable finish on the primary.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-checkpoint');");
+$primary->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-checkpoint');");
+
+# The launcher exits only after the checkpoint it requests completes,
+# so a checkpoint record now follows the transition in WAL.
+$primary->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_catalog.pg_stat_activity "
+ . "WHERE backend_type = 'datachecksums launcher';");
+
+my $log_offset = -s $standby->logfile;
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+
+# The nodes legitimately diverged, which replayed checkpoints report.
+$standby->wait_for_log(
+ qr/data checksum state "off" of this node does not match the state "on" in the replayed WAL/,
+ $log_offset);
+
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/020_cascade_divergence.pl b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
new file mode 100644
index 00000000000..8f2816ac405
--- /dev/null
+++ b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
@@ -0,0 +1,127 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# An offline data checksum state change made on one node of a cascading setup
+# is reported all the way down the chain.
+#
+# CheckReplayedDataChecksumState() is reached from xlog_redo() when a
+# checkpoint record (XLOG_CHECKPOINT_SHUTDOWN or XLOG_CHECKPOINT_REDO) is
+# replayed after consistency has been reached. Those records are written by
+# the root primary only and are relayed verbatim by every intermediate
+# standby, so a cascaded standby that was changed offline learns about the
+# divergence too. An intermediate standby's own restartpoints are not
+# WAL-logged and cannot surface it.
+#
+# Note that this makes detection checkpoint-driven: an idle or read-only
+# primary surfaces the divergence no sooner than its next checkpoint, so up to
+# checkpoint_timeout may pass before the operator sees anything. Closing that
+# window would require comparing the state when a walreceiver connects.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+# checkpoint_timeout is deliberately long so that the only checkpoint record in
+# this test is the one requested explicitly below.
+my $conf = qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+);
+
+my $primary = PostgreSQL::Test::Cluster->new('cascade_primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', $conf);
+$primary->start;
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,1000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('cascade_standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf('postgresql.conf', $conf);
+$standby->start;
+
+$standby->backup('backup');
+my $cascade = PostgreSQL::Test::Cluster->new('cascade_cascade');
+$cascade->init_from_backup($standby, 'backup', has_streaming => 1);
+$cascade->append_conf('postgresql.conf', $conf);
+$cascade->start;
+
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+test_checksum_state($cascade, 'off');
+
+# Change the cascaded standby offline, and only it.
+$cascade->stop;
+$cascade->command_ok([ 'pg_checksums', '--enable', '-D', $cascade->data_dir ],
+ 'pg_checksums enables checksums on the cascaded standby only');
+
+my $logstart = -s $cascade->logfile;
+$cascade->start;
+
+test_checksum_state($cascade, 'on');
+
+# Ordinary WAL traffic carries no state, so the cascaded standby keeps
+# streaming from a chain whose state is "off" without noticing.
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (1);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+is($cascade->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '1001', 'cascaded standby keeps streaming from a divergent chain');
+
+# Restartpoints on the intermediate standby are not WAL-logged, so this does
+# not surface anything on the cascaded standby either.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (2);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+my $log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+unlike(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'an intermediate restartpoint does not surface the divergence');
+
+# A checkpoint on the root primary does: the record is relayed down the whole
+# chain, so both the direct and the cascaded standby compare it against their
+# own state. wait_for_catchup() below waits for replay, not just receipt, so
+# the record has been through xlog_redo() on both nodes once it returns.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+$log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+like(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'the cascaded standby reports the divergence on the primary checkpoint');
+
+$cascade->stop;
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
new file mode 100644
index 00000000000..3f660623bd3
--- /dev/null
+++ b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
@@ -0,0 +1,205 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind across online data checksum transitions on both sides of a
+# divergence, with the two nodes trading roles between the scenarios.
+#
+# Scenario 1: after a switchover the new primary enables checksums
+# online, while the old primary restarts on its old timeline, advances
+# its WAL beyond the enable records and runs an online enable/disable
+# cycle of its own. The old primary ends "off" with a checksum
+# watermark numerically above every checksum record the new primary has
+# written. pg_rewind clamps the watermark it keeps to the divergence
+# point; without the clamp, replay on the rewound node would skip the
+# source's enable as already applied and stay "off" under an "on"
+# primary.
+#
+# Scenario 2: checksums are disabled again, and after another
+# switchover both nodes enable them online independently, so the
+# divergence checkpoint carries "off" while both control files say
+# "on". The rewind is allowed, the target keeps its own "on" state,
+# and replay re-walks the source's enable onto the already enabled
+# node.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum_extended' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum_extended\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+sub controldata_watermark
+{
+ my ($node) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $node->data_dir ]);
+ $stdout =~ /^Data checksum watermark:\s*([0-9A-F]+)\/([0-9A-F]+)$/m
+ or die "watermark missing from pg_controldata output";
+ return (hex($1) << 32) + hex($2);
+}
+
+# Wait until the standby has replayed the shutdown checkpoint of the
+# stopped primary, so that a subsequent promotion diverges after it and
+# the shutdown checkpoint becomes the last common checkpoint.
+sub wait_for_shutdown_checkpoint_replay
+{
+ my ($primary, $standby) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $primary->data_dir ]);
+ $stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+ my $shutdown_checkpoint = $1;
+
+ $standby->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_checkpoint'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+}
+
+# Old primary, checksums off. wal_log_hints is required by pg_rewind
+# on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+# Clean switchover to B; enable checksums online on it.
+$node_a->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_a, $node_b);
+$node_b->promote;
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+$node_b->stop('fast');
+my $watermark_b = controldata_watermark($node_b);
+$node_b->start;
+
+# Accidental restart of the old primary on the old timeline. Advance
+# its WAL beyond the enable watermark of B, then run an online enable
+# and disable cycle: the node ends "off" with a watermark above every
+# checksum record B has written.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t_pad AS SELECT generate_series(1,200000) AS a;");
+enable_data_checksums($node_a, wait => 'on');
+disable_data_checksums($node_a, wait => 'off');
+test_checksum_state($node_a, 'off');
+$node_a->stop('fast');
+
+my $watermark_a = controldata_watermark($node_a);
+die "test broken: target watermark not above the source's enable"
+ unless $watermark_a > $watermark_b;
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the promoted node');
+
+# Start the rewound node as a standby of B. Replay from the last
+# common checkpoint runs through B's online enable, which the target
+# never saw, so the rewound node must converge to "on".
+my $connstr_b = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr_b application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+# Scenario 2, reusing the pair with the roles reversed. Disable
+# checksums online so the next divergence point carries "off", and let
+# A replay the change.
+disable_data_checksums($node_b, wait => 'off');
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'off');
+
+# Clean switchover back to A; enable checksums online on it.
+$node_b->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_b, $node_a);
+$node_a->promote;
+enable_data_checksums($node_a, wait => 'on');
+test_checksum_state($node_a, 'on');
+my $source_enable_watermark = controldata_watermark($node_a);
+
+# The old primary restarts on its old timeline and enables checksums
+# online independently: both control files say "on", the divergence
+# checkpoint says "off".
+$node_b->start;
+$node_b->safe_psql('postgres', "INSERT INTO t_div VALUES (2);");
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+$node_b->stop('fast');
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_b->data_dir,
+ '--source-server' => $node_a->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled on both nodes');
+
+my $connstr_a = $node_a->connstr;
+$node_b->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_b->port]}
+primary_conninfo = '$connstr_a application_name=@{[$node_b->name]}'
+]);
+$node_b->set_standby_mode;
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_b, 'on');
+
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_b->stop('fast');
+is(controldata_watermark($node_b), $source_enable_watermark,
+ 'rewound node replayed the source checksum transition');
+command_ok([ 'pg_checksums', '--check', '-D', $node_b->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the twice-rewound node');
+
+done_testing();
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v13-0002-pg_checksums-Refuse-interrupted-transitions-note.patch (7.9K, ../../97139167-3B7B-4B00-B623-2DD5C0656DDB@yesql.se/3-v13-0002-pg_checksums-Refuse-interrupted-transitions-note.patch)
download | inline diff:
From 5dbdafcd2a03fe553ba691d02b5bebc53390ac08 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 3 Sep 2026 23:31:33 +0200
Subject: [PATCH v13 2/4] pg_checksums: Refuse interrupted transitions, note
change is local
An inprogress-on or inprogress-off control file means an online
transition was cut short; running the offline tool on top of it mixes
two procedures. Refuse it in every mode, with a hint pointing at the
start-stop cycle (or, on a standby, at letting replication finish)
that resets the state. Also print that the change applies to one
data directory only, with standby-aware wording, since in a
replication setup the same change must be made on every node.
On a primary, inprogress-on cannot survive a graceful stop: the
checksums launcher resolves it from its exit cleanup, and a crashed
primary is already rejected by the existing "cluster must be shut
down" check. inprogress-off can, however: it is set by the backend
running pg_disable_data_checksums(), and a fast shutdown arriving
between its two barriers leaves it behind in a cleanly shut down
control file, where the next start-stop cycle resolves it at end of
recovery. A standby can be stopped with either state, since it has
no launcher and only carries forward whatever state the last replayed
WAL record left it in, with a restartpoint persisting that as-is.
The standby is also the deterministic way to reach the new guard,
which is why the test coverage uses one; the primary window would
need an injection point between the two barriers.
The pg_checksums page said the tool still processes all relation files
regardless of an interrupted online transition; describe the refusal
instead.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Reviewed-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Discussion: https://postgr.es/m/anwm6UPxoVS41QA2@bdtpg
---
doc/src/sgml/ref/pg_checksums.sgml | 11 ++++---
src/bin/pg_checksums/pg_checksums.c | 21 +++++++++++++
.../test_checksums/t/012_offline_standby.pl | 31 ++++++++++++++-----
3 files changed, 51 insertions(+), 12 deletions(-)
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index 2d2057a8c19..abc035e5409 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -49,11 +49,12 @@ PostgreSQL documentation
</para>
<para>
- When enabling checksums with <application>pg_checksums</application>, if
- checksums were in the process of being enabled using
- <xref linkend="checksums-online-enable-disable"/> when the cluster was shut
- down, <application>pg_checksums</application> will still process all
- relation files regardless of the progress of online checksum processing.
+ If checksums were in the process of being enabled or disabled using
+ <xref linkend="checksums-online-enable-disable"/> when the cluster was
+ shut down, the control file still records that interrupted state, and
+ <application>pg_checksums</application> refuses to run in any mode.
+ Start the cluster and shut it down cleanly to reset the state, then
+ retry; on a standby, let replication complete the transition first.
</para>
<para>
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 5d6ea318784..3b58c6ca608 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -586,6 +586,21 @@ main(int argc, char *argv[])
ControlFile->state != DB_SHUTDOWNED_IN_RECOVERY)
pg_fatal("cluster must be shut down");
+ /*
+ * An inprogress state means an online transition was cut short. A
+ * standby stopped mid-transition carries either state; a cleanly shut
+ * down primary can still carry inprogress-off, which a fast shutdown
+ * during pg_disable_data_checksums() leaves behind, while inprogress-on
+ * is always resolved by the launcher's exit cleanup.
+ */
+ if (ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted");
+ pg_log_error_hint("Start and cleanly shut down the cluster once to reset the data checksum state, then retry. On a standby, let replication complete the transition first.");
+ exit(1);
+ }
+
if (ControlFile->data_checksum_version != PG_DATA_CHECKSUM_VERSION &&
mode == PG_MODE_CHECK)
pg_fatal("data checksums are not enabled in cluster");
@@ -673,6 +688,12 @@ main(int argc, char *argv[])
printf(_("Checksums enabled in cluster\n"));
else
printf(_("Checksums disabled in cluster\n"));
+
+ printf(_("This change applies to this data directory only.\n"));
+ if (ControlFile->state == DB_SHUTDOWNED_IN_RECOVERY)
+ printf(_("This node appears to be a standby; apply the same change to the primary and all other standbys.\n"));
+ else
+ printf(_("In a replication setup, apply the same change to every node while all are stopped, before restarting any of them.\n"));
}
return 0;
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
index c0ec374c8be..ba1f14e0399 100644
--- a/src/test/modules/test_checksums/t/012_offline_standby.pl
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -164,7 +164,10 @@ $newnode->stop('immediate');
# Converge the cluster: enable offline on the standby too.
$standby->stop;
-$standby->checksum_enable_offline;
+command_checks_all(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ 0, [qr/appears to be a standby/],
+ [], 'standby-role notice on offline enable');
$standby->start;
test_checksum_state($standby, 'on');
$primary->wait_for_catchup($standby);
@@ -232,11 +235,11 @@ unlike(
);
# Scenario 4: a standby stopped while replaying an interrupted online
-# transition keeps the interrupted state in its own control file. A
-# primary is never caught this way, as its checksums launcher resolves
-# inprogress-on back to off from its exit cleanup. A standby has no
-# launcher; it carries forward whatever the last replayed record left
-# it in.
+# transition keeps the interrupted state in its own control file, and
+# pg_checksums must refuse to touch it. A primary is never caught this
+# way, as its checksums launcher resolves inprogress-on back to off from
+# its exit cleanup. A standby has no launcher; it carries forward
+# whatever the last replayed record left it in.
# Block an online enable on the primary at inprogress-on with a
# blocking temp table, same trick as in 004_offline.pl.
@@ -250,6 +253,20 @@ wait_for_checksum_state($standby, 'inprogress-on');
# Stop the standby cleanly; its restartpoint persists inprogress-on to
# its own control file, since nothing on a standby resolves it away.
$standby->stop;
+
+command_fails_like(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --enable refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --check refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--disable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --disable refuses a standby stopped mid-transition');
+
$standby->start;
wait_for_checksum_state($standby, 'inprogress-on');
@@ -272,7 +289,7 @@ wait_for_checksum_state($primary, 'on');
$primary->wait_for_catchup($standby);
wait_for_checksum_state($standby, 'on');
-is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
'10001', 'standby readable once the transition completes');
# Scenario 4 continued: the backup taken mid-transition must also
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v13-0003-pg_rewind-Check-the-data-checksum-states-of-sour.patch (22.5K, ../../97139167-3B7B-4B00-B623-2DD5C0656DDB@yesql.se/4-v13-0003-pg_rewind-Check-the-data-checksum-states-of-sour.patch)
download | inline diff:
From 0f07c58d03bd761fef6a1f45cc0e486ec0f805e0 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 3 Sep 2026 23:31:52 +0200
Subject: [PATCH v13 3/4] pg_rewind: Check the data checksum states of source
and target
pg_rewind installs the source's control file on the target, but every
block it does not copy keeps the target's content. With checksums
enabled on the target and disabled on the source, the rewound server
claims enabled checksums while the blocks copied from the source have
none, and fails checksum verification as soon as it reads them; in the
worst case every connection attempt dies on an unverifiable catalog
page.
Comparing the control files is not enough. Replay on the rewound
server resumes from the last common checkpoint and adopts the data
checksum state recorded there, so an offline disable on the source
after the divergence leaves both control files saying "off" while the
rewound server still resumes with verification enabled. Check the
state carried by the divergence checkpoint as well, which pg_rewind
already reads.
Refuse both cases, and an interrupted online transition on either
side, same as pg_checksums does. The opposite mismatch, checksums on
the source only, is allowed with a warning: recovery under the backup
label adopts the divergence state, so the rewound server keeps
checksums disabled (or converges through the replayed WAL if the
source enabled them online), and no page can fail verification.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Reviewed-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Discussion: https://postgr.es/m/anwm6UPxoVS41QA2@bdtpg
---
doc/src/sgml/ref/pg_rewind.sgml | 14 ++
src/bin/pg_rewind/parsexlog.c | 4 +-
src/bin/pg_rewind/pg_rewind.c | 93 ++++++++++-
src/bin/pg_rewind/pg_rewind.h | 1 +
src/test/modules/test_checksums/meson.build | 2 +
.../test_checksums/t/022_rewind_state.pl | 145 +++++++++++++++++
.../t/023_rewind_standby_target.pl | 153 ++++++++++++++++++
7 files changed, 410 insertions(+), 2 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/022_rewind_state.pl
create mode 100644 src/test/modules/test_checksums/t/023_rewind_standby_target.pl
diff --git a/doc/src/sgml/ref/pg_rewind.sgml b/doc/src/sgml/ref/pg_rewind.sgml
index b95e3868e9d..ac3d0c9328f 100644
--- a/doc/src/sgml/ref/pg_rewind.sgml
+++ b/doc/src/sgml/ref/pg_rewind.sgml
@@ -124,6 +124,20 @@ PostgreSQL documentation
<literal>on</literal>, but is enabled by default.
</para>
+ <para>
+ The data checksum states of the source and target server must be
+ compatible. <application>pg_rewind</application> refuses to run if an
+ online checksum state transition was interrupted on either server, or
+ if data checksums are enabled on the target server, or were enabled at
+ the point of divergence, while the source server runs without them: the
+ rewound server would fail checksum verification on the blocks copied
+ from the source. The opposite combination is allowed with a warning;
+ the rewound server keeps data checksums disabled unless the source
+ server enabled them online after the point of divergence. See
+ <xref linkend="app-pgchecksums"/> for how to change the state
+ consistently across a replication setup.
+ </para>
+
<warning>
<title>Warning: Failures While Rewinding</title>
<para>
diff --git a/src/bin/pg_rewind/parsexlog.c b/src/bin/pg_rewind/parsexlog.c
index 023e23b063c..6e87b00f8c2 100644
--- a/src/bin/pg_rewind/parsexlog.c
+++ b/src/bin/pg_rewind/parsexlog.c
@@ -167,7 +167,8 @@ readOneRecord(const char *datadir, XLogRecPtr ptr, int tliIndex,
void
findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
- XLogRecPtr *lastchkptredo, const char *restoreCommand)
+ XLogRecPtr *lastchkptredo, uint32 *lastchkptdatachecksums,
+ const char *restoreCommand)
{
/* Walk backwards, starting from the given record */
XLogRecord *record;
@@ -255,6 +256,7 @@ findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
*lastchkptrec = searchptr;
*lastchkpttli = checkPoint.ThisTimeLineID;
*lastchkptredo = checkPoint.redo;
+ *lastchkptdatachecksums = checkPoint.dataChecksumState;
break;
}
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 0e4c2df3f4b..d2521dab333 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -145,6 +145,7 @@ main(int argc, char **argv)
XLogRecPtr chkptrec;
TimeLineID chkpttli;
XLogRecPtr chkptredo;
+ uint32 chkptdatachecksums;
TimeLineID source_tli;
TimeLineID target_tli;
XLogRecPtr target_wal_endrec;
@@ -472,10 +473,51 @@ main(int argc, char **argv)
keepwal_init();
findLastCheckpoint(datadir_target, divergerec, lastcommontliIndex,
- &chkptrec, &chkpttli, &chkptredo, restore_command);
+ &chkptrec, &chkpttli, &chkptredo, &chkptdatachecksums,
+ restore_command);
pg_log_info("rewinding from last common checkpoint at %X/%08X on timeline %u",
LSN_FORMAT_ARGS(chkptrec), chkpttli);
+ /*
+ * Replay on the rewound server resumes from the last common checkpoint
+ * and adopts the data checksum state recorded there, not the state in the
+ * control file installed from the source. If checksums were enabled at
+ * the divergence point but the source runs without them, the rewound
+ * server would verify checksums while the blocks copied from the source
+ * have none. The control files cannot reveal this: an offline disable on
+ * the source after the divergence leaves both of them saying "off".
+ *
+ * Only the fully enabled state needs checking. An in-progress state at
+ * the divergence point means a transition was still running there, and
+ * all of its page rewrites are logged after that point, so replay brings
+ * the rewound server to whatever state the copied WAL ends in.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the source server or recreate the target server from a base backup.");
+ exit(1);
+ }
+
+ /*
+ * The same comparison is needed against the target: every block the
+ * rewind does not copy keeps the target's content. The control files
+ * cannot reveal this case either, in the other direction: a standby's
+ * checkpoints are written by its upstream primary, so a standby whose
+ * checksums were disabled offline still has "on" checkpoints in its WAL,
+ * while both control files may agree.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_target.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the target server");
+ pg_log_error_detail("Blocks kept from the target would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the target server with pg_checksums or recreate the target server from a base backup.");
+ exit(1);
+ }
+
/* Initialize the hash table to track the status of each file */
filehash_init();
@@ -799,6 +841,55 @@ sanityChecks(void)
pg_fatal("target server needs to use either data checksums or \"wal_log_hints = on\"");
}
+ /*
+ * The rewound target keeps its control file fields from the source, but
+ * every block the rewind does not copy keeps the target's content. If
+ * checksums are enabled on the target and disabled on the source, the
+ * result would claim enabled checksums while the blocks copied from the
+ * source have none, and the target would fail checksum verification as
+ * soon as it reads them. Refuse that combination, and refuse an
+ * interrupted online transition on either side, same as pg_checksums.
+ *
+ * The opposite mismatch is allowed: recovery under the backup label
+ * written by pg_rewind adopts the checksum state as of the divergence
+ * point, so a target whose own checkpoints carry no checksums stays
+ * without them, or converges through the replayed WAL if the source
+ * enabled them online. Only warn about it, so that an offline change on
+ * the source is not overlooked.
+ *
+ * That reasoning does not hold when the target is a standby, whose
+ * checkpoints were written by its upstream primary; see the checks after
+ * findLastCheckpoint().
+ */
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted on the target server");
+ pg_log_error_hint("Start the server and shut it down cleanly to reset the state, then retry; on a standby, let replication complete the transition first.");
+ exit(1);
+ }
+ if (ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition is incomplete on the source server");
+ pg_log_error_hint("Let the transition complete, or reset the state with a clean restart, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_OFF)
+ {
+ pg_log_error("data checksums are enabled on the target server but disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, and the target would fail checksum verification after the rewind.");
+ pg_log_error_hint("Either disable data checksums on the target server or enable them on the source server, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_OFF &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_warning("data checksums are disabled on the target server but enabled on the source server");
+ pg_log_warning_detail("The rewound server will keep data checksums disabled unless the source server enabled them online after the point of divergence.");
+ }
+
/*
* Target cluster better not be running. This doesn't guard against
* someone starting the cluster concurrently. Also, this is probably more
diff --git a/src/bin/pg_rewind/pg_rewind.h b/src/bin/pg_rewind/pg_rewind.h
index 9a981f7f246..9187c3bf72f 100644
--- a/src/bin/pg_rewind/pg_rewind.h
+++ b/src/bin/pg_rewind/pg_rewind.h
@@ -39,6 +39,7 @@ extern void findLastCheckpoint(const char *datadir, XLogRecPtr forkptr,
int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
XLogRecPtr *lastchkptredo,
+ uint32 *lastchkptdatachecksums,
const char *restoreCommand);
extern XLogRecPtr readOneRecord(const char *datadir, XLogRecPtr ptr,
int tliIndex, const char *restoreCommand);
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index e5d38fafb7d..7d07c757052 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -45,6 +45,8 @@ tests += {
't/019_standby_shutdown_catchup.pl',
't/020_cascade_divergence.pl',
't/021_rewind_divergent_transitions.pl',
+ 't/022_rewind_state.pl',
+ 't/023_rewind_standby_target.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/022_rewind_state.pl b/src/test/modules/test_checksums/t/022_rewind_state.pl
new file mode 100644
index 00000000000..e6d9c645a0b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/022_rewind_state.pl
@@ -0,0 +1,145 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind checks the data checksum states of source and target.
+# Checksums enabled on the target, or at the point of divergence, with
+# a source running without them are refused: the rewound server would
+# verify checksums on blocks copied from a source that has none. The
+# opposite mismatch only warns; replay keeps the target's own state.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+# Scenarios 1 and 2: checksums on from initdb. wal_log_hints keeps the
+# target eligible for pg_rewind after checksums are disabled on it.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b);
+
+# Failover to B, and divergence on A.
+$node_b->promote;
+$node_b->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_a->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_a->stop('fast');
+
+# Offline disable on the source only: enabled target, disabled source.
+$node_b->stop;
+$node_b->checksum_disable_offline;
+$node_b->start;
+test_checksum_state($node_b, 'off');
+
+my @rewind_cmd = (
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'));
+
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums are enabled on the target server but disabled on the source server/,
+ 'refuses an enabled target with a disabled source');
+
+# The refusal happens before any modification: the target still starts
+# on its own timeline.
+$node_a->start;
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'target untouched by the refused rewind');
+$node_a->stop('fast');
+
+# Scenario 2: disabling the target too makes the control files match,
+# but checksums were still enabled at the point of divergence, which is
+# the state replay on the rewound server would resume with.
+$node_a->checksum_disable_offline;
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums were enabled at the point of divergence but are disabled on the source server/,
+ 'refuses when checksums were enabled at the point of divergence');
+
+# Scenario 3: fresh pair without checksums; offline enable on the
+# source only. Allowed with a warning, and the rewound server keeps
+# checksums disabled.
+my $node_c = PostgreSQL::Test::Cluster->new('node_c');
+$node_c->init(allows_streaming => 1, no_data_checksums => 1);
+$node_c->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_c->start;
+$node_c->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_c->backup('backup');
+my $node_d = PostgreSQL::Test::Cluster->new('node_d');
+$node_d->init_from_backup($node_c, 'backup', has_streaming => 1);
+$node_d->start;
+$node_c->wait_for_catchup($node_d);
+
+$node_d->promote;
+$node_d->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_c->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_c->stop('fast');
+
+$node_d->stop;
+$node_d->checksum_enable_offline;
+$node_d->start;
+test_checksum_state($node_d, 'on');
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_c->data_dir,
+ '--source-server' => $node_d->connstr('postgres'),
+ ]);
+like(
+ $stderr,
+ qr/data checksums are disabled on the target server but enabled on the source server/,
+ 'warns for a disabled target with an enabled source');
+like($stderr, qr/Done!/, 'rewind completed despite the warning');
+
+# The rewound server follows D and keeps its own state.
+$node_c->append_conf('postgresql.conf', 'port = ' . $node_c->port);
+$node_c->enable_streaming($node_d);
+$node_c->set_standby_mode;
+$node_c->start;
+$node_d->wait_for_catchup($node_c);
+is($node_c->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+test_checksum_state($node_c, 'off');
+
+$node_c->stop;
+$node_d->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/023_rewind_standby_target.pl b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
new file mode 100644
index 00000000000..9071114a9ae
--- /dev/null
+++ b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
@@ -0,0 +1,153 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_rewind compares the data checksum state at the point of
+# divergence against the target as well as the source.
+#
+# A standby's checkpoints are written by its upstream primary, so a standby
+# whose checksums were disabled offline still has "on" checkpoints in its
+# WAL. Rewinding it onto a promoted sibling passes every control file
+# check (target off + source on only warns), yet recovery after the rewind
+# adopts the divergence checkpoint's state and the rewound server would
+# verify checksums over the pages it wrote while it was locally "off".
+# pg_rewind must refuse; after an offline enable on the target, which
+# rewrites all of its pages, the rewind goes through.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum_extended' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum_extended\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t0 AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+
+my $lagging = PostgreSQL::Test::Cluster->new('lagging');
+$lagging->init_from_backup($primary, 'backup', has_streaming => 1);
+$lagging->start;
+
+my $failover = PostgreSQL::Test::Cluster->new('failover');
+$failover->init_from_backup($primary, 'backup', has_streaming => 1);
+$failover->start;
+
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+
+test_checksum_state($primary, 'on');
+test_checksum_state($lagging, 'on');
+test_checksum_state($failover, 'on');
+
+# Offline-disable checksums on the lagging standby only. This is a
+# divergence 012_offline_standby.pl declares survivable: the standby keeps
+# its own state, warns, and stays readable.
+$lagging->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $lagging->data_dir);
+$lagging->start;
+test_checksum_state($lagging, 'off');
+
+# Give the lagging standby plenty of pages to write while it is locally
+# "off", and force a restartpoint so they reach disk without checksums.
+# The other standby replays the same WAL with checksums on.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+$lagging->safe_psql('postgres', "CHECKPOINT;");
+
+# Take the "failover" standby out of the picture. It is promoted from
+# this position later, so everything replayed from here on is the part of
+# the WAL that diverges and that pg_rewind will roll back. t1 is already
+# replayed everywhere, so its blocks are *not* rolled back.
+$failover->stop;
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$primary->wait_for_catchup($lagging);
+
+# Failover: promote the other standby, which forks the timeline behind the
+# replay position of the lagging standby.
+$primary->stop('fast');
+$failover->start;
+$failover->promote;
+$failover->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$failover->safe_psql('postgres', "CHECKPOINT;");
+
+# Re-attaching the lagging standby with pg_rewind must be refused: the
+# divergence checkpoint says "on" while the target is "off", so the rewound
+# server would verify checksums over the blocks it keeps.
+$lagging->stop;
+
+command_fails_like(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ],
+ qr/data checksums were enabled at the point of divergence but are disabled on the target server/,
+ 'pg_rewind refuses a target whose divergence checkpoint has checksums');
+
+# Enabling checksums offline on the target rewrites all of its pages, after
+# which the rewind is safe.
+system_or_bail('pg_checksums', '--enable', '--pgdata', $lagging->data_dir);
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ]);
+like($stderr, qr/Done!/, 'pg_rewind completes after the offline enable');
+
+$lagging->append_conf('postgresql.conf', 'port = ' . $lagging->port);
+$lagging->enable_streaming($failover);
+$lagging->set_standby_mode;
+$lagging->start;
+
+my ($rc, $out, $err) =
+ $lagging->psql('postgres',
+ "SELECT setting FROM pg_settings " . "WHERE name = 'data_checksums';");
+is($rc, 0, 'rewound standby accepts connections') or diag("stderr: $err");
+is($out, 'on', 'rewound standby resumes with checksums on');
+
+($rc, $out, $err) = $lagging->psql('postgres', "SELECT count(*) FROM t1;");
+is($rc, 0, 'blocks kept from the target stay readable')
+ or diag("stderr: $err");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($lagging->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the rewound standby');
+
+$lagging->stop('immediate');
+$failover->stop;
+done_testing();
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v13-0004-pg_combinebackup-Refuse-mixed-data-checksum-stat.patch (10.0K, ../../97139167-3B7B-4B00-B623-2DD5C0656DDB@yesql.se/5-v13-0004-pg_combinebackup-Refuse-mixed-data-checksum-stat.patch)
download | inline diff:
From 40a344fe895e88b076c8940fb2fece9582393544 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 3 Sep 2026 23:32:04 +0200
Subject: [PATCH v13 4/4] pg_combinebackup: Refuse mixed data checksum states
in a backup chain
check_control_files() detected a chain whose backups were taken under
different data checksum states, warned, and proceeded. The output
directory keeps the last backup's control file, so a full backup taken
with checksums off combined with an incremental taken after an offline
enable produces a cluster whose control file says "on" while most of
its blocks carry no checksums; it starts, and then every connection
dies on the first unchecksummed catalog page. An offline enable
between two backups of a chain is all it takes, since it rewrites
every page without logging anything, so the incremental backup does
not re-ship the pages.
Turn the warning into an error, matching what pg_rewind does for the
equivalent combinations. The check stays asymmetric on purpose: when
the last backup was taken with checksums off, stale checksums from an
earlier backup are never verified, and that chain remains usable.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Reviewed-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Discussion: https://postgr.es/m/anwm6UPxoVS41QA2@bdtpg
---
doc/src/sgml/ref/pg_combinebackup.sgml | 19 +--
src/bin/pg_combinebackup/pg_combinebackup.c | 14 +-
src/test/modules/test_checksums/meson.build | 1 +
.../t/024_combinebackup_mixed.pl | 146 ++++++++++++++++++
4 files changed, 167 insertions(+), 13 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
diff --git a/doc/src/sgml/ref/pg_combinebackup.sgml b/doc/src/sgml/ref/pg_combinebackup.sgml
index 9a6d201e0b8..f5d7e177b36 100644
--- a/doc/src/sgml/ref/pg_combinebackup.sgml
+++ b/doc/src/sgml/ref/pg_combinebackup.sgml
@@ -306,18 +306,19 @@ PostgreSQL documentation
<para>
<literal>pg_combinebackup</literal> does not recompute page checksums when
- writing the output directory. Therefore, if any of the backups used for
- reconstruction were taken with checksums disabled, but the final backup was
- taken with checksums enabled, the resulting directory may contain pages
- with invalid checksums.
+ writing the output directory. It therefore refuses a chain in which some
+ of the backups used for reconstruction were taken with checksums disabled
+ but the final backup was taken with checksums enabled: the resulting
+ directory would contain pages with invalid checksums, and the cluster
+ would fail checksum verification as soon as it reads them. Take a new
+ full backup after enabling data checksums with
+ <xref linkend="app-pgchecksums"/>.
</para>
<para>
- To avoid this problem, taking a new full backup after changing the checksum
- state of the cluster using <xref linkend="app-pgchecksums"/> is
- recommended. Otherwise, you can disable and then optionally reenable
- checksums on the directory produced by <literal>pg_combinebackup</literal>
- in order to correct the problem.
+ The reverse case is accepted: when the final backup was taken with
+ checksums disabled, stale checksums copied from an earlier backup are
+ never verified.
</para>
</refsect1>
diff --git a/src/bin/pg_combinebackup/pg_combinebackup.c b/src/bin/pg_combinebackup/pg_combinebackup.c
index 86e2ee37c40..254a27b125b 100644
--- a/src/bin/pg_combinebackup/pg_combinebackup.c
+++ b/src/bin/pg_combinebackup/pg_combinebackup.c
@@ -670,13 +670,19 @@ check_control_files(int n_backups, char **backup_dirs)
pg_log_debug("system identifier is %" PRIu64, system_identifier);
/*
- * Warn the user if not all backups are in the same state with regards to
- * checksums.
+ * Reject a chain whose backups were not all in the same state with
+ * regards to checksums. The loop above only flags that when the last
+ * backup has checksums enabled: the blocks taken from the older backups
+ * have no checksums, and pg_combinebackup does not recompute them, so the
+ * combined cluster would fail verification as soon as it reads them. The
+ * other direction is harmless, since stale checksums are never verified
+ * while checksums are disabled.
*/
if (data_checksum_mismatch)
{
- pg_log_warning("only some backups have checksums enabled");
- pg_log_warning_hint("Disable, and optionally reenable, checksums on the output directory to avoid failures.");
+ pg_log_error("only some backups have checksums enabled");
+ pg_log_error_hint("Take a new full backup after changing the data checksum state with pg_checksums.");
+ exit(1);
}
return system_identifier;
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index 7d07c757052..7eccd5156b9 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -47,6 +47,7 @@ tests += {
't/021_rewind_divergent_transitions.pl',
't/022_rewind_state.pl',
't/023_rewind_standby_target.pl',
+ 't/024_combinebackup_mixed.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
new file mode 100644
index 00000000000..2f51c7fd9dc
--- /dev/null
+++ b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
@@ -0,0 +1,146 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_combinebackup refuses a chain whose backups were taken under
+# different data checksum states when the final state has them enabled.
+#
+# The output directory keeps the last backup's control file, so a full
+# backup taken with checksums off plus an incremental taken after an
+# offline enable would produce a directory whose control file says "on"
+# while most of its blocks have no checksums. An offline enable between
+# the two backups is enough to get there: it rewrites every page but logs
+# nothing, so the incremental backup does not re-ship the pages.
+#
+# The reverse order stays allowed: with the final backup taken with
+# checksums off, stale checksums from an earlier backup are never verified.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('node');
+$node->init(
+ no_data_checksums => 1,
+ has_archiving => 1,
+ allows_streaming => 1);
+$node->append_conf('postgresql.conf', 'summarize_wal = on');
+$node->append_conf('postgresql.conf', 'autovacuum = off');
+$node->start;
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+test_checksum_state($node, 'off');
+
+$node->backup('full');
+
+# Offline enable: rewrites every page, logs nothing.
+$node->stop;
+system_or_bail('pg_checksums', '--enable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'on');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full/backup_manifest',
+ ],
+ 'incremental backup with checksums on');
+
+# Combining the chain must be refused: the result would say "on" while the
+# blocks inherited from the full backup have no checksums.
+command_fails_like(
+ [
+ 'pg_combinebackup',
+ $node->backup_dir . '/full',
+ $node->backup_dir . '/incr',
+ '--output' => $node->backup_dir . '/combined',
+ ],
+ qr/only some backups have checksums enabled/,
+ 'pg_combinebackup refuses a chain crossing an offline enable');
+
+# The other direction: full backup with checksums on, offline disable, then
+# an incremental. The last backup wins, so the combined cluster comes up off.
+$node->backup('full2');
+
+$node->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr2',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full2/backup_manifest',
+ ],
+ 'incremental backup with checksums off');
+
+my $restored = PostgreSQL::Test::Cluster->new('restored');
+$restored->init_from_backup(
+ $node, 'incr2',
+ combine_with_prior => ['full2'],
+ has_restoring => 1,
+ standby => 0);
+
+$restored->start;
+
+my ($rc, $stdout, $stderr) = $restored->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'combined backup accepts connections') or diag("stderr: $stderr");
+is($stdout, 'off', 'combined backup follows the last backup state');
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t3;');
+is($rc, 0, 'blocks from the incremental backup are readable')
+ or diag("stderr: $stderr");
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t1;');
+is($rc, 0, 'blocks inherited from the full backup are readable')
+ or diag("stderr: $stderr");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($restored->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures in the combined backup');
+
+$restored->stop('immediate');
+$node->stop;
+
+done_testing();
--
2.39.3 (Apple Git-146)
=
[text/plain] commits.txt (5.4K, ../../97139167-3B7B-4B00-B623-2DD5C0656DDB@yesql.se/6-commits.txt)
download | inline:
Bugs
----
* 20260406: d771b0a907e: Handle checksumworker startup wait race
Handle the race condition that data checksums are fully enabled and the worker
finishes before the launcher starts waiting for it.
* 20260430: 25b922ec582: Fix invalid checksum state transition in checkpoints
There was a race condition around updating the controlfile leading to a backend
starting at exactly the wrong time reading the previous state. A related issue
was if multiple checksum state changes happened between a backend initializing
procsignals and finishing XLOG startup. Somewhat hard to hit since it requires
a backend magnitudes slower than other background workers in the same cluster.
* 20260430: 8fb8ded8895: Handle data_checksum state changes during launcher_exit
Silly bug, the state transition for reverting from inprogress-on to off via
inprogress-off in one error path was missing the state machince allowed state
table.
* 20260430: b120358c612: Prevent pg_enable/disable_data_checksums() on standby
The SQL functions for disabling and enabling checksums could be executed on a
hot standby.
* 20260506: 9a39056c418: Apply data-checksum worker throttling parameters
The API for for cost parameters had changed between when I wrote the code
(years ago) and now and I had missed that, so the cost settings were not
properly applied.
* 20260529: 5fee7cab1b8: Fix checksum state transition during promotion
When a primary crashed during checksum enabling, the standby didn't emit the
procsignalbarrier during promotion to revert the cluster state to off.
* 20260804: 01805b7d16b: Do not reuse rd_smgr in fork loop when enabling data checksums
The worker incorrectly accessed rd_smgr which can be set to NULL at a relcache
invalidation.
* 20260817: 3a18526e8d6: Make data checksums launcher cancel its worker at SIGINT
The launcher waited for the worker to complete before terminating at SIGINT
which cause pointless work and can take some time. Fix to make it faster by
having the launcher terminate and worker at SIGINT and then exit itself.
* 20260430: 1df361e3d82: Improve database detection logic in datachecksumsworker
* 20260804: 343d98c3601: Don't skip invalid databases when enabling data checksums
Commit 1df361e3d82 attempted to improve the logic for handling invalid
databases, later fixed in 343d98c3601 to handle databases which have had a DROP
DATABASE operation crash.
* 20260828: e469e4784ea: Handle invalid and dropped databases during checksum enable
Detect invalid databases earlier, and also detect when a database is dropped
while checksums are being enabled in that very database, previously this case
resulted in a process error with the state rolled back to off.
* 20260818: 0907112d388: basebackup: do not verify checksums on pages from before enabling
A base backup which started before a checksum state transition, or when
checksums are disabled and re-enabled while the backup is running, would
produce false negatives.
Improvements and Optimizations
------------------------------
* 20260430: bf25e5571b3: Improve handling of concurrent checksum requests
Made the logic for handling concurrent invocations of enabling or disabling
checksums not require creating a new launcher for detection, which reduce
cluster process usage.
* 20260506: 2018bd61679: Skip WAL for unlogged main fork during online checksum enable
Pages for unlogged relations were not exempt from processing which caused
needless WAL traffic. All pages were still properly checksummed, it "just"
generated too much WAL traffic.
* 20260818: 397f0fd06ed: Add data_page_checksum_version to pg_control_checkpoint
pg_controldata shows the checksum version from the checkpoint but system view
pg_control_checkpoint omitted it.
Minor Source Code changes
-------------------------
* 20260406: b3a37ffbc5b: Use PG_DATA_CHECKSUM_OFF instead of hardcoded value
Replace 0 with the PG_DATA_CHECKSUM_OFF label in the code. Readability
improvement over using the hardcoded zero which has been used since forever.
* 20260529: 0ca1b301059: Use correct datatype for PID
Use pid_t instead of int.
* 20260817: abac86c7a27: Reorder function prototypes to match definition order
Clearly not required, but there was an ask to do it in HEAD and backpatching
seemed a cheap insurance against backpatching conflicts.
Pre-existing bugs
-----------------
* 20260818: aaf8b9989f7: Record initial state of data checksums in controlfile
Fixes a longstanding bug in offline checksums where pg_control_init is defined
to list the initial cluster state, but pg_checksums just overwrites and
effectively made it show the current state.
Test Stabilization
-------------------
* 20260406: 07009121c23: Test stabilization for online checksums
* 20260901: 5313fc415d9: Wait for checksum state transition in test
Documentation, Error messages and Code comments etc
---------------------------------------------------
* 20260408: b364828f825: doc: Fix data_checksums data type
* 20260430: 381d19da153: Typo and spelling fixups for online checksums
* 20260529: cd857dec0e0: Improve comments in online checksums code
* 20260529: 5ab239c9a90: Constistent naming for datacheckusms processes
* 20260605: 4ae3e98c02c: doc: Mention online checksum enabling in pg_checksums docs
* 20260605: e5e1f6dc795: Reword activity message to avoid truncation
* 20260618: 8d22f523245: Fix comments on data checksum cost settings
* 20260716: e3a27cad462: doc: Fix link text for data checksums
* 20260801: 602f19c84ca: doc: Fix glossary entry for data checksums workers
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-07 11:08 Heikki Linnakangas <hlinnaka@iki.fi>
parent: Daniel Gustafsson <daniel@yesql.se>
0 siblings, 2 replies; 43+ messages in thread
From: Heikki Linnakangas @ 2026-09-07 11:08 UTC (permalink / raw)
To: Daniel Gustafsson <daniel@yesql.se>; Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-hackers@lists.postgresql.org
On 04/09/2026 02:08, Daniel Gustafsson wrote:
>> On 3 Sep 2026, at 13:54, Bertrand Drouvot <bertranddrouvot.pg@gmail.com> wrote:
>> On Thu, Sep 03, 2026 at 12:06:42PM +0100, Zsolt Parragi wrote:
>
>>> v12 addresses these, otherwise it is unchanged to compared 11.
>>
>> Thanks! v12 LGTM.
>
> Thanks for review. I've attached a v13 where I've moved most of the new tests
> under PG_TEST_EXTRA to keep test times down. I placed most tests under
> 'checksum' and 18, 21 and 23 under 'checksum_extended', but the exact split may
> be tweaked further. Since the origin of this open item is missing test
> coverage, I prefer to add all these tests even though they aren't executed
> during normal testruns. There are at least one BF animal running the full
> suite which ensures timely execution of the tests.
>
> This concludes the only open item left (thus far). Being able to error
> standbys out of mismatched clusters would be nice, and is a potential
> development area for 20, but it's not a showstopper if we never add it IMHO.
>
> My current plan is to commit this to master only either tomorrow or Monday
> after staring at it a little bit more, to a) give it exposure in the buildfarm
> before an eventual backpatching; b) allow time for the revert discussion. If
> we decide to revert I prefer to avoid more v19 churn.
Thanks, I started to review this now. I'm still at patch 0001, haven't
looked at the rest yet, but some quick comments on that one:
> diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
> index ec62d17fbcc..dbbfdae5e4f 100644
> --- a/doc/src/sgml/wal.sgml
> +++ b/doc/src/sgml/wal.sgml
> @@ -317,6 +317,29 @@
> verify checksums, on an offline cluster.
> </para>
>
> + <para>
> + An offline change provides durability differently from an
> + <link linkend="checksums-online-enable-disable">online change</link>.
> + An online transition is WAL-logged: it is ordered against all other
> + WAL records, it is replayed after a crash, and it propagates to
> + standbys. An offline change is recorded only in the cluster's
> + control file: it writes no WAL, it is invisible to replication, and
> + it has no defined ordering against WAL the node has not replayed
> + yet. When a node later replays WAL that contains an online checksum
> + state change, that change takes effect on the node even if it was
> + written before the offline change was made.
> + </para>
> +
> + <para>
> + An offline change only affects the data directory it is run on; the
> + new state does not propagate over replication. In a replication setup
> + the same change must be applied to all nodes while all of them are
> + stopped, as described in <xref linkend="app-pgchecksums"/>. A standby
> + whose state diverges logs a warning but keeps its local setting. The
> + mismatch persists until the states are brought together again, with
> + the offline procedure or with an online transition; do this promptly.
> + </para>
> +
> </sect2>
>
> <sect2 id="checksums-online-enable-disable" xreflabel="Online Enabling of Checksums">
Let's add a new 'sect2' for this explanation, and move it after the
"Online Enabling of Checksums" section. It's currently placed under
"Offline Enabling of Checksums", but it actually goes into a lot of
details of how *online* checksumming works, but "Online Enabling of
Checksums" is covered in the following paragraph. If you read this in
order like a novel, it feels weird.
I think these paragraphs could use some copy-editing too. It feels like
a pretty deep technical explanation, not very accessible to a DBA. Maybe
start with "The primary server and replica can have different checksum
states".
(Not new with this patch, but: )
The placement of the states in the state diagram on that page looks
bizarre. I know it's auto-generated so not sure there's much we can do
about it.. but could we, please? Maybe it'd get more clear if you leave
'initdb' out of the diagram. Or consider some completely different
representation.
I'm still trying to understand all the different states and interactions
between online and offline changes. It's really complicated :-(. I know
it's a tall order, but is there something we could do to make it
simpler? Would it help if there was a separate flag in the control file
for "checksums enabled in primary" and "checksums enabled in this
replica", for example?
- Heikki
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-07 12:26 Daniel Gustafsson <daniel@yesql.se>
parent: Heikki Linnakangas <hlinnaka@iki.fi>
1 sibling, 1 reply; 43+ messages in thread
From: Daniel Gustafsson @ 2026-09-07 12:26 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-hackers@lists.postgresql.org
> On 7 Sep 2026, at 13:08, Heikki Linnakangas <hlinnaka@iki.fi> wrote:
> Thanks, I started to review this now. I'm still at patch 0001, haven't looked at the rest yet, but some quick comments on that one:
Thanks for reviewing!
> Let's add a new 'sect2' for this explanation, and move it after the "Online Enabling of Checksums" section. It's currently placed under "Offline Enabling of Checksums", but it actually goes into a lot of details of how *online* checksumming works, but "Online Enabling of Checksums" is covered in the following paragraph. If you read this in order like a novel, it feels weird.
>
> I think these paragraphs could use some copy-editing too. It feels like a pretty deep technical explanation, not very accessible to a DBA. Maybe start with "The primary server and replica can have different checksum states".
>
> (Not new with this patch, but: )
Can, but really really shouldn't =) I'll try to rework the documentation here
to make it less dense.
> The placement of the states in the state diagram on that page looks bizarre. I know it's auto-generated so not sure there's much we can do about it.. but could we, please? Maybe it'd get more clear if you leave 'initdb' out of the diagram. Or consider some completely different representation.
I can try, maybe breaking it up into multiple diagrams could help?
> I'm still trying to understand all the different states and interactions between online and offline changes. It's really complicated :-(. I know it's a tall order, but is there something we could do to make it simpler?
If there was I'd love to try it, but across the many alteratives tried during
this open item there hasn't been anyhing less complicated which also solves the
problem. Combining a WAL logged procedure with one that can rewrite the data
directory without any WAL entries at all is inherently complicated.
> Would it help if there was a separate flag in the control file for "checksums enabled in primary" and "checksums enabled in this replica", for example?
Not sure I follow, should pg_checksums maintan such a flag or the StartupXLOG?
--
Daniel Gustafsson
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-07 15:57 Heikki Linnakangas <hlinnaka@iki.fi>
parent: Daniel Gustafsson <daniel@yesql.se>
0 siblings, 1 reply; 43+ messages in thread
From: Heikki Linnakangas @ 2026-09-07 15:57 UTC (permalink / raw)
To: Daniel Gustafsson <daniel@yesql.se>; +Cc: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-hackers@lists.postgresql.org
On 07/09/2026 15:26, Daniel Gustafsson wrote:
>> On 7 Sep 2026, at 13:08, Heikki Linnakangas <hlinnaka@iki.fi> wrote:
>> I think these paragraphs could use some copy-editing too. It feels like a pretty deep technical explanation, not very accessible to a DBA. Maybe start with "The primary server and replica can have different checksum states".
>
> Can, but really really shouldn't =)
:-D
>> I'm still trying to understand all the different states and interactions between online and offline changes. It's really complicated :-(. I know it's a tall order, but is there something we could do to make it simpler?
>
> If there was I'd love to try it, but across the many alteratives tried during
> this open item there hasn't been anyhing less complicated which also solves the
> problem. Combining a WAL logged procedure with one that can rewrite the data
> directory without any WAL entries at all is inherently complicated.
So, if you can have different state in primary and a replica, there are
four combinations:
1: Primary on, replica on
2: Primary off, replica off
These are straightforward
3: Primary on, replica off
You end up in this situation, if you turn on run pg_checksums to turn on
checksums in primary. You stay that you really really shouldn't stay for
long in this state, but why? What's the harm?
One harm is that it's confusing, but if we have to deal with it anyway,
why is it so bad?
4: Primary off, replica on
Is this possible? Does it make sense?
>> Would it help if there was a separate flag in the control file for "checksums enabled in primary" and "checksums enabled in this replica", for example?
>
> Not sure I follow, should pg_checksums maintan such a flag or the StartupXLOG?
I don't know. It was just thought, probably a bad one...
- Heikki
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-08 08:17 Daniel Gustafsson <daniel@yesql.se>
parent: Heikki Linnakangas <hlinnaka@iki.fi>
0 siblings, 1 reply; 43+ messages in thread
From: Daniel Gustafsson @ 2026-09-08 08:17 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-hackers@lists.postgresql.org
> On 7 Sep 2026, at 17:57, Heikki Linnakangas <hlinnaka@iki.fi> wrote:
>>> I'm still trying to understand all the different states and interactions between online and offline changes. It's really complicated :-(. I know it's a tall order, but is there something we could do to make it simpler?
>> If there was I'd love to try it, but across the many alteratives tried during
>> this open item there hasn't been anyhing less complicated which also solves the
>> problem. Combining a WAL logged procedure with one that can rewrite the data
>> directory without any WAL entries at all is inherently complicated.
>
> So, if you can have different state in primary and a replica, there are four combinations:
>
> 1: Primary on, replica on
> 2: Primary off, replica off
>
> These are straightforward
>
> 3: Primary on, replica off
>
> You end up in this situation, if you turn on run pg_checksums to turn on checksums in primary. You stay that you really really shouldn't stay for long in this state, but why? What's the harm?
>
> One harm is that it's confusing, but if we have to deal with it anyway, why is it so bad?
Tools like pg_rewind etc rely on the fact that nodes are equal in checksum
state. What if the replica is promoted? I'm personally unconvinced that there
is a good usecase for per-node settings, but I've spent X years thinking about
checksums being replicated so I am clearly biased.
Considering how complicated it was to get replicated states right, if we want
to make it per-node I think we need to go back to the drawing board and
re-think properly rather than settle for that it seems to work. (Not that I
think that's what you're advocating, I just expect getting it work will be
complicated.)
> 4: Primary off, replica on
>
> Is this possible? Does it make sense?
This is possible using pg_checksums, but like the inverse case I don't think it
makes sense to use different settings across replication.
--
Daniel Gustafsson
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-08 08:58 Heikki Linnakangas <hlinnaka@iki.fi>
parent: Daniel Gustafsson <daniel@yesql.se>
0 siblings, 2 replies; 43+ messages in thread
From: Heikki Linnakangas @ 2026-09-08 08:58 UTC (permalink / raw)
To: Daniel Gustafsson <daniel@yesql.se>; +Cc: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-hackers@lists.postgresql.org
On 08/09/2026 11:17, Daniel Gustafsson wrote:
>> On 7 Sep 2026, at 17:57, Heikki Linnakangas <hlinnaka@iki.fi> wrote:
>
>> 3: Primary on, replica off
>>
>> You end up in this situation, if you turn on run pg_checksums to turn on checksums in primary. You stay that you really really shouldn't stay for long in this state, but why? What's the harm?
>>
>> One harm is that it's confusing, but if we have to deal with it anyway, why is it so bad?
>
> Tools like pg_rewind etc rely on the fact that nodes are equal in checksum
> state. What if the replica is promoted? I'm personally unconvinced that there
> is a good usecase for per-node settings, but I've spent X years thinking about
> checksums being replicated so I am clearly biased.
>
> Considering how complicated it was to get replicated states right, if we want
> to make it per-node I think we need to go back to the drawing board and
> re-think properly rather than settle for that it seems to work. (Not that I
> think that's what you're advocating, I just expect getting it work will be
> complicated.)
Sure, I don't see any point in this setup either. It's just something
that you can end up with. But as long as you can end up with it, we need
to deal with it gracefully.
If the replica is promoted, that seems fine. Checksums will be off.
In principle, I think pg_rewind would still work as long as
wal_log_hints=on. But I don't think we need to cater for that, erroring
out is fine.
>> 4: Primary off, replica on
>>
>> Is this possible? Does it make sense?
>
> This is possible using pg_checksums, but like the inverse case I don't think it
> makes sense to use different settings across replication.
How about we forbid this completely? Forbid running pg_checksums on
replica, unless the primary already has checksums on.
I guess that would make it impossible to run pg_checksums concurrently
in the primary and the standby. You'd have to run pg_checksums on
primary first, wait for it to finish, and only then launch it in the
standby.
Maybe print a warning if you run pg_checksums in a replica:
WARNING: unless you run pg_checksums on the primary at the same time,
the checksums will be immediately disabled again after startup
- Heikki
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-08 09:12 Daniel Gustafsson <daniel@yesql.se>
parent: Heikki Linnakangas <hlinnaka@iki.fi>
1 sibling, 0 replies; 43+ messages in thread
From: Daniel Gustafsson @ 2026-09-08 09:12 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-hackers@lists.postgresql.org
> On 8 Sep 2026, at 10:58, Heikki Linnakangas <hlinnaka@iki.fi> wrote:
>
> On 08/09/2026 11:17, Daniel Gustafsson wrote:
>>> On 7 Sep 2026, at 17:57, Heikki Linnakangas <hlinnaka@iki.fi> wrote:
>>> 3: Primary on, replica off
>>>
>>> You end up in this situation, if you turn on run pg_checksums to turn on checksums in primary. You stay that you really really shouldn't stay for long in this state, but why? What's the harm?
>>>
>>> One harm is that it's confusing, but if we have to deal with it anyway, why is it so bad?
>> Tools like pg_rewind etc rely on the fact that nodes are equal in checksum
>> state. What if the replica is promoted? I'm personally unconvinced that there
>> is a good usecase for per-node settings, but I've spent X years thinking about
>> checksums being replicated so I am clearly biased.
>> Considering how complicated it was to get replicated states right, if we want
>> to make it per-node I think we need to go back to the drawing board and
>> re-think properly rather than settle for that it seems to work. (Not that I
>> think that's what you're advocating, I just expect getting it work will be
>> complicated.)
>
> Sure, I don't see any point in this setup either. It's just something that you can end up with. But as long as you can end up with it, we need to deal with it gracefully.
Agreed, that's the gist of the patch in this thread. Identify, issue WARNINGs
on the secondary and make sure tools like pg_rewind don't break things. The
cluster will continue to operate and the state can be unified.
> If the replica is promoted, that seems fine. Checksums will be off.
>
> In principle, I think pg_rewind would still work as long as wal_log_hints=on. But I don't think we need to cater for that, erroring out is fine.
>
>>> 4: Primary off, replica on
>>>
>>> Is this possible? Does it make sense?
>> This is possible using pg_checksums, but like the inverse case I don't think it
>> makes sense to use different settings across replication.
>
> How about we forbid this completely? Forbid running pg_checksums on replica, unless the primary already has checksums on.
>
> I guess that would make it impossible to run pg_checksums concurrently in the primary and the standby. You'd have to run pg_checksums on primary first, wait for it to finish, and only then launch it in the standby.
Since pg_checksums work on offline clusters in isolation it (currently) doesn't
know about other nodes or replication at all.
> Maybe print a warning if you run pg_checksums in a replica:
>
> WARNING: unless you run pg_checksums on the primary at the same time, the checksums will be immediately disabled again after startup
Regardless of the fate of online checksums in 19 I think it would be good to
issue a warning in pg_checksums that all nodes need to be modified.
--
Daniel Gustafsson
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-08 21:21 Zsolt Parragi <zsolt.parragi@percona.com>
parent: Heikki Linnakangas <hlinnaka@iki.fi>
1 sibling, 0 replies; 43+ messages in thread
From: Zsolt Parragi @ 2026-09-08 21:21 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: pgsql-hackers@lists.postgresql.org, Daniel Gustafsson <daniel@yesql.se>; Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
On Tue, 08 Sep 2026, Heikki Linnakangas
<heikki.linnakangas@enterprisedb.com> wrote:
> In principle, I think pg_rewind would still work as long as
> wal_log_hints=on. But I don't think we need to cater for that, erroring
> out is fine.
The problem with rewind/backups/etc in a mixed cluster is all the
corner cases. Everything should work, except that we can end up with
checksum errors because an instance that supposedly has checksums
enabled ends up with a few pages without checksums. Most of the time
the issue is immediately visible, but we can construct scenarios where
users only get those checksum errors much later, when it is no longer
clear if its a real corruption or a leftover issue from some tool use.
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-09 21:45 Daniel Gustafsson <daniel@yesql.se>
parent: Heikki Linnakangas <hlinnaka@iki.fi>
1 sibling, 1 reply; 43+ messages in thread
From: Daniel Gustafsson @ 2026-09-09 21:45 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-hackers@lists.postgresql.org
> On 7 Sep 2026, at 13:08, Heikki Linnakangas <hlinnaka@iki.fi> wrote:
> Let's add a new 'sect2' for this explanation, and move it after the "Online Enabling of Checksums" section. It's currently placed under "Offline Enabling of Checksums", but it actually goes into a lot of details of how *online* checksumming works, but "Online Enabling of Checksums" is covered in the following paragraph. If you read this in order like a novel, it feels weird.
>
> I think these paragraphs could use some copy-editing too. It feels like a pretty deep technical explanation, not very accessible to a DBA. Maybe start with "The primary server and replica can have different checksum states".
>
> (Not new with this patch, but: )
>
> The placement of the states in the state diagram on that page looks bizarre. I know it's auto-generated so not sure there's much we can do about it.. but could we, please? Maybe it'd get more clear if you leave 'initdb' out of the diagram. Or consider some completely different representation.
The attached v14 rewords and reorganizes the docs, and attempts to tidy uo the
state diagram. There are no new code changes in this version.
--
Daniel Gustafsson
Attachments:
[application/octet-stream] v14-0005-doc-Documentation-updates-for-online-checksums.patch (16.6K, ../../166AD46C-25A4-4AF3-B65F-095E3746A96E@yesql.se/2-v14-0005-doc-Documentation-updates-for-online-checksums.patch)
download | inline diff:
From 9e7058af0f7ab70996bf2bd0bd6ea719e2a7ab9c Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Wed, 9 Sep 2026 23:38:11 +0200
Subject: [PATCH v14 5/5] doc: Documentation updates for online checksums
Tidy up the state diagram to be more readable and move the para about
mismatched states to its own sect2 section.
Reported-by: Heikki Linnakangas <hlinnaka@iki.fi>
Discussion: https://postgr.es/m/5b08b4d1-0982-4b77-bac5-3bdffc6583f5@iki.fi
---
doc/src/sgml/images/datachecksums.gv | 61 ++++++++--
doc/src/sgml/images/datachecksums.svg | 163 ++++++++++++++++----------
doc/src/sgml/images/meson.build | 1 +
doc/src/sgml/wal.sgml | 60 ++++++----
4 files changed, 189 insertions(+), 96 deletions(-)
diff --git a/doc/src/sgml/images/datachecksums.gv b/doc/src/sgml/images/datachecksums.gv
index dff3ff7340a..f032555db3e 100644
--- a/doc/src/sgml/images/datachecksums.gv
+++ b/doc/src/sgml/images/datachecksums.gv
@@ -1,14 +1,49 @@
-digraph G {
- A -> B [label="SELECT pg_enable_data_checksums()"];
- B -> C;
- D -> A;
- C -> D [label="SELECT pg_disable_data_checksums()"];
- E -> A [label=" --no-data-checksums"];
- E -> C [label=" --data-checksums"];
-
- A [label="off"];
- B [label="inprogress-on"];
- C [label="on"];
- D [label="inprogress-off"];
- E [label="initdb"];
+digraph "datachecksums_states" {
+ layout=dot;
+ node [label="", shape=box, style=filled, fillcolor=gray, width=1.0, fontname="sans-serif"];
+
+ m1 [label="initdb", shape=Mdiamond];
+
+ subgraph cluster01 {
+ label="Online Checksums";
+ subgraph agroup1 {
+ rank=same;
+ a1;
+ a3;
+ }
+
+ subgraph agroup2 {
+ rank=same;
+ a2;
+ a4;
+ }
+
+ a1 -> a2;
+ a2 -> a3;
+ a3 -> a4;
+ a4 -> a1;
+
+ a1 [fillcolor=lightblue, label="off"];
+ a2 [fillcolor=lightgreen, label="inprogress-on"];
+ a3 [fillcolor=lightblue, label="on"];
+ a4 [fillcolor=lightgreen, label="inprogress-off"];
+ }
+
+ subgraph cluster05 {
+ label="Offline Checksums";
+ subgraph bgroup2 {
+ rank=same;
+ b1 -> b2;
+ }
+
+ b2 -> b1;
+
+ b1 [fillcolor=lightblue, label="off"];
+ b2 [fillcolor=lightblue, label="on"];
+ }
+
+ m1 -> a1;
+ m1 -> a3;
+ m1 -> b1;
+ m1 -> b2;
}
diff --git a/doc/src/sgml/images/datachecksums.svg b/doc/src/sgml/images/datachecksums.svg
index 8c58f42922e..eee57c393c0 100644
--- a/doc/src/sgml/images/datachecksums.svg
+++ b/doc/src/sgml/images/datachecksums.svg
@@ -1,81 +1,126 @@
-<?xml version="1.0" encoding="UTF-8" standalone="no"?>
+<?xml version="1.0"?>
<!-- Generated by graphviz version 14.0.5 (20251129.0259)
-->
-<!-- Title: G Pages: 1 -->
-<svg width="409pt" height="383pt"
- viewBox="0.00 0.00 409.00 383.00" xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink">
-<g id="graph0" class="graph" transform="scale(1 1) rotate(0) translate(4 378.5)">
-<title>G</title>
-<polygon fill="white" stroke="none" points="-4,4 -4,-378.5 404.74,-378.5 404.74,4 -4,4"/>
-<!-- A -->
+<!-- Title: datachecksums_states Pages: 1 -->
+<svg xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" width="441pt" height="209pt" viewBox="0.00 0.00 441.00 209.00">
+<g id="graph0" class="graph" transform="scale(1 1) rotate(0) translate(4 204.5)">
+<title>datachecksums_states</title>
+<polygon fill="white" stroke="none" points="-4,4 -4,-204.5 437,-204.5 437,4 -4,4"/>
+<g id="clust1" class="cluster">
+<title>cluster01</title>
+<polygon fill="none" stroke="black" points="8,-8 8,-156.5 239,-156.5 239,-8 8,-8"/>
+<text xml:space="preserve" text-anchor="middle" x="123.5" y="-139.2" font-family="Times,serif" font-size="14.00">Online Checksums</text>
+</g>
+<g id="clust4" class="cluster">
+<title>cluster05</title>
+<polygon fill="none" stroke="black" points="247,-80 247,-156.5 425,-156.5 425,-80 247,-80"/>
+<text xml:space="preserve" text-anchor="middle" x="336" y="-139.2" font-family="Times,serif" font-size="14.00">Offline Checksums</text>
+</g>
+<!-- m1 -->
<g id="node1" class="node">
-<title>A</title>
-<ellipse fill="none" stroke="black" cx="80.12" cy="-268" rx="27" ry="18"/>
-<text xml:space="preserve" text-anchor="middle" x="80.12" y="-262.95" font-family="Times,serif" font-size="14.00">off</text>
+<title>m1</title>
+<polygon fill="gray" stroke="black" points="242,-200.5 197.65,-182.5 242,-164.5 286.35,-182.5 242,-200.5"/>
+<polyline fill="none" stroke="black" points="208.77,-187.01 208.77,-177.99"/>
+<polyline fill="none" stroke="black" points="230.88,-169.01 253.12,-169.01"/>
+<polyline fill="none" stroke="black" points="275.23,-177.99 275.23,-187.01"/>
+<polyline fill="none" stroke="black" points="253.12,-195.99 230.88,-195.99"/>
+<text xml:space="preserve" text-anchor="middle" x="242" y="-176.7" font-family="sans-serif" font-size="14.00">initdb</text>
</g>
-<!-- B -->
+<!-- a1 -->
<g id="node2" class="node">
-<title>B</title>
-<ellipse fill="none" stroke="black" cx="137.12" cy="-179.5" rx="61.59" ry="18"/>
-<text xml:space="preserve" text-anchor="middle" x="137.12" y="-174.45" font-family="Times,serif" font-size="14.00">inprogress-on</text>
+<title>a1</title>
+<polygon fill="lightblue" stroke="black" points="134,-124 62,-124 62,-88 134,-88 134,-124"/>
+<text xml:space="preserve" text-anchor="middle" x="98" y="-100.2" font-family="sans-serif" font-size="14.00">off</text>
</g>
-<!-- A->B -->
-<g id="edge1" class="edge">
-<title>A->B</title>
-<path fill="none" stroke="black" d="M76.5,-249.68C75.22,-239.14 75.3,-225.77 81.12,-215.5 84.2,-210.08 88.49,-205.38 93.35,-201.34"/>
-<polygon fill="black" stroke="black" points="95.22,-204.31 101.33,-195.66 91.16,-198.61 95.22,-204.31"/>
-<text xml:space="preserve" text-anchor="middle" x="187.62" y="-218.7" font-family="Times,serif" font-size="14.00">SELECT pg_enable_data_checksums()</text>
+<!-- m1->a1 -->
+<g id="edge7" class="edge">
+<title>m1->a1</title>
+<path fill="none" stroke="black" d="M208.02,-177.91C187.98,-174.61 162.76,-168.35 143,-156.5 132.95,-150.47 123.82,-141.5 116.46,-132.85"/>
+<polygon fill="black" stroke="black" points="119.38,-130.89 110.41,-125.26 113.91,-135.26 119.38,-130.89"/>
</g>
-<!-- C -->
+<!-- a3 -->
<g id="node3" class="node">
-<title>C</title>
-<ellipse fill="none" stroke="black" cx="137.12" cy="-106.5" rx="27" ry="18"/>
-<text xml:space="preserve" text-anchor="middle" x="137.12" y="-101.45" font-family="Times,serif" font-size="14.00">on</text>
+<title>a3</title>
+<polygon fill="lightblue" stroke="black" points="224,-124 152,-124 152,-88 224,-88 224,-124"/>
+<text xml:space="preserve" text-anchor="middle" x="188" y="-100.2" font-family="sans-serif" font-size="14.00">on</text>
</g>
-<!-- B->C -->
-<g id="edge2" class="edge">
-<title>B->C</title>
-<path fill="none" stroke="black" d="M137.12,-161.31C137.12,-153.73 137.12,-144.6 137.12,-136.04"/>
-<polygon fill="black" stroke="black" points="140.62,-136.04 137.12,-126.04 133.62,-136.04 140.62,-136.04"/>
+<!-- m1->a3 -->
+<g id="edge8" class="edge">
+<title>m1->a3</title>
+<path fill="none" stroke="black" d="M232.35,-168.18C225.36,-158.54 215.68,-145.18 207.14,-133.41"/>
+<polygon fill="black" stroke="black" points="210.21,-131.68 201.51,-125.64 204.54,-135.79 210.21,-131.68"/>
+</g>
+<!-- b1 -->
+<g id="node6" class="node">
+<title>b1</title>
+<polygon fill="lightblue" stroke="black" points="327,-124 255,-124 255,-88 327,-88 327,-124"/>
+<text xml:space="preserve" text-anchor="middle" x="291" y="-100.2" font-family="sans-serif" font-size="14.00">off</text>
+</g>
+<!-- m1->b1 -->
+<g id="edge9" class="edge">
+<title>m1->b1</title>
+<path fill="none" stroke="black" d="M250.99,-167.84C257.29,-158.25 265.92,-145.13 273.55,-133.53"/>
+<polygon fill="black" stroke="black" points="276.25,-135.79 278.82,-125.51 270.4,-131.95 276.25,-135.79"/>
+</g>
+<!-- b2 -->
+<g id="node7" class="node">
+<title>b2</title>
+<polygon fill="lightblue" stroke="black" points="417,-124 345,-124 345,-88 417,-88 417,-124"/>
+<text xml:space="preserve" text-anchor="middle" x="381" y="-100.2" font-family="sans-serif" font-size="14.00">on</text>
+</g>
+<!-- m1->b2 -->
+<g id="edge10" class="edge">
+<title>m1->b2</title>
+<path fill="none" stroke="black" d="M275.15,-177.42C294.03,-173.98 317.53,-167.73 336,-156.5 346.02,-150.41 355.13,-141.42 362.49,-132.78"/>
+<polygon fill="black" stroke="black" points="365.04,-135.2 368.56,-125.2 359.57,-130.82 365.04,-135.2"/>
</g>
-<!-- D -->
+<!-- a2 -->
<g id="node4" class="node">
-<title>D</title>
-<ellipse fill="none" stroke="black" cx="63.12" cy="-18" rx="63.12" ry="18"/>
-<text xml:space="preserve" text-anchor="middle" x="63.12" y="-12.95" font-family="Times,serif" font-size="14.00">inprogress-off</text>
+<title>a2</title>
+<polygon fill="lightgreen" stroke="black" points="114.25,-52 15.75,-52 15.75,-16 114.25,-16 114.25,-52"/>
+<text xml:space="preserve" text-anchor="middle" x="65" y="-28.2" font-family="sans-serif" font-size="14.00">inprogress-on</text>
</g>
-<!-- C->D -->
-<g id="edge4" class="edge">
-<title>C->D</title>
-<path fill="none" stroke="black" d="M124.23,-90.43C113.36,-77.73 97.58,-59.28 84.77,-44.31"/>
-<polygon fill="black" stroke="black" points="87.78,-42.44 78.62,-37.12 82.46,-46.99 87.78,-42.44"/>
-<text xml:space="preserve" text-anchor="middle" x="214.75" y="-57.2" font-family="Times,serif" font-size="14.00">SELECT pg_disable_data_checksums()</text>
+<!-- a1->a2 -->
+<g id="edge1" class="edge">
+<title>a1->a2</title>
+<path fill="none" stroke="black" d="M89.84,-87.7C86.25,-80.07 81.93,-70.92 77.92,-62.4"/>
+<polygon fill="black" stroke="black" points="81.14,-61.03 73.71,-53.47 74.81,-64.01 81.14,-61.03"/>
+</g>
+<!-- a4 -->
+<g id="node5" class="node">
+<title>a4</title>
+<polygon fill="lightgreen" stroke="black" points="231.25,-52 132.75,-52 132.75,-16 231.25,-16 231.25,-52"/>
+<text xml:space="preserve" text-anchor="middle" x="182" y="-28.2" font-family="sans-serif" font-size="14.00">inprogress-off</text>
</g>
-<!-- D->A -->
+<!-- a3->a4 -->
<g id="edge3" class="edge">
-<title>D->A</title>
-<path fill="none" stroke="black" d="M62.52,-36.28C61.62,-68.21 60.54,-138.57 66.12,-197.5 67.43,-211.24 70.27,-226.28 73.06,-238.85"/>
-<polygon fill="black" stroke="black" points="69.64,-239.59 75.32,-248.54 76.46,-238 69.64,-239.59"/>
+<title>a3->a4</title>
+<path fill="none" stroke="black" d="M186.52,-87.7C185.89,-80.41 185.15,-71.73 184.45,-63.54"/>
+<polygon fill="black" stroke="black" points="187.94,-63.28 183.6,-53.61 180.96,-63.87 187.94,-63.28"/>
</g>
-<!-- E -->
-<g id="node5" class="node">
-<title>E</title>
-<ellipse fill="none" stroke="black" cx="198.12" cy="-356.5" rx="32.41" ry="18"/>
-<text xml:space="preserve" text-anchor="middle" x="198.12" y="-351.45" font-family="Times,serif" font-size="14.00">initdb</text>
+<!-- a2->a3 -->
+<g id="edge2" class="edge">
+<title>a2->a3</title>
+<path fill="none" stroke="black" d="M95.33,-52.26C111.09,-61.23 130.56,-72.31 147.59,-82"/>
+<polygon fill="black" stroke="black" points="145.54,-84.86 155.96,-86.77 149,-78.78 145.54,-84.86"/>
+</g>
+<!-- a4->a1 -->
+<g id="edge4" class="edge">
+<title>a4->a1</title>
+<path fill="none" stroke="black" d="M161.18,-52.35C150.98,-60.85 138.51,-71.24 127.36,-80.53"/>
+<polygon fill="black" stroke="black" points="125.37,-77.64 119.93,-86.73 129.85,-83.01 125.37,-77.64"/>
</g>
-<!-- E->A -->
+<!-- b1->b2 -->
<g id="edge5" class="edge">
-<title>E->A</title>
-<path fill="none" stroke="black" d="M179.16,-341.6C159.64,-327.29 129.05,-304.86 107.03,-288.72"/>
-<polygon fill="black" stroke="black" points="109.23,-286 99.1,-282.91 105.09,-291.64 109.23,-286"/>
-<text xml:space="preserve" text-anchor="middle" x="208.57" y="-307.2" font-family="Times,serif" font-size="14.00"> --no-data-checksums</text>
+<title>b1->b2</title>
+<path fill="none" stroke="black" d="M327.21,-93.01C329.33,-92.77 331.44,-92.61 333.55,-92.54"/>
+<polygon fill="black" stroke="black" points="333.21,-96.03 343.35,-92.96 333.51,-89.03 333.21,-96.03"/>
</g>
-<!-- E->C -->
+<!-- b2->b1 -->
<g id="edge6" class="edge">
-<title>E->C</title>
-<path fill="none" stroke="black" d="M227.13,-348.04C242.29,-342.72 259.95,-334.06 271.12,-320.5 301.5,-283.62 316.36,-257.78 294.12,-215.5 268.41,-166.6 209.42,-135.53 171.52,-119.85"/>
-<polygon fill="black" stroke="black" points="172.96,-116.65 162.37,-116.21 170.37,-123.16 172.96,-116.65"/>
-<text xml:space="preserve" text-anchor="middle" x="350.87" y="-218.7" font-family="Times,serif" font-size="14.00"> --data-checksums</text>
+<title>b2->b1</title>
+<path fill="none" stroke="black" d="M344.86,-118.98C342.75,-119.23 340.63,-119.39 338.52,-119.46"/>
+<polygon fill="black" stroke="black" points="338.86,-115.97 328.72,-119.05 338.57,-122.96 338.86,-115.97"/>
</g>
</g>
</svg>
diff --git a/doc/src/sgml/images/meson.build b/doc/src/sgml/images/meson.build
index 220e3eaafb8..54e80fb402c 100644
--- a/doc/src/sgml/images/meson.build
+++ b/doc/src/sgml/images/meson.build
@@ -11,6 +11,7 @@ image_targets = []
fixup_svg_xsl = files('fixup-svg.xsl')
all_files = [
+ 'datachecksums.gv',
'genetic-algorithm.gv',
'gin.gv',
'pagelayout.txt',
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index dbbfdae5e4f..edd39ee437f 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -316,30 +316,6 @@
application can be used to enable or disable data checksums, as well as
verify checksums, on an offline cluster.
</para>
-
- <para>
- An offline change provides durability differently from an
- <link linkend="checksums-online-enable-disable">online change</link>.
- An online transition is WAL-logged: it is ordered against all other
- WAL records, it is replayed after a crash, and it propagates to
- standbys. An offline change is recorded only in the cluster's
- control file: it writes no WAL, it is invisible to replication, and
- it has no defined ordering against WAL the node has not replayed
- yet. When a node later replays WAL that contains an online checksum
- state change, that change takes effect on the node even if it was
- written before the offline change was made.
- </para>
-
- <para>
- An offline change only affects the data directory it is run on; the
- new state does not propagate over replication. In a replication setup
- the same change must be applied to all nodes while all of them are
- stopped, as described in <xref linkend="app-pgchecksums"/>. A standby
- whose state diverges logs a warning but keeps its local setting. The
- mismatch persists until the states are brought together again, with
- the offline procedure or with an online transition; do this promptly.
- </para>
-
</sect2>
<sect2 id="checksums-online-enable-disable" xreflabel="Online Enabling of Checksums">
@@ -456,7 +432,43 @@
still required.
</para>
</sect3>
+ </sect2>
+
+ <sect2 id="checksums-mismatched-states">
+ <title>Mismatched Data Checksums States in a Replicated Cluster</title>
+ <para>
+ The primary and secondaries can end up with different data checksums
+ states due to the different durability models between offline and online
+ checksum operations.
+ </para>
+ <para>
+ An online transition is WAL-logged, which means that it is replayed after a
+ crash, and the transition along with all checksum updates are propagated to
+ standbys. An offline change is recorded only in the cluster's control
+ file, it writes no WAL and is invisible to replication. The changes made
+ to the datafiles to write checksums are not WAL logged. This means that it
+ has no defined ordering against WAL the node has not replayed yet. When a
+ node later replays WAL that contains an online checksum state change, that
+ change takes effect on the node even if it was written before the offline
+ change was made.
+ </para>
+ <para>
+ An offline change only affects the data directory it is run on; the new
+ state does not propagate over replication. This means that all nodes can
+ be operated on in parallel, and no additional network traffic or WAL
+ archive traffic will occur. In a replication setup the same change must be
+ applied to all nodes while all of them are stopped, as described in
+ <xref linkend="app-pgchecksums"/>. A standby whose state diverges logs a
+ warning but keeps its local setting. The mismatch persists until the
+ states are brought together again, with the offline procedure or with an
+ online transition.
+ </para>
+
+ <para>
+ Running a replicated cluster with different data checksum states on the
+ different nodes is not supported or recommended.
+ </para>
</sect2>
</sect1>
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v14-0004-pg_combinebackup-Refuse-mixed-data-checksum-stat.patch (10.0K, ../../166AD46C-25A4-4AF3-B65F-095E3746A96E@yesql.se/3-v14-0004-pg_combinebackup-Refuse-mixed-data-checksum-stat.patch)
download | inline diff:
From cbc71b7d64320f511ae91d44c75122151644ed9f Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 3 Sep 2026 23:32:04 +0200
Subject: [PATCH v14 4/5] pg_combinebackup: Refuse mixed data checksum states
in a backup chain
check_control_files() detected a chain whose backups were taken under
different data checksum states, warned, and proceeded. The output
directory keeps the last backup's control file, so a full backup taken
with checksums off combined with an incremental taken after an offline
enable produces a cluster whose control file says "on" while most of
its blocks carry no checksums; it starts, and then every connection
dies on the first unchecksummed catalog page. An offline enable
between two backups of a chain is all it takes, since it rewrites
every page without logging anything, so the incremental backup does
not re-ship the pages.
Turn the warning into an error, matching what pg_rewind does for the
equivalent combinations. The check stays asymmetric on purpose: when
the last backup was taken with checksums off, stale checksums from an
earlier backup are never verified, and that chain remains usable.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Reviewed-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Discussion: https://postgr.es/m/anwm6UPxoVS41QA2@bdtpg
---
doc/src/sgml/ref/pg_combinebackup.sgml | 19 +--
src/bin/pg_combinebackup/pg_combinebackup.c | 14 +-
src/test/modules/test_checksums/meson.build | 1 +
.../t/024_combinebackup_mixed.pl | 146 ++++++++++++++++++
4 files changed, 167 insertions(+), 13 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
diff --git a/doc/src/sgml/ref/pg_combinebackup.sgml b/doc/src/sgml/ref/pg_combinebackup.sgml
index 9a6d201e0b8..f5d7e177b36 100644
--- a/doc/src/sgml/ref/pg_combinebackup.sgml
+++ b/doc/src/sgml/ref/pg_combinebackup.sgml
@@ -306,18 +306,19 @@ PostgreSQL documentation
<para>
<literal>pg_combinebackup</literal> does not recompute page checksums when
- writing the output directory. Therefore, if any of the backups used for
- reconstruction were taken with checksums disabled, but the final backup was
- taken with checksums enabled, the resulting directory may contain pages
- with invalid checksums.
+ writing the output directory. It therefore refuses a chain in which some
+ of the backups used for reconstruction were taken with checksums disabled
+ but the final backup was taken with checksums enabled: the resulting
+ directory would contain pages with invalid checksums, and the cluster
+ would fail checksum verification as soon as it reads them. Take a new
+ full backup after enabling data checksums with
+ <xref linkend="app-pgchecksums"/>.
</para>
<para>
- To avoid this problem, taking a new full backup after changing the checksum
- state of the cluster using <xref linkend="app-pgchecksums"/> is
- recommended. Otherwise, you can disable and then optionally reenable
- checksums on the directory produced by <literal>pg_combinebackup</literal>
- in order to correct the problem.
+ The reverse case is accepted: when the final backup was taken with
+ checksums disabled, stale checksums copied from an earlier backup are
+ never verified.
</para>
</refsect1>
diff --git a/src/bin/pg_combinebackup/pg_combinebackup.c b/src/bin/pg_combinebackup/pg_combinebackup.c
index 86e2ee37c40..254a27b125b 100644
--- a/src/bin/pg_combinebackup/pg_combinebackup.c
+++ b/src/bin/pg_combinebackup/pg_combinebackup.c
@@ -670,13 +670,19 @@ check_control_files(int n_backups, char **backup_dirs)
pg_log_debug("system identifier is %" PRIu64, system_identifier);
/*
- * Warn the user if not all backups are in the same state with regards to
- * checksums.
+ * Reject a chain whose backups were not all in the same state with
+ * regards to checksums. The loop above only flags that when the last
+ * backup has checksums enabled: the blocks taken from the older backups
+ * have no checksums, and pg_combinebackup does not recompute them, so the
+ * combined cluster would fail verification as soon as it reads them. The
+ * other direction is harmless, since stale checksums are never verified
+ * while checksums are disabled.
*/
if (data_checksum_mismatch)
{
- pg_log_warning("only some backups have checksums enabled");
- pg_log_warning_hint("Disable, and optionally reenable, checksums on the output directory to avoid failures.");
+ pg_log_error("only some backups have checksums enabled");
+ pg_log_error_hint("Take a new full backup after changing the data checksum state with pg_checksums.");
+ exit(1);
}
return system_identifier;
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index 7d07c757052..7eccd5156b9 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -47,6 +47,7 @@ tests += {
't/021_rewind_divergent_transitions.pl',
't/022_rewind_state.pl',
't/023_rewind_standby_target.pl',
+ 't/024_combinebackup_mixed.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
new file mode 100644
index 00000000000..2f51c7fd9dc
--- /dev/null
+++ b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
@@ -0,0 +1,146 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_combinebackup refuses a chain whose backups were taken under
+# different data checksum states when the final state has them enabled.
+#
+# The output directory keeps the last backup's control file, so a full
+# backup taken with checksums off plus an incremental taken after an
+# offline enable would produce a directory whose control file says "on"
+# while most of its blocks have no checksums. An offline enable between
+# the two backups is enough to get there: it rewrites every page but logs
+# nothing, so the incremental backup does not re-ship the pages.
+#
+# The reverse order stays allowed: with the final backup taken with
+# checksums off, stale checksums from an earlier backup are never verified.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('node');
+$node->init(
+ no_data_checksums => 1,
+ has_archiving => 1,
+ allows_streaming => 1);
+$node->append_conf('postgresql.conf', 'summarize_wal = on');
+$node->append_conf('postgresql.conf', 'autovacuum = off');
+$node->start;
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+test_checksum_state($node, 'off');
+
+$node->backup('full');
+
+# Offline enable: rewrites every page, logs nothing.
+$node->stop;
+system_or_bail('pg_checksums', '--enable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'on');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full/backup_manifest',
+ ],
+ 'incremental backup with checksums on');
+
+# Combining the chain must be refused: the result would say "on" while the
+# blocks inherited from the full backup have no checksums.
+command_fails_like(
+ [
+ 'pg_combinebackup',
+ $node->backup_dir . '/full',
+ $node->backup_dir . '/incr',
+ '--output' => $node->backup_dir . '/combined',
+ ],
+ qr/only some backups have checksums enabled/,
+ 'pg_combinebackup refuses a chain crossing an offline enable');
+
+# The other direction: full backup with checksums on, offline disable, then
+# an incremental. The last backup wins, so the combined cluster comes up off.
+$node->backup('full2');
+
+$node->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr2',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full2/backup_manifest',
+ ],
+ 'incremental backup with checksums off');
+
+my $restored = PostgreSQL::Test::Cluster->new('restored');
+$restored->init_from_backup(
+ $node, 'incr2',
+ combine_with_prior => ['full2'],
+ has_restoring => 1,
+ standby => 0);
+
+$restored->start;
+
+my ($rc, $stdout, $stderr) = $restored->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'combined backup accepts connections') or diag("stderr: $stderr");
+is($stdout, 'off', 'combined backup follows the last backup state');
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t3;');
+is($rc, 0, 'blocks from the incremental backup are readable')
+ or diag("stderr: $stderr");
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t1;');
+is($rc, 0, 'blocks inherited from the full backup are readable')
+ or diag("stderr: $stderr");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($restored->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures in the combined backup');
+
+$restored->stop('immediate');
+$node->stop;
+
+done_testing();
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v14-0003-pg_rewind-Check-the-data-checksum-states-of-sour.patch (22.5K, ../../166AD46C-25A4-4AF3-B65F-095E3746A96E@yesql.se/4-v14-0003-pg_rewind-Check-the-data-checksum-states-of-sour.patch)
download | inline diff:
From e01ff6bf8dcc219815c56130c294f11fe848620e Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 3 Sep 2026 23:31:52 +0200
Subject: [PATCH v14 3/5] pg_rewind: Check the data checksum states of source
and target
pg_rewind installs the source's control file on the target, but every
block it does not copy keeps the target's content. With checksums
enabled on the target and disabled on the source, the rewound server
claims enabled checksums while the blocks copied from the source have
none, and fails checksum verification as soon as it reads them; in the
worst case every connection attempt dies on an unverifiable catalog
page.
Comparing the control files is not enough. Replay on the rewound
server resumes from the last common checkpoint and adopts the data
checksum state recorded there, so an offline disable on the source
after the divergence leaves both control files saying "off" while the
rewound server still resumes with verification enabled. Check the
state carried by the divergence checkpoint as well, which pg_rewind
already reads.
Refuse both cases, and an interrupted online transition on either
side, same as pg_checksums does. The opposite mismatch, checksums on
the source only, is allowed with a warning: recovery under the backup
label adopts the divergence state, so the rewound server keeps
checksums disabled (or converges through the replayed WAL if the
source enabled them online), and no page can fail verification.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Reviewed-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Discussion: https://postgr.es/m/anwm6UPxoVS41QA2@bdtpg
---
doc/src/sgml/ref/pg_rewind.sgml | 14 ++
src/bin/pg_rewind/parsexlog.c | 4 +-
src/bin/pg_rewind/pg_rewind.c | 93 ++++++++++-
src/bin/pg_rewind/pg_rewind.h | 1 +
src/test/modules/test_checksums/meson.build | 2 +
.../test_checksums/t/022_rewind_state.pl | 145 +++++++++++++++++
.../t/023_rewind_standby_target.pl | 153 ++++++++++++++++++
7 files changed, 410 insertions(+), 2 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/022_rewind_state.pl
create mode 100644 src/test/modules/test_checksums/t/023_rewind_standby_target.pl
diff --git a/doc/src/sgml/ref/pg_rewind.sgml b/doc/src/sgml/ref/pg_rewind.sgml
index b95e3868e9d..ac3d0c9328f 100644
--- a/doc/src/sgml/ref/pg_rewind.sgml
+++ b/doc/src/sgml/ref/pg_rewind.sgml
@@ -124,6 +124,20 @@ PostgreSQL documentation
<literal>on</literal>, but is enabled by default.
</para>
+ <para>
+ The data checksum states of the source and target server must be
+ compatible. <application>pg_rewind</application> refuses to run if an
+ online checksum state transition was interrupted on either server, or
+ if data checksums are enabled on the target server, or were enabled at
+ the point of divergence, while the source server runs without them: the
+ rewound server would fail checksum verification on the blocks copied
+ from the source. The opposite combination is allowed with a warning;
+ the rewound server keeps data checksums disabled unless the source
+ server enabled them online after the point of divergence. See
+ <xref linkend="app-pgchecksums"/> for how to change the state
+ consistently across a replication setup.
+ </para>
+
<warning>
<title>Warning: Failures While Rewinding</title>
<para>
diff --git a/src/bin/pg_rewind/parsexlog.c b/src/bin/pg_rewind/parsexlog.c
index 023e23b063c..6e87b00f8c2 100644
--- a/src/bin/pg_rewind/parsexlog.c
+++ b/src/bin/pg_rewind/parsexlog.c
@@ -167,7 +167,8 @@ readOneRecord(const char *datadir, XLogRecPtr ptr, int tliIndex,
void
findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
- XLogRecPtr *lastchkptredo, const char *restoreCommand)
+ XLogRecPtr *lastchkptredo, uint32 *lastchkptdatachecksums,
+ const char *restoreCommand)
{
/* Walk backwards, starting from the given record */
XLogRecord *record;
@@ -255,6 +256,7 @@ findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
*lastchkptrec = searchptr;
*lastchkpttli = checkPoint.ThisTimeLineID;
*lastchkptredo = checkPoint.redo;
+ *lastchkptdatachecksums = checkPoint.dataChecksumState;
break;
}
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 0e4c2df3f4b..d2521dab333 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -145,6 +145,7 @@ main(int argc, char **argv)
XLogRecPtr chkptrec;
TimeLineID chkpttli;
XLogRecPtr chkptredo;
+ uint32 chkptdatachecksums;
TimeLineID source_tli;
TimeLineID target_tli;
XLogRecPtr target_wal_endrec;
@@ -472,10 +473,51 @@ main(int argc, char **argv)
keepwal_init();
findLastCheckpoint(datadir_target, divergerec, lastcommontliIndex,
- &chkptrec, &chkpttli, &chkptredo, restore_command);
+ &chkptrec, &chkpttli, &chkptredo, &chkptdatachecksums,
+ restore_command);
pg_log_info("rewinding from last common checkpoint at %X/%08X on timeline %u",
LSN_FORMAT_ARGS(chkptrec), chkpttli);
+ /*
+ * Replay on the rewound server resumes from the last common checkpoint
+ * and adopts the data checksum state recorded there, not the state in the
+ * control file installed from the source. If checksums were enabled at
+ * the divergence point but the source runs without them, the rewound
+ * server would verify checksums while the blocks copied from the source
+ * have none. The control files cannot reveal this: an offline disable on
+ * the source after the divergence leaves both of them saying "off".
+ *
+ * Only the fully enabled state needs checking. An in-progress state at
+ * the divergence point means a transition was still running there, and
+ * all of its page rewrites are logged after that point, so replay brings
+ * the rewound server to whatever state the copied WAL ends in.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the source server or recreate the target server from a base backup.");
+ exit(1);
+ }
+
+ /*
+ * The same comparison is needed against the target: every block the
+ * rewind does not copy keeps the target's content. The control files
+ * cannot reveal this case either, in the other direction: a standby's
+ * checkpoints are written by its upstream primary, so a standby whose
+ * checksums were disabled offline still has "on" checkpoints in its WAL,
+ * while both control files may agree.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_target.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the target server");
+ pg_log_error_detail("Blocks kept from the target would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the target server with pg_checksums or recreate the target server from a base backup.");
+ exit(1);
+ }
+
/* Initialize the hash table to track the status of each file */
filehash_init();
@@ -799,6 +841,55 @@ sanityChecks(void)
pg_fatal("target server needs to use either data checksums or \"wal_log_hints = on\"");
}
+ /*
+ * The rewound target keeps its control file fields from the source, but
+ * every block the rewind does not copy keeps the target's content. If
+ * checksums are enabled on the target and disabled on the source, the
+ * result would claim enabled checksums while the blocks copied from the
+ * source have none, and the target would fail checksum verification as
+ * soon as it reads them. Refuse that combination, and refuse an
+ * interrupted online transition on either side, same as pg_checksums.
+ *
+ * The opposite mismatch is allowed: recovery under the backup label
+ * written by pg_rewind adopts the checksum state as of the divergence
+ * point, so a target whose own checkpoints carry no checksums stays
+ * without them, or converges through the replayed WAL if the source
+ * enabled them online. Only warn about it, so that an offline change on
+ * the source is not overlooked.
+ *
+ * That reasoning does not hold when the target is a standby, whose
+ * checkpoints were written by its upstream primary; see the checks after
+ * findLastCheckpoint().
+ */
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted on the target server");
+ pg_log_error_hint("Start the server and shut it down cleanly to reset the state, then retry; on a standby, let replication complete the transition first.");
+ exit(1);
+ }
+ if (ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition is incomplete on the source server");
+ pg_log_error_hint("Let the transition complete, or reset the state with a clean restart, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_OFF)
+ {
+ pg_log_error("data checksums are enabled on the target server but disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, and the target would fail checksum verification after the rewind.");
+ pg_log_error_hint("Either disable data checksums on the target server or enable them on the source server, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_OFF &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_warning("data checksums are disabled on the target server but enabled on the source server");
+ pg_log_warning_detail("The rewound server will keep data checksums disabled unless the source server enabled them online after the point of divergence.");
+ }
+
/*
* Target cluster better not be running. This doesn't guard against
* someone starting the cluster concurrently. Also, this is probably more
diff --git a/src/bin/pg_rewind/pg_rewind.h b/src/bin/pg_rewind/pg_rewind.h
index 9a981f7f246..9187c3bf72f 100644
--- a/src/bin/pg_rewind/pg_rewind.h
+++ b/src/bin/pg_rewind/pg_rewind.h
@@ -39,6 +39,7 @@ extern void findLastCheckpoint(const char *datadir, XLogRecPtr forkptr,
int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
XLogRecPtr *lastchkptredo,
+ uint32 *lastchkptdatachecksums,
const char *restoreCommand);
extern XLogRecPtr readOneRecord(const char *datadir, XLogRecPtr ptr,
int tliIndex, const char *restoreCommand);
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index e5d38fafb7d..7d07c757052 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -45,6 +45,8 @@ tests += {
't/019_standby_shutdown_catchup.pl',
't/020_cascade_divergence.pl',
't/021_rewind_divergent_transitions.pl',
+ 't/022_rewind_state.pl',
+ 't/023_rewind_standby_target.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/022_rewind_state.pl b/src/test/modules/test_checksums/t/022_rewind_state.pl
new file mode 100644
index 00000000000..e6d9c645a0b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/022_rewind_state.pl
@@ -0,0 +1,145 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind checks the data checksum states of source and target.
+# Checksums enabled on the target, or at the point of divergence, with
+# a source running without them are refused: the rewound server would
+# verify checksums on blocks copied from a source that has none. The
+# opposite mismatch only warns; replay keeps the target's own state.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+# Scenarios 1 and 2: checksums on from initdb. wal_log_hints keeps the
+# target eligible for pg_rewind after checksums are disabled on it.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b);
+
+# Failover to B, and divergence on A.
+$node_b->promote;
+$node_b->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_a->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_a->stop('fast');
+
+# Offline disable on the source only: enabled target, disabled source.
+$node_b->stop;
+$node_b->checksum_disable_offline;
+$node_b->start;
+test_checksum_state($node_b, 'off');
+
+my @rewind_cmd = (
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'));
+
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums are enabled on the target server but disabled on the source server/,
+ 'refuses an enabled target with a disabled source');
+
+# The refusal happens before any modification: the target still starts
+# on its own timeline.
+$node_a->start;
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'target untouched by the refused rewind');
+$node_a->stop('fast');
+
+# Scenario 2: disabling the target too makes the control files match,
+# but checksums were still enabled at the point of divergence, which is
+# the state replay on the rewound server would resume with.
+$node_a->checksum_disable_offline;
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums were enabled at the point of divergence but are disabled on the source server/,
+ 'refuses when checksums were enabled at the point of divergence');
+
+# Scenario 3: fresh pair without checksums; offline enable on the
+# source only. Allowed with a warning, and the rewound server keeps
+# checksums disabled.
+my $node_c = PostgreSQL::Test::Cluster->new('node_c');
+$node_c->init(allows_streaming => 1, no_data_checksums => 1);
+$node_c->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_c->start;
+$node_c->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_c->backup('backup');
+my $node_d = PostgreSQL::Test::Cluster->new('node_d');
+$node_d->init_from_backup($node_c, 'backup', has_streaming => 1);
+$node_d->start;
+$node_c->wait_for_catchup($node_d);
+
+$node_d->promote;
+$node_d->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_c->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_c->stop('fast');
+
+$node_d->stop;
+$node_d->checksum_enable_offline;
+$node_d->start;
+test_checksum_state($node_d, 'on');
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_c->data_dir,
+ '--source-server' => $node_d->connstr('postgres'),
+ ]);
+like(
+ $stderr,
+ qr/data checksums are disabled on the target server but enabled on the source server/,
+ 'warns for a disabled target with an enabled source');
+like($stderr, qr/Done!/, 'rewind completed despite the warning');
+
+# The rewound server follows D and keeps its own state.
+$node_c->append_conf('postgresql.conf', 'port = ' . $node_c->port);
+$node_c->enable_streaming($node_d);
+$node_c->set_standby_mode;
+$node_c->start;
+$node_d->wait_for_catchup($node_c);
+is($node_c->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+test_checksum_state($node_c, 'off');
+
+$node_c->stop;
+$node_d->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/023_rewind_standby_target.pl b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
new file mode 100644
index 00000000000..9071114a9ae
--- /dev/null
+++ b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
@@ -0,0 +1,153 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_rewind compares the data checksum state at the point of
+# divergence against the target as well as the source.
+#
+# A standby's checkpoints are written by its upstream primary, so a standby
+# whose checksums were disabled offline still has "on" checkpoints in its
+# WAL. Rewinding it onto a promoted sibling passes every control file
+# check (target off + source on only warns), yet recovery after the rewind
+# adopts the divergence checkpoint's state and the rewound server would
+# verify checksums over the pages it wrote while it was locally "off".
+# pg_rewind must refuse; after an offline enable on the target, which
+# rewrites all of its pages, the rewind goes through.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum_extended' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum_extended\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t0 AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+
+my $lagging = PostgreSQL::Test::Cluster->new('lagging');
+$lagging->init_from_backup($primary, 'backup', has_streaming => 1);
+$lagging->start;
+
+my $failover = PostgreSQL::Test::Cluster->new('failover');
+$failover->init_from_backup($primary, 'backup', has_streaming => 1);
+$failover->start;
+
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+
+test_checksum_state($primary, 'on');
+test_checksum_state($lagging, 'on');
+test_checksum_state($failover, 'on');
+
+# Offline-disable checksums on the lagging standby only. This is a
+# divergence 012_offline_standby.pl declares survivable: the standby keeps
+# its own state, warns, and stays readable.
+$lagging->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $lagging->data_dir);
+$lagging->start;
+test_checksum_state($lagging, 'off');
+
+# Give the lagging standby plenty of pages to write while it is locally
+# "off", and force a restartpoint so they reach disk without checksums.
+# The other standby replays the same WAL with checksums on.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+$lagging->safe_psql('postgres', "CHECKPOINT;");
+
+# Take the "failover" standby out of the picture. It is promoted from
+# this position later, so everything replayed from here on is the part of
+# the WAL that diverges and that pg_rewind will roll back. t1 is already
+# replayed everywhere, so its blocks are *not* rolled back.
+$failover->stop;
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$primary->wait_for_catchup($lagging);
+
+# Failover: promote the other standby, which forks the timeline behind the
+# replay position of the lagging standby.
+$primary->stop('fast');
+$failover->start;
+$failover->promote;
+$failover->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$failover->safe_psql('postgres', "CHECKPOINT;");
+
+# Re-attaching the lagging standby with pg_rewind must be refused: the
+# divergence checkpoint says "on" while the target is "off", so the rewound
+# server would verify checksums over the blocks it keeps.
+$lagging->stop;
+
+command_fails_like(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ],
+ qr/data checksums were enabled at the point of divergence but are disabled on the target server/,
+ 'pg_rewind refuses a target whose divergence checkpoint has checksums');
+
+# Enabling checksums offline on the target rewrites all of its pages, after
+# which the rewind is safe.
+system_or_bail('pg_checksums', '--enable', '--pgdata', $lagging->data_dir);
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ]);
+like($stderr, qr/Done!/, 'pg_rewind completes after the offline enable');
+
+$lagging->append_conf('postgresql.conf', 'port = ' . $lagging->port);
+$lagging->enable_streaming($failover);
+$lagging->set_standby_mode;
+$lagging->start;
+
+my ($rc, $out, $err) =
+ $lagging->psql('postgres',
+ "SELECT setting FROM pg_settings " . "WHERE name = 'data_checksums';");
+is($rc, 0, 'rewound standby accepts connections') or diag("stderr: $err");
+is($out, 'on', 'rewound standby resumes with checksums on');
+
+($rc, $out, $err) = $lagging->psql('postgres', "SELECT count(*) FROM t1;");
+is($rc, 0, 'blocks kept from the target stay readable')
+ or diag("stderr: $err");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($lagging->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the rewound standby');
+
+$lagging->stop('immediate');
+$failover->stop;
+done_testing();
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v14-0002-pg_checksums-Refuse-interrupted-transitions-note.patch (7.9K, ../../166AD46C-25A4-4AF3-B65F-095E3746A96E@yesql.se/5-v14-0002-pg_checksums-Refuse-interrupted-transitions-note.patch)
download | inline diff:
From 3bf1a096ef82f61150bf64db3878dbd941201dd9 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 3 Sep 2026 23:31:33 +0200
Subject: [PATCH v14 2/5] pg_checksums: Refuse interrupted transitions, note
change is local
An inprogress-on or inprogress-off control file means an online
transition was cut short; running the offline tool on top of it mixes
two procedures. Refuse it in every mode, with a hint pointing at the
start-stop cycle (or, on a standby, at letting replication finish)
that resets the state. Also print that the change applies to one
data directory only, with standby-aware wording, since in a
replication setup the same change must be made on every node.
On a primary, inprogress-on cannot survive a graceful stop: the
checksums launcher resolves it from its exit cleanup, and a crashed
primary is already rejected by the existing "cluster must be shut
down" check. inprogress-off can, however: it is set by the backend
running pg_disable_data_checksums(), and a fast shutdown arriving
between its two barriers leaves it behind in a cleanly shut down
control file, where the next start-stop cycle resolves it at end of
recovery. A standby can be stopped with either state, since it has
no launcher and only carries forward whatever state the last replayed
WAL record left it in, with a restartpoint persisting that as-is.
The standby is also the deterministic way to reach the new guard,
which is why the test coverage uses one; the primary window would
need an injection point between the two barriers.
The pg_checksums page said the tool still processes all relation files
regardless of an interrupted online transition; describe the refusal
instead.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Reviewed-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Discussion: https://postgr.es/m/anwm6UPxoVS41QA2@bdtpg
---
doc/src/sgml/ref/pg_checksums.sgml | 11 ++++---
src/bin/pg_checksums/pg_checksums.c | 21 +++++++++++++
.../test_checksums/t/012_offline_standby.pl | 31 ++++++++++++++-----
3 files changed, 51 insertions(+), 12 deletions(-)
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index 2d2057a8c19..abc035e5409 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -49,11 +49,12 @@ PostgreSQL documentation
</para>
<para>
- When enabling checksums with <application>pg_checksums</application>, if
- checksums were in the process of being enabled using
- <xref linkend="checksums-online-enable-disable"/> when the cluster was shut
- down, <application>pg_checksums</application> will still process all
- relation files regardless of the progress of online checksum processing.
+ If checksums were in the process of being enabled or disabled using
+ <xref linkend="checksums-online-enable-disable"/> when the cluster was
+ shut down, the control file still records that interrupted state, and
+ <application>pg_checksums</application> refuses to run in any mode.
+ Start the cluster and shut it down cleanly to reset the state, then
+ retry; on a standby, let replication complete the transition first.
</para>
<para>
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 5d6ea318784..3b58c6ca608 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -586,6 +586,21 @@ main(int argc, char *argv[])
ControlFile->state != DB_SHUTDOWNED_IN_RECOVERY)
pg_fatal("cluster must be shut down");
+ /*
+ * An inprogress state means an online transition was cut short. A
+ * standby stopped mid-transition carries either state; a cleanly shut
+ * down primary can still carry inprogress-off, which a fast shutdown
+ * during pg_disable_data_checksums() leaves behind, while inprogress-on
+ * is always resolved by the launcher's exit cleanup.
+ */
+ if (ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted");
+ pg_log_error_hint("Start and cleanly shut down the cluster once to reset the data checksum state, then retry. On a standby, let replication complete the transition first.");
+ exit(1);
+ }
+
if (ControlFile->data_checksum_version != PG_DATA_CHECKSUM_VERSION &&
mode == PG_MODE_CHECK)
pg_fatal("data checksums are not enabled in cluster");
@@ -673,6 +688,12 @@ main(int argc, char *argv[])
printf(_("Checksums enabled in cluster\n"));
else
printf(_("Checksums disabled in cluster\n"));
+
+ printf(_("This change applies to this data directory only.\n"));
+ if (ControlFile->state == DB_SHUTDOWNED_IN_RECOVERY)
+ printf(_("This node appears to be a standby; apply the same change to the primary and all other standbys.\n"));
+ else
+ printf(_("In a replication setup, apply the same change to every node while all are stopped, before restarting any of them.\n"));
}
return 0;
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
index c0ec374c8be..ba1f14e0399 100644
--- a/src/test/modules/test_checksums/t/012_offline_standby.pl
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -164,7 +164,10 @@ $newnode->stop('immediate');
# Converge the cluster: enable offline on the standby too.
$standby->stop;
-$standby->checksum_enable_offline;
+command_checks_all(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ 0, [qr/appears to be a standby/],
+ [], 'standby-role notice on offline enable');
$standby->start;
test_checksum_state($standby, 'on');
$primary->wait_for_catchup($standby);
@@ -232,11 +235,11 @@ unlike(
);
# Scenario 4: a standby stopped while replaying an interrupted online
-# transition keeps the interrupted state in its own control file. A
-# primary is never caught this way, as its checksums launcher resolves
-# inprogress-on back to off from its exit cleanup. A standby has no
-# launcher; it carries forward whatever the last replayed record left
-# it in.
+# transition keeps the interrupted state in its own control file, and
+# pg_checksums must refuse to touch it. A primary is never caught this
+# way, as its checksums launcher resolves inprogress-on back to off from
+# its exit cleanup. A standby has no launcher; it carries forward
+# whatever the last replayed record left it in.
# Block an online enable on the primary at inprogress-on with a
# blocking temp table, same trick as in 004_offline.pl.
@@ -250,6 +253,20 @@ wait_for_checksum_state($standby, 'inprogress-on');
# Stop the standby cleanly; its restartpoint persists inprogress-on to
# its own control file, since nothing on a standby resolves it away.
$standby->stop;
+
+command_fails_like(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --enable refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --check refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--disable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --disable refuses a standby stopped mid-transition');
+
$standby->start;
wait_for_checksum_state($standby, 'inprogress-on');
@@ -272,7 +289,7 @@ wait_for_checksum_state($primary, 'on');
$primary->wait_for_catchup($standby);
wait_for_checksum_state($standby, 'on');
-is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
'10001', 'standby readable once the transition completes');
# Scenario 4 continued: the backup taken mid-transition must also
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v14-0001-Do-not-adopt-data-checksum-state-from-another-no.patch (137.1K, ../../166AD46C-25A4-4AF3-B65F-095E3746A96E@yesql.se/6-v14-0001-Do-not-adopt-data-checksum-state-from-another-no.patch)
download | inline diff:
From 0f4182826f73fe5181d4ef093c5e8f2fc092e25a Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 3 Sep 2026 23:29:17 +0200
Subject: [PATCH v14 1/5] Do not adopt data checksum state from another node
during replay
Offline pg_checksums changes are local to one node, but replay adopted
the data checksum state carried by checkpoint records unconditionally.
After an offline enable on the primary, a standby whose pages were
never rewritten started verifying checksums it does not have; after an
offline change on a standby, the next replayed checkpoint silently
reverted it.
To fix, make the control file track this node's state alone, and have
replay cross-check the replayed state against it instead of adopting
it, warning once per divergent value and reporting when the states
agree again. pg_control gains a watermark, normally the end LSN of
the newest XLOG2_CHECKSUMS record the node has written or applied, so
that replay can skip transition records whose effect the control file
already contains, and a flag marking a state last written by
pg_checksums, which recovery must never overwrite with a replayed one.
Since the control file may only claim "on" once every page on disk
carries a checksum, persisting the state is tied to flushes:
XLOG2_CHECKSUMS replay persists every state but "on" immediately and
leaves "on" to the next restartpoint, checkpoints and restartpoints
only persist a state their flush ran under from beginning to end, and
transitions publish their state under the new
DataChecksumTransitionLock so that the states carried by WAL records
match their WAL order. The comments in xlog.c spell out the
individual rules.
Recovery from a base backup may be an exception to not adopting: its
control file was copied at an arbitrary moment, so the state carried
by the starting checkpoint is the one the WAL from there on was
written under. That holds for a backup taken from a primary, and only
when the copied state is not node local and its watermark is below the
starting checkpoint; a backup taken from a standby keeps the copied
state. pg_rewind keeps the target's own state and clamps the
watermark to the divergence point, since a watermark set on the
target's abandoned fork could numerically cover transition records the
source wrote after the divergence.
Bump PG_CONTROL_VERSION.
Document the offline procedure for replication setups, the lockstep
one: stop all nodes, run pg_checksums on each of them, and only then
restart. An offline change writes no WAL and has no ordering against
WAL a node has not replayed yet, so a node must not stop before
replaying all WAL of its upstream, or the change is overridden on
restart.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Author: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reported-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Discussion: https://postgr.es/m/anwm6UPxoVS41QA2@bdtpg
---
doc/src/sgml/ref/pg_checksums.sgml | 87 ++-
doc/src/sgml/wal.sgml | 23 +
src/backend/access/transam/xlog.c | 571 ++++++++++++++++--
src/backend/postmaster/datachecksum_state.c | 26 +-
.../utils/activity/wait_event_names.txt | 1 +
src/bin/pg_checksums/pg_checksums.c | 10 +
src/bin/pg_controldata/pg_controldata.c | 4 +
src/bin/pg_resetwal/pg_resetwal.c | 7 +
src/bin/pg_rewind/pg_rewind.c | 35 +-
src/bin/pg_upgrade/controldata.c | 2 +-
src/include/catalog/pg_control.h | 29 +-
src/include/storage/lwlocklist.h | 1 +
src/test/modules/test_checksums/Makefile | 2 +-
src/test/modules/test_checksums/meson.build | 10 +
.../test_checksums/t/012_offline_standby.pl | 304 ++++++++++
.../modules/test_checksums/t/013_rewind.pl | 201 ++++++
.../modules/test_checksums/t/014_lockstep.pl | 182 ++++++
.../t/015_standby_crash_after_disable.pl | 137 +++++
.../t/016_promote_enable_crash.pl | 146 +++++
.../test_checksums/t/017_restartpoint_race.pl | 151 +++++
.../t/018_enable_crash_windows.pl | 508 ++++++++++++++++
.../t/019_standby_shutdown_catchup.pl | 138 +++++
.../t/020_cascade_divergence.pl | 127 ++++
.../t/021_rewind_divergent_transitions.pl | 205 +++++++
24 files changed, 2823 insertions(+), 84 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/012_offline_standby.pl
create mode 100644 src/test/modules/test_checksums/t/013_rewind.pl
create mode 100644 src/test/modules/test_checksums/t/014_lockstep.pl
create mode 100644 src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
create mode 100644 src/test/modules/test_checksums/t/016_promote_enable_crash.pl
create mode 100644 src/test/modules/test_checksums/t/017_restartpoint_race.pl
create mode 100644 src/test/modules/test_checksums/t/018_enable_crash_windows.pl
create mode 100644 src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
create mode 100644 src/test/modules/test_checksums/t/020_cascade_divergence.pl
create mode 100644 src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index ae66fad3f0f..2d2057a8c19 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -242,15 +242,79 @@ PostgreSQL documentation
data directory must not be started or else data loss may occur.
</para>
<para>
- When using a replication setup with tools which perform direct copies
- of relation file blocks (for example <xref linkend="app-pgrewind"/>),
- enabling or disabling checksums can lead to page corruptions in the
- shape of incorrect checksums if the operation is not done consistently
- across all nodes. When enabling or disabling checksums in a replication
- setup, it is thus recommended to stop all the clusters before switching
- them all consistently. Destroying all standbys, performing the operation
- on the primary and finally recreating the standbys from scratch is also
- safe.
+ Enabling or disabling checksums with
+ <application>pg_checksums</application> changes only the local data
+ directory; the new state is not replicated to any other node. In a
+ replication setup the same change must be applied to every node:
+ </para>
+
+ <procedure>
+ <step id="shut-down-nodes">
+ <title>Shut down all nodes</title>
+ <para>
+ All nodes participating in the replication must be stopped with a
+ clean shutdown; <application>pg_checksums</application> refuses to run
+ on a data directory left behind by an immediate shutdown. Before
+ stopping a standby, make sure it has replayed all WAL of the primary:
+ stop the primary first, read its <quote>Latest checkpoint
+ location</quote> with <xref linkend="app-pgcontroldata"/>, and check
+ that <function>pg_last_wal_replay_lsn()</function> on the standby has
+ advanced past it. Comparing the replay position with
+ <function>pg_last_wal_receive_lsn()</function> is not enough, as it
+ only shows that the WAL the standby received has been replayed.
+ </para>
+ </step>
+
+ <step>
+ <title>Enable or disable data checksums on each node</title>
+ <para>
+ Run <application>pg_checksums</application> on the data directory of
+ each node in the replication setup. Nodes can be processed in
+ parallel while they are shut down. Processing must complete
+ successfully on all nodes before continuing.
+ </para>
+ </step>
+
+ <step>
+ <title>Restart all nodes</title>
+ <para>
+ Start the nodes normally, verify that
+ <xref linkend="guc-data-checksums"/> matches on all of them, and
+ monitor the logs of the standbys for data checksum state mismatch
+ warnings.
+ </para>
+ </step>
+ </procedure>
+
+ <para>
+ The replay requirement in <xref linkend="shut-down-nodes"/>exists
+ because an offline change is recorded only in the control file and
+ has no defined ordering against WAL the node has not replayed yet; see
+ <xref linkend="checksums-offline-enable-disable"/>. A node stopped
+ before replaying an online checksum state change applies that change
+ when it is restarted, overriding the offline change, and the states of
+ the nodes silently diverge until a later checkpoint record triggers
+ the warning described below. Because of this it is best not to mix
+ the two mechanisms: change the state of a replication setup either
+ with the offline procedure above or with an online transition, and
+ make sure the previous change has reached every node before starting
+ the next one.
+ </para>
+ <para>
+ If the change is applied inconsistently, each node keeps its own state,
+ and a standby logs a warning when the state recorded in the replayed WAL
+ differs from its own. The same warning can appear transiently while a
+ standby catches up over WAL written before a consistent change; it stops
+ once a checkpoint record carrying the new state has been replayed.
+ </para>
+ <para>
+ A standby whose data directory was never checksummed must not have
+ checksums enabled by catching up this way. Converge the cluster by
+ running <application>pg_checksums</application> on it while stopped as
+ outlined above, by enabling checksums online, or by recreating it from
+ a base backup. Note
+ that an online enable only starts from a primary whose checksums are off,
+ so if they are already enabled there, disable them online first.
</para>
<para>
If <application>pg_checksums</application> is aborted or killed while
@@ -258,6 +322,11 @@ PostgreSQL documentation
remains unchanged, and <application>pg_checksums</application> can be
re-run to perform the same operation.
</para>
+ <para>
+ Tools that copy relation file blocks directly between nodes, such as
+ <xref linkend="app-pgrewind"/>, require all nodes to be in
+ the same data checksum state, else there is risk for data corruption.
+ </para>
<para>
The target cluster must have the same major version as
<application>pg_checksums</application>.
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index ec62d17fbcc..dbbfdae5e4f 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -317,6 +317,29 @@
verify checksums, on an offline cluster.
</para>
+ <para>
+ An offline change provides durability differently from an
+ <link linkend="checksums-online-enable-disable">online change</link>.
+ An online transition is WAL-logged: it is ordered against all other
+ WAL records, it is replayed after a crash, and it propagates to
+ standbys. An offline change is recorded only in the cluster's
+ control file: it writes no WAL, it is invisible to replication, and
+ it has no defined ordering against WAL the node has not replayed
+ yet. When a node later replays WAL that contains an online checksum
+ state change, that change takes effect on the node even if it was
+ written before the offline change was made.
+ </para>
+
+ <para>
+ An offline change only affects the data directory it is run on; the
+ new state does not propagate over replication. In a replication setup
+ the same change must be applied to all nodes while all of them are
+ stopped, as described in <xref linkend="app-pgchecksums"/>. A standby
+ whose state diverges logs a warning but keeps its local setting. The
+ mismatch persists until the states are brought together again, with
+ the offline procedure or with an online transition; do this promptly.
+ </para>
+
</sect2>
<sect2 id="checksums-online-enable-disable" xreflabel="Online Enabling of Checksums">
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index 3203f2fd4ee..1b4f62c96e5 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -556,14 +556,22 @@ typedef struct XLogCtlData
*/
XLogRecPtr lastFpwDisableRecPtr;
- /* last data_checksum_version we've seen */
+ /* current data checksum state of this node */
uint32 data_checksum_version;
+ /*
+ * Copies of control file fields with the same names, see pg_control.h for
+ * an in-depth description of these fields. Must be updated together with
+ * data_checksum_version under info_lck.
+ */
+ XLogRecPtr data_checksum_lsn;
+ bool data_checksum_is_local;
+
slock_t info_lck; /* locks shared variables shown above */
/*
* lastChecksumChangeRecPtr points to the end of the last XLOG2_CHECKSUMS
- * record inserted or replayed, i.e. the last change of
+ * record inserted or replayed which corresponds to the last change of
* data_checksum_version. InvalidXLogRecPtr if the state hasn't changed
* since the server started.
*/
@@ -690,6 +698,20 @@ static ChecksumStateType LocalDataChecksumState = 0;
*/
int data_checksums = 0;
+/*
+ * Whether replay of the next checkpoint-family record must adopt the data
+ * checksum state it carries. Set when recovery starts from a base backup,
+ * where the state at the redo point takes precedence over the control file
+ * copied with the backup later.
+ */
+static bool adoptChecksumStateFromNextCheckpoint = false;
+
+/*
+ * Sentinel for the data checksum mismatch warning tracking in
+ * CheckReplayedDataChecksumState(): no warning is outstanding.
+ */
+#define NO_WARNING_ISSUED PG_UINT32_MAX
+
/* For WALInsertLockAcquire/Release functions */
static int MyLockNo = 0;
static bool holdingAllLocks = false;
@@ -730,6 +752,8 @@ static void ValidateXLOGDirectoryStructure(void);
static void CleanupBackupHistory(void);
static void UpdateMinRecoveryPoint(XLogRecPtr lsn, bool force);
static bool PerformRecoveryXLogAction(void);
+static void CheckReplayedDataChecksumState(uint32 replayed_version);
+static void AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn);
static void InitControlFile(uint64 sysidentifier, uint32 data_checksum_version);
static void WriteControlFile(void);
static void ReadControlFile(void);
@@ -757,7 +781,7 @@ static void WALInsertLockAcquireExclusive(void);
static void WALInsertLockRelease(void);
static void WALInsertLockUpdateInsertingAt(XLogRecPtr insertingAt);
-static void XLogChecksums(uint32 new_type);
+static XLogRecPtr XLogChecksums(uint32 new_type);
/*
* Insert an XLOG record represented by an already-constructed chain of data
@@ -4292,6 +4316,8 @@ InitControlFile(uint64 sysidentifier, uint32 data_checksum_version)
ControlFile->track_commit_timestamp = track_commit_timestamp;
ControlFile->data_checksum_version = data_checksum_version;
ControlFile->data_checksum_version_init = data_checksum_version;
+ ControlFile->data_checksum_is_local = false;
+ ControlFile->data_checksum_lsn = InvalidXLogRecPtr;
/*
* Set the data_checksum_version value into XLogCtl, which is where all
@@ -4774,6 +4800,7 @@ void
SetDataChecksumsOnInProgress(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
/*
* The state transition is performed in a critical section with
@@ -4782,14 +4809,12 @@ SetDataChecksumsOnInProgress(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4827,6 +4852,8 @@ void
SetDataChecksumsOn(void)
{
uint64 barrier;
+ bool persist;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4847,23 +4874,11 @@ SetDataChecksumsOn(void)
SpinLockRelease(&XLogCtl->info_lck);
INJECTION_POINT("datachecksums-enable-checksums-delay", NULL);
+ INJECTION_POINT_LOAD("datachecksums-on-before-publish");
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_VERSION);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- SpinLockRelease(&XLogCtl->info_lck);
-
- /*
- * Update the controlfile before waiting since if we have an immediate
- * shutdown while waiting we want to come back up with checksums enabled.
- */
- LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- UpdateControlFile();
- LWLockRelease(ControlFileLock);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_VERSION);
barrier = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_CHECKSUM_ON);
@@ -4873,6 +4888,41 @@ SetDataChecksumsOn(void)
INJECTION_POINT("datachecksums-on-before-checkpoint", NULL);
RequestCheckpoint(CHECKPOINT_FORCE | CHECKPOINT_WAIT | CHECKPOINT_FAST);
+
+ INJECTION_POINT("datachecksums-on-after-checkpoint", NULL);
+
+ /*
+ * Persist "on" only now that the checkpoint has flushed the pages the
+ * transition rewrote. Crash recovery initializes verification from the
+ * control file but resumes from a checkpoint that can predate the
+ * transition, so an "on" persisted earlier would verify pages during
+ * replay whose rewrite never reached the disk: the pages found by the
+ * data checksums worker in shared buffers are not written back by its
+ * ring buffer.
+ *
+ * The checkpoint above normally persists the state itself; this covers
+ * the case where it started before the record written above and left the
+ * field alone. Crashing before this point is safe, as replay then
+ * re-establishes "on" from the full page images of the rewrite. Skip the
+ * write if the state moved on meanwhile, since whatever moved it persists
+ * its own. Compare the watermark rather than the state: a state
+ * comparison could not tell our transition from a later round trip back
+ * to "on".
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ persist = (XLogCtl->data_checksum_lsn == recptr);
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (persist)
+ {
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+ }
+
WaitForProcSignalBarrier(barrier);
}
@@ -4893,6 +4943,7 @@ void
SetDataChecksumsOff(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4918,14 +4969,12 @@ SetDataChecksumsOff(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4956,14 +5005,12 @@ SetDataChecksumsOff(void)
/* Ensure that we don't incur a checkpoint during disabling checksums */
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -5001,6 +5048,138 @@ SetLocalDataChecksumState(uint32 data_checksum_version)
data_checksums = data_checksum_version;
}
+/*
+ * CheckReplayedDataChecksumState
+ * Cross-check the data checksum state carried by a replayed checkpoint
+ * record against the state of this node.
+ *
+ * The state in the record belongs to whichever node wrote the WAL, and must
+ * not be adopted: an offline state change made with pg_checksums on one node
+ * of a replication set generates no WAL, and adopting would leak it into the
+ * other nodes through replay. Backup label recovery is the one exception,
+ * see AdoptReplayedDataChecksumState(). XLOG_CHECKPOINT_ONLINE needs no call
+ * here, as an online checkpoint's state already traveled in the preceding
+ * XLOG_CHECKPOINT_REDO record.
+ *
+ * Only archive recovery can see a lasting mismatch, as only there can the WAL
+ * and the control file come from different nodes or different times. In
+ * crash recovery a mismatch means replay resumed from a restartpoint
+ * predating an already-applied XLOG2_CHECKSUMS record, and replaying forward
+ * re-establishes the same state.
+ */
+static void
+CheckReplayedDataChecksumState(uint32 replayed_version)
+{
+ /*
+ * Warn once per remote value, so a lasting mismatch does not flood the
+ * log. Matching states re-arm the warning. Backend-local state is
+ * enough: replay only runs in the startup process, and restarting it
+ * re-arms as well.
+ */
+ static uint32 last_warned_version = NO_WARNING_ISSUED;
+ uint32 local_version;
+
+ if (!ArchiveRecoveryRequested)
+ return;
+
+ /*
+ * Re-replayed WAL below the consistency point was already cross-checked
+ * before minRecoveryPoint was last persisted, and the persisted state can
+ * legitimately be newer than what checkpoint records there carry:
+ * XLOG2_CHECKSUMS replay persists most states ahead of the restartpoint
+ * horizon. In particular the checkpoint record recovery restarts from is
+ * such a re-replay.
+ */
+ if (!reachedConsistency)
+ return;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ local_version = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (replayed_version == local_version)
+ {
+ /*
+ * Report convergence if this process warned before. Nothing else
+ * tells the operator that running pg_checksums on the other nodes, or
+ * a rebuild, took effect.
+ */
+ if (last_warned_version != NO_WARNING_ISSUED)
+ ereport(LOG,
+ errmsg("data checksum state \"%s\" of this node now agrees with the replayed WAL",
+ get_checksum_state_string(local_version)));
+
+ last_warned_version = NO_WARNING_ISSUED;
+ return;
+ }
+
+ /* the nodes legitimately differ while an online transition runs */
+ if (replayed_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ replayed_version == PG_DATA_CHECKSUM_INPROGRESS_OFF ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ return;
+
+ if (replayed_version == last_warned_version)
+ return;
+ last_warned_version = replayed_version;
+
+ ereport(WARNING,
+ errmsg("data checksum state \"%s\" of this node does not match the state \"%s\" in the replayed WAL",
+ get_checksum_state_string(local_version),
+ get_checksum_state_string(replayed_version)),
+ errdetail("The data checksum state may have been changed with pg_checksums on another node."),
+ errhint("Apply the same change with pg_checksums on the primary and all standby servers, or rebuild this server from a base backup."));
+}
+
+/*
+ * AdoptReplayedDataChecksumState
+ * Adopt the data checksum state at the redo point of backup label
+ * recovery.
+ *
+ * The state is persisted immediately so that a crash before the first
+ * restartpoint does not resurrect the state copied with the backup; a crash
+ * at this point restarts from the same redo point, so the control file does
+ * not run ahead of the replay position. If the value is unchanged the
+ * control file already carries it, so both the barrier and the persist are
+ * skipped.
+ *
+ * lsn is the location of the checkpoint-family record the state was taken
+ * from and becomes the new watermark: the adopted state covers everything
+ * below the redo point, which replay never revisits. For the same reason
+ * the old watermark can stay when the value is unchanged: the records
+ * between the two positions are never replayed, so nothing depends on
+ * which one is recorded.
+ */
+static void
+AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn)
+{
+ bool changed = false;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (XLogCtl->data_checksum_version != new_version)
+ {
+ XLogCtl->data_checksum_version = new_version;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(new_version);
+ changed = true;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (!changed)
+ return;
+
+ EmitAndWaitDataChecksumsBarrier(new_version);
+
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = new_version;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+}
+
/* guc hook */
const char *
show_data_checksums(void)
@@ -5454,6 +5633,8 @@ XLOGShmemInit(void *arg)
/* Use the checksum info from control file */
XLogCtl->data_checksum_version = ControlFile->data_checksum_version;
+ XLogCtl->data_checksum_lsn = ControlFile->data_checksum_lsn;
+ XLogCtl->data_checksum_is_local = ControlFile->data_checksum_is_local;
SetLocalDataChecksumState(XLogCtl->data_checksum_version);
SpinLockInit(&XLogCtl->Insert.insertpos_lck);
@@ -6031,6 +6212,51 @@ StartupXLOG(void)
SetCommitTsLimit(checkPoint.oldestCommitTsXid,
checkPoint.newestCommitTsXid);
+ /*
+ * When recovery starts from a base backup, the control file was copied at
+ * an arbitrary moment and its data checksum state may differ from the
+ * state at the redo point, which is what the WAL from there on was
+ * written under. Adopt the state of the starting checkpoint: a shutdown
+ * checkpoint is not replayed, so take it from the record read above; the
+ * redo point of an online checkpoint is its CHECKPOINT_REDO record, so
+ * let the replay of that record adopt it. Check backupStartPoint in
+ * addition to the label: on a crash restart during backup recovery the
+ * label file is already renamed away, but the start point persists until
+ * the backup end record.
+ *
+ * Not for a base backup taken from a standby, though. Its starting
+ * checkpoint is the standby's last restartpoint, a record written by the
+ * upstream primary, whose state is not the one the copied files were
+ * written under. The copied control file is already correct: a standby
+ * persists its state only at restartpoint horizons and never claims more
+ * than what reached disk. Such backups are recognized by backupEndPoint
+ * together with backupEndRequired; backupEndPoint is only set for "BACKUP
+ * FROM: standby" labels and persists across a crash restart. pg_rewind
+ * writes a standby label as well, but no backupEndPoint, and its recovery
+ * keeps adopting: the control file it installs carries the target's own
+ * checksum state, which can lag the redo point of the last common
+ * checkpoint the same way a restartpoint horizon can.
+ *
+ * Never adopt over a state the control file's watermark or local flag
+ * marks as newer than the starting checkpoint. A pg_checksums change is
+ * local to the node and generates no WAL, so nothing in the replayed WAL
+ * could ever restore it once overwritten; and a watermark above the redo
+ * point means the control file already contains the effect of every
+ * transition record up to there.
+ */
+ if ((haveBackupLabel || XLogRecPtrIsValid(ControlFile->backupStartPoint)) &&
+ !(XLogRecPtrIsValid(ControlFile->backupEndPoint) &&
+ ControlFile->backupEndRequired) &&
+ !ControlFile->data_checksum_is_local &&
+ checkPoint.redo > ControlFile->data_checksum_lsn)
+ {
+ if (wasShutdown)
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ checkPoint.redo);
+ else
+ adoptChecksumStateFromNextCheckpoint = true;
+ }
+
/*
* Clear out any old relcache cache files. This is *necessary* if we do
* any WAL replay, since that would probably result in the cache files
@@ -6634,11 +6860,7 @@ StartupXLOG(void)
if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
ereport(WARNING,
@@ -6655,11 +6877,7 @@ StartupXLOG(void)
else if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
}
@@ -6684,6 +6902,8 @@ StartupXLOG(void)
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
XLogCtl->SharedRecoveryState = RECOVERY_STATE_DONE;
SpinLockRelease(&XLogCtl->info_lck);
@@ -6815,6 +7035,23 @@ static bool
PerformRecoveryXLogAction(void)
{
bool promoted = false;
+ bool flushForChecksums;
+ uint32 checksum_state;
+
+ /*
+ * The end-of-recovery record persists the data checksum state without
+ * flushing the buffer pool, but the control file may only claim "on" once
+ * every page on disk carries a checksum. If replay entered that state
+ * without a restartpoint following it, the pages rewritten by the
+ * transition are still only in the buffer pool, so take the full
+ * checkpoint below instead of the lightweight record.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ flushForChecksums = (checksum_state == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile->data_checksum_version != checksum_state);
/*
* Perform a checkpoint to update all our recovery activity to disk.
@@ -6830,7 +7067,7 @@ PerformRecoveryXLogAction(void)
* fully out of recovery mode and already accepting queries.
*/
if (ArchiveRecoveryRequested && IsUnderPostmaster &&
- PromoteIsTriggered())
+ PromoteIsTriggered() && !flushForChecksums)
{
promoted = true;
@@ -7437,6 +7674,7 @@ CreateCheckPoint(int flags)
uint32 freespace;
XLogRecPtr PriorRedoPtr;
XLogRecPtr last_important_lsn;
+ XLogRecPtr checksumLsn;
VirtualTransactionId *vxids;
int nvxids;
int oldXLogAllowed = 0;
@@ -7549,11 +7787,14 @@ CreateCheckPoint(int flags)
checkPoint.wal_level = wal_level;
/*
- * Get the current data_checksum_version value from xlogctl, valid at the
- * time of the checkpoint.
+ * Get the current data_checksum_version value from xlogctl. This is
+ * final only for a shutdown checkpoint, where no concurrent transition is
+ * possible; an online checkpoint resamples it together with the redo
+ * record below.
*/
SpinLockAcquire(&XLogCtl->info_lck);
checkPoint.dataChecksumState = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
if (shutdown)
@@ -7610,10 +7851,21 @@ CreateCheckPoint(int flags)
{
xl_checkpoint_redo redo_rec;
+ /*
+ * Sample the data checksum state and insert the redo record under
+ * DataChecksumTransitionLock, so that a concurrent transition cannot
+ * insert its XLOG2_CHECKSUMS record between the sampling and the
+ * insertion below. Without this, the redo record could follow the
+ * transition record in WAL while carrying the pre-transition state,
+ * and recovery resuming here would never learn about the transition.
+ * See XLogChecksums().
+ */
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
WALInsertLockAcquire();
redo_rec.wal_level = wal_level;
SpinLockAcquire(&XLogCtl->info_lck);
redo_rec.data_checksum_version = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
WALInsertLockRelease();
@@ -7621,6 +7873,14 @@ CreateCheckPoint(int flags)
XLogBeginInsert();
XLogRegisterData(&redo_rec, sizeof(xl_checkpoint_redo));
(void) XLogInsert(RM_XLOG_ID, XLOG_CHECKPOINT_REDO);
+ LWLockRelease(DataChecksumTransitionLock);
+
+ /*
+ * The checkpoint record must carry the same state as the redo record
+ * just inserted: the sample taken before redo determination can be
+ * stale by now, and the pair would otherwise disagree.
+ */
+ checkPoint.dataChecksumState = redo_rec.data_checksum_version;
/*
* XLogInsertRecord will have updated XLogCtl->Insert.RedoRecPtr in
@@ -7824,6 +8084,41 @@ CreateCheckPoint(int flags)
ControlFile->minRecoveryPoint = InvalidXLogRecPtr;
ControlFile->minRecoveryPointTLI = 0;
+ /*
+ * Persist the data checksum state this node runs under into the control
+ * file. Only the top-level field tracks this node, checkPointCopy is a
+ * historical record used to resume replay.
+ *
+ * checkPoint.dataChecksumState was sampled while holding the
+ * DataChecksumTransitionLock together with the redo record, so it is the
+ * state in effect at the redo point. If it was "on", the XLOG2_CHECKSUMS
+ * record announcing that precedes the redo point and every page the
+ * transition rewrote was dirtied before it, so CheckPointGuts() has just
+ * written all of them out. Recording the state here is what keeps a
+ * finished transition from being resolved as interrupted when this
+ * checkpoint is the one crash recovery resumes from: replay never sees
+ * the record announcing it.
+ *
+ * Persist it only if the state did not change while the flush was in
+ * progress. If it changed in between, the pages written out straddle two
+ * states, and the newer one could claim checksums that pages already on
+ * disk do not carry; leave the field to the next checkpoint then, the
+ * transition itself has already persisted every state that is safe
+ * without a flush.
+ *
+ * Compare the watermark rather than the state: record positions are
+ * unique, so a full round trip back to the sampled state cannot alias,
+ * while its flushed pages straddle the intermediate states all the same.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksumLsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checkPoint.dataChecksumState;
+ ControlFile->data_checksum_lsn = checksumLsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
/*
* Persist unloggedLSN value. It's reset on crash recovery, so this goes
* unused on non-shutdown checkpoints, but seems useful to store it always
@@ -7968,9 +8263,11 @@ CreateEndOfRecoveryRecord(void)
ControlFile->minRecoveryPoint = recptr;
ControlFile->minRecoveryPointTLI = xlrec.ThisTimeLineID;
- /* start with the latest checksum version (as of the end of recovery) */
+ /* persist the data checksum state this node ended recovery with */
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
@@ -8175,6 +8472,9 @@ CreateRestartPoint(int flags)
XLogRecPtr endptr;
XLogSegNo _logSegNo;
TimestampTz xtime;
+ uint32 checksum_state;
+ XLogRecPtr checksum_lsn;
+ bool checksum_is_local;
/* Concurrent checkpoint/restartpoint cannot happen */
Assert(!IsUnderPostmaster || MyBackendType == B_CHECKPOINTER);
@@ -8221,8 +8521,47 @@ CreateRestartPoint(int flags)
UpdateMinRecoveryPoint(InvalidXLogRecPtr, true);
if (flags & CHECKPOINT_IS_SHUTDOWN)
{
+ bool catchUpChecksums;
+
+ /*
+ * There is no new restartpoint to persist the data checksum state
+ * with, but a cleanly stopped node should not leave the control
+ * file behind the state replay reached: pg_checksums and
+ * pg_rewind read it, and an in-progress state there makes them
+ * refuse to run. Catching it up needs the same guarantee a
+ * restartpoint gives: that every page on disk carries a checksum,
+ * so flush the buffer pool first. Only a transition to "on" that
+ * no restartpoint followed can get here; the other states are
+ * already persisted by XLOG2_CHECKSUMS replay. Replay has ended
+ * by now, so the state cannot change under us.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ checksum_is_local = XLogCtl->data_checksum_is_local;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockAcquire(ControlFileLock, LW_SHARED);
+ catchUpChecksums =
+ (checksum_lsn != ControlFile->data_checksum_lsn &&
+ XLogRecPtrIsValid(lastCheckPointRecPtr));
+ LWLockRelease(ControlFileLock);
+
+ if (catchUpChecksums)
+ {
+ MemSet(&CheckpointStats, 0, sizeof(CheckpointStats));
+ CheckpointStats.ckpt_start_t = GetCurrentTimestamp();
+ CheckPointGuts(lastCheckPoint.redo, flags);
+ }
+
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
+ if (catchUpChecksums)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = checksum_is_local;
+ }
UpdateControlFile();
LWLockRelease(ControlFileLock);
}
@@ -8263,6 +8602,17 @@ CreateRestartPoint(int flags)
/* Update the process title */
update_checkpoint_display(flags, true, false);
+ /*
+ * Note the data checksum state the flush below starts under. Replay runs
+ * concurrently and can change the state while the flush is in progress,
+ * in which case the flush covers pages written under both states; see
+ * where the state is persisted further down.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
CheckPointGuts(lastCheckPoint.redo, flags);
/*
@@ -8322,8 +8672,26 @@ CreateRestartPoint(int flags)
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
}
- /* we shall start with the latest checksum version */
- ControlFile->data_checksum_version = lastCheckPoint.dataChecksumState;
+ /*
+ * Persist the data checksum state of this node. Not the state of the
+ * replayed checkpoint: that one belongs to the node that wrote it and
+ * may differ after an offline change on either side.
+ * ControlFile->checkPointCopy above keeps the replayed value on
+ * purpose, being a historical record used to resume replay rather
+ * than a tracker of node state.
+ *
+ * Persist only if the flush above ran under one state throughout; see
+ * CreateCheckPoint() for why, including why this compares the
+ * watermark and not the state.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksum_lsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
}
@@ -8764,9 +9132,22 @@ XLogReportParameters(void)
}
/*
- * Log the new state of checksums
+ * XLogChecksums
+ * Log and publish the new state of checksums
+ *
+ * Inserting the record and publishing the new state must be atomic with
+ * respect to a checkpoint sampling the state for its XLOG_CHECKPOINT_REDO
+ * record: without that, a checkpoint could read the old state after the
+ * record is already in WAL and insert a redo record that both follows the
+ * transition in WAL order and carries the pre-transition state. Recovery
+ * resuming from such a redo point would never replay the transition record
+ * and resolve the finished transition as interrupted.
+ * DataChecksumTransitionLock serializes the two; see CreateCheckPoint().
+ *
+ * Returns the end LSN of the inserted record, which the caller persists
+ * together with the new state as the data checksum watermark.
*/
-static void
+static XLogRecPtr
XLogChecksums(uint32 new_type)
{
xl_checksum_state xlrec;
@@ -8774,12 +9155,28 @@ XLogChecksums(uint32 new_type)
xlrec.new_checksum_state = new_type;
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
+
XLogBeginInsert();
XLogRegisterData((char *) &xlrec, sizeof(xl_checksum_state));
recptr = XLogInsert(RM_XLOG2_ID, XLOG2_CHECKSUMS);
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, recptr);
+
+ /* only loaded by SetDataChecksumsOn(), a no-op for the other callers */
+ INJECTION_POINT_CACHED("datachecksums-on-before-publish", NULL);
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ XLogCtl->data_checksum_version = new_type;
+ XLogCtl->data_checksum_lsn = recptr;
+ XLogCtl->data_checksum_is_local = false;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockRelease(DataChecksumTransitionLock);
+
XLogFlush(recptr);
+
+ return recptr;
}
/*
@@ -8967,11 +9364,19 @@ xlog_redo(XLogReaderState *record)
/* ControlFile->checkPointCopy always tracks the latest ckpt XID */
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->checkPointCopy.nextXid = checkPoint.nextXid;
- ControlFile->data_checksum_version = checkPoint.dataChecksumState;
UpdateControlFile();
LWLockRelease(ControlFileLock);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(checkPoint.dataChecksumState);
+
/*
* We should've already switched to the new TLI before replaying this
* record.
@@ -9207,19 +9612,17 @@ xlog_redo(XLogReaderState *record)
else if (info == XLOG_CHECKPOINT_REDO)
{
xl_checkpoint_redo redo_rec;
- bool new_state = false;
memcpy(&redo_rec, XLogRecGetData(record), sizeof(xl_checkpoint_redo));
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = redo_rec.data_checksum_version;
- SetLocalDataChecksumState(redo_rec.data_checksum_version);
- if (redo_rec.data_checksum_version != ControlFile->data_checksum_version)
- new_state = true;
- SpinLockRelease(&XLogCtl->info_lck);
-
- if (new_state)
- EmitAndWaitDataChecksumsBarrier(redo_rec.data_checksum_version);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(redo_rec.data_checksum_version,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(redo_rec.data_checksum_version);
}
else if (info == XLOG_LOGICAL_DECODING_STATUS_CHANGE)
{
@@ -9281,25 +9684,63 @@ xlog2_redo(XLogReaderState *record)
{
xl_checksum_state state;
XLogRecPtr lsn = record->EndRecPtr;
+ XLogRecPtr watermark;
memcpy(&state, XLogRecGetData(record), sizeof(xl_checksum_state));
+ SpinLockAcquire(&XLogCtl->info_lck);
+ watermark = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ /*
+ * Skip records this node has already applied. The control file
+ * carries the watermark, so this holds across restarts: recovery
+ * resuming below a record whose effect the control file already
+ * contains must not re-apply it, or it would revert a state change
+ * made with pg_checksums in between, which moves the state without
+ * writing any record of its own.
+ */
+ if (lsn <= watermark)
+ return;
+
/* advertise the location before the new state becomes visible */
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, lsn);
SpinLockAcquire(&XLogCtl->info_lck);
XLogCtl->data_checksum_version = state.new_checksum_state;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(state.new_checksum_state);
SpinLockRelease(&XLogCtl->info_lck);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = state.new_checksum_state;
+
+ /*
+ * Persist the new state, except when it is "on". Only "on" verifies
+ * checksums during reads, and between the last restartpoint and this
+ * record there may be pages on disk flushed under the old state; a
+ * crash-restart initializes verification from the control file and
+ * replay reads those pages back, so the control file may only say
+ * "on" once everything written under the transition has been flushed,
+ * as restartpoints and the end of recovery do. The opposite direction
+ * cannot wait for the restartpoint: once this record is replayed,
+ * evicted pages are written without checksums, and a control file
+ * still saying "on" would fail verification on exactly those pages
+ * after a crash.
+ */
+ if (state.new_checksum_state != PG_DATA_CHECKSUM_VERSION)
+ {
+ ControlFile->data_checksum_version = state.new_checksum_state;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ }
/*
* Update minRecoveryPoint to ensure that if recovery is aborted, we
* recover back up to this point before allowing hot standby again.
- * The new state is durable in pg_control while its location is only
- * tracked in shared memory; a standby becoming consistent below this
- * record would let base backups resume checksum verification with the
+ * The change location is only tracked in shared memory and is lost
+ * over a restart; a standby becoming consistent below this record
+ * would let base backups resume checksum verification with the
* location unknown. The local copies cannot be updated as long as
* crash recovery is happening and we expect all the WAL to be
* replayed.
diff --git a/src/backend/postmaster/datachecksum_state.c b/src/backend/postmaster/datachecksum_state.c
index f69258bc33d..099a6b4fe2e 100644
--- a/src/backend/postmaster/datachecksum_state.c
+++ b/src/backend/postmaster/datachecksum_state.c
@@ -94,9 +94,9 @@
*
* If processing is started in an online cluster then all backends are in Bd.
* If processing was halted by the cluster shutting down (due to a crash or
- * intentional restart), the controlfile state "inprogress-on" will be observed
- * on system startup and all backends will be placed in Bd. The controlfile
- * state will also be set to "off".
+ * intentional restart), the control file state "inprogress-on" will be
+ * observed on system startup and all backends will be placed in Bd. The
+ * control file state will also be set to "off".
*
* Backends transition Bd -> Bi via a procsignalbarrier which is emitted by the
* DataChecksumsWorkerLauncherMain. When all backends have acknowledged the
@@ -146,7 +146,25 @@
* stop writing data checksums as no backend is enforcing data checksum
* validation any longer.
*
- * 4. Future opportunities for optimizations
+ * 4. Interaction with offline data checksum changes
+ * -------------------------------------------------
+ * Enabling or disabling checksums offline with pg_checksums uses none of the
+ * machinery in this file, but the two mechanisms share the state kept in the
+ * control file, so their interaction is documented here.
+ *
+ * pg_checksums writes the new state to the control file and sets
+ * data_checksum_is_local, marking a state that no WAL record accounts for.
+ * Recovery then does not adopt the state carried by a replayed checkpoint
+ * record over it. The control file also carries a watermark, the WAL
+ * position through which data checksum transitions are covered. Replay skips
+ * transition records ending at or below the watermark, as their effect is
+ * already contained in the control file, and applies records above it as
+ * usual, whether they were written before or after an offline change. This
+ * is why an offline change in a replicated setup must be made on every node
+ * while all of them are stopped and caught up; see the pg_checksums
+ * documentation for the procedure.
+ *
+ * 5. Future opportunities for optimizations
* -----------------------------------------
* Below are some potential optimizations and improvements which were brought
* up during reviews of this feature, but which weren't implemented in the
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 0a70cffa081..3d366fd1114 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -371,6 +371,7 @@ WaitLSN "Waiting to read or update shared Wait-for-LSN state."
LogicalDecodingControl "Waiting to read or update logical decoding status information."
DataChecksumsWorker "Waiting for data checksums worker."
AioWorkerControl "Waiting to update AIO worker information."
+DataChecksumTransition "Waiting for a data checksum state transition to be written to WAL."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 3b3ae23f1a6..5d6ea318784 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -648,6 +648,16 @@ main(int argc, char *argv[])
ControlFile->data_checksum_version =
(mode == PG_MODE_ENABLE) ? PG_DATA_CHECKSUM_VERSION : PG_DATA_CHECKSUM_OFF;
+ /*
+ * Mark the state as changed locally, without a WAL record. Recovery
+ * then does not let a replayed checkpoint overwrite it, as no record
+ * could restore the change afterwards. The watermark is left alone:
+ * XLOG2_CHECKSUMS records at or below it stay covered, while records
+ * above it, which this node has not applied yet, still take effect on
+ * replay no matter when they were written.
+ */
+ ControlFile->data_checksum_is_local = true;
+
if (do_sync)
{
pg_log_info("syncing data directory");
diff --git a/src/bin/pg_controldata/pg_controldata.c b/src/bin/pg_controldata/pg_controldata.c
index 6fc87ed114d..c363ce3dbb7 100644
--- a/src/bin/pg_controldata/pg_controldata.c
+++ b/src/bin/pg_controldata/pg_controldata.c
@@ -349,6 +349,10 @@ main(int argc, char *argv[])
(ControlFile->float8ByVal ? _("by value") : _("by reference")));
printf(_("Data page checksum version: %u\n"),
ControlFile->data_checksum_version);
+ printf(_("Data checksum watermark: %X/%08X\n"),
+ LSN_FORMAT_ARGS(ControlFile->data_checksum_lsn));
+ printf(_("Data checksum state is node-local: %s\n"),
+ (ControlFile->data_checksum_is_local ? _("yes") : _("no")));
printf(_("Default char data signedness: %s\n"),
(ControlFile->default_char_signedness ? _("signed") : _("unsigned")));
printf(_("Mock authentication nonce: %s\n"),
diff --git a/src/bin/pg_resetwal/pg_resetwal.c b/src/bin/pg_resetwal/pg_resetwal.c
index 79f3085d769..5efd487f318 100644
--- a/src/bin/pg_resetwal/pg_resetwal.c
+++ b/src/bin/pg_resetwal/pg_resetwal.c
@@ -923,6 +923,13 @@ RewriteControlFile(void)
ControlFile.backupEndPoint = InvalidXLogRecPtr;
ControlFile.backupEndRequired = false;
+ /*
+ * The old WAL is gone and the new position may lie below the old
+ * watermark, which would make replay ignore future checksum transition
+ * records. The state itself is kept.
+ */
+ ControlFile.data_checksum_lsn = InvalidXLogRecPtr;
+
/*
* Force the defaults for max_* settings. The values don't really matter
* as long as wal_level='minimal'; the postmaster will reset these fields
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 2e86fd158d0..0e4c2df3f4b 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -37,7 +37,8 @@ static void usage(const char *progname);
static void perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo);
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec);
static void createBackupLabel(XLogRecPtr startpoint, TimeLineID starttli,
XLogRecPtr checkpointloc);
@@ -531,7 +532,8 @@ main(int argc, char **argv)
* This is the point of no return. Once we start copying things, there is
* no turning back!
*/
- perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo);
+ perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo,
+ divergerec);
if (showprogress)
pg_log_info("syncing target data directory");
@@ -566,7 +568,8 @@ static void
perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo)
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec)
{
XLogRecPtr endrec;
TimeLineID endtli;
@@ -738,6 +741,32 @@ perform_rewind(filemap_t *filemap, rewind_source *source,
ControlFile_new.minRecoveryPoint = endrec;
ControlFile_new.minRecoveryPointTLI = endtli;
ControlFile_new.state = DB_IN_ARCHIVE_RECOVERY;
+
+ /*
+ * Keep the target's own data checksum state. Most of the data directory
+ * is still the target's: only blocks it changed since the divergence were
+ * copied from the source, so the source's state says nothing about the
+ * pages that stay. Replay from the last common checkpoint applies any
+ * WAL-logged transition the target has not seen (the watermark tells them
+ * apart), which converges the rewound server to the source's state
+ * whenever the WAL carries it.
+ */
+ ControlFile_new.data_checksum_version = ControlFile_target.data_checksum_version;
+ ControlFile_new.data_checksum_lsn = ControlFile_target.data_checksum_lsn;
+ ControlFile_new.data_checksum_is_local = ControlFile_target.data_checksum_is_local;
+
+ /*
+ * The watermark is only meaningful within the history the node replays.
+ * Records at or below the divergence point are common to both histories
+ * and stay covered, but a watermark above it was set by a transition
+ * record on the target's own abandoned fork: numerically it can cover
+ * transition records the source wrote after the divergence, and replay
+ * would skip them as already applied. Clamp it to the divergence point,
+ * so that every transition record on the source's history takes effect.
+ */
+ if (ControlFile_new.data_checksum_lsn > divergerec)
+ ControlFile_new.data_checksum_lsn = divergerec;
+
if (!dry_run)
update_controlfile(datadir_target, &ControlFile_new, do_sync);
}
diff --git a/src/bin/pg_upgrade/controldata.c b/src/bin/pg_upgrade/controldata.c
index 7543a988045..7259538de39 100644
--- a/src/bin/pg_upgrade/controldata.c
+++ b/src/bin/pg_upgrade/controldata.c
@@ -431,7 +431,7 @@ get_control_data(ClusterInfo *cluster)
cluster->controldata.date_is_int = strstr(p, "64-bit integers") != NULL;
got_date_is_int = true;
}
- else if ((p = strstr(bufin, "checksum")) != NULL)
+ else if ((p = strstr(bufin, "Data page checksum version:")) != NULL)
{
p = strchr(p, ':');
diff --git a/src/include/catalog/pg_control.h b/src/include/catalog/pg_control.h
index 7b5404460ec..3a55228d180 100644
--- a/src/include/catalog/pg_control.h
+++ b/src/include/catalog/pg_control.h
@@ -22,7 +22,7 @@
/* Version identifier for this pg_control format */
-#define PG_CONTROL_VERSION 1903
+#define PG_CONTROL_VERSION 1904
/* Nonce key length, see below */
#define MOCK_AUTH_NONCE_LEN 32
@@ -237,6 +237,33 @@ typedef struct ControlFileData
/* Current data checksums state */
uint32 data_checksum_version;
+ /*
+ * WAL position through which data checksum transitions are covered.
+ * Replay ignores XLOG2_CHECKSUMS records ending at or below this point:
+ * their effect is already contained in data_checksum_version, or an
+ * offline pg_checksums change made after they were first applied
+ * supersedes them. Ordinarily this is the end of the newest such record
+ * this node has written or applied, but a tool may store any position
+ * that covers the same set of records. If the node has never written or
+ * applied such a record this field shall be set to InvalidXLogRecPtr.
+ *
+ * The comparison has no timeline context, so the value is only valid
+ * within the WAL history this node replays. A tool that moves the node
+ * to another history must clamp the watermark to the point where the
+ * histories fork, as pg_rewind does, or reset it, as pg_resetwal does.
+ */
+ XLogRecPtr data_checksum_lsn;
+
+ /*
+ * True when data_checksum_version was last set by pg_checksums in an
+ * offline operation rather than by an online, WAL-logged transition. Such
+ * a state is local to this node and not derived from WAL, so recovery
+ * must not replace it with a state taken from a checkpoint record;
+ * nothing in the WAL could restore the change once it is overwritten.
+ * Cleared by the next WAL-logged transition.
+ */
+ bool data_checksum_is_local;
+
/*
* True if the default signedness of char is "signed" on a platform where
* the cluster is initialized.
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index d7eb648bd27..8d858be9927 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -89,6 +89,7 @@ PG_LWLOCK(54, WaitLSN)
PG_LWLOCK(55, LogicalDecodingControl)
PG_LWLOCK(56, DataChecksumsWorker)
PG_LWLOCK(57, AioWorkerControl)
+PG_LWLOCK(58, DataChecksumTransition)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/test/modules/test_checksums/Makefile b/src/test/modules/test_checksums/Makefile
index 71455cd5577..80f54bf6d8c 100644
--- a/src/test/modules/test_checksums/Makefile
+++ b/src/test/modules/test_checksums/Makefile
@@ -9,7 +9,7 @@
#
#-------------------------------------------------------------------------
-EXTRA_INSTALL = src/test/modules/injection_points
+EXTRA_INSTALL = contrib/pg_buffercache src/test/modules/injection_points
export enable_injection_points
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index fb7129d796f..e5d38fafb7d 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -35,6 +35,16 @@ tests += {
't/009_fpi.pl',
't/010_backup_straddle.pl',
't/011_standby_straddle.pl',
+ 't/012_offline_standby.pl',
+ 't/013_rewind.pl',
+ 't/014_lockstep.pl',
+ 't/015_standby_crash_after_disable.pl',
+ 't/016_promote_enable_crash.pl',
+ 't/017_restartpoint_race.pl',
+ 't/018_enable_crash_windows.pl',
+ 't/019_standby_shutdown_catchup.pl',
+ 't/020_cascade_divergence.pl',
+ 't/021_rewind_divergent_transitions.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
new file mode 100644
index 00000000000..c0ec374c8be
--- /dev/null
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -0,0 +1,304 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Offline checksum changes with pg_checksums are local to one node. A
+# standby must neither adopt the state of the primary from replayed
+# checkpoint records, nor lose its own offline change to them.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Scenario 1: enable offline on the primary only. The standby must
+# stay off, warn about the mismatch, and remain readable.
+$standby->stop;
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+test_checksum_state($primary, 'on');
+test_checksum_state($standby, 'off');
+
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after offline enable on the primary');
+
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+my @warnings = $log =~ /(does not match the state)/g;
+is(scalar(@warnings), 1, 'mismatch warned once per remote value');
+
+# Matching states re-arm the warning: undo the divergence on the primary,
+# then diverge again to the same value, all without restarting the standby.
+$primary->stop;
+$primary->checksum_disable_offline;
+$logstart = -s $standby->logfile;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no warning while the states match again');
+like($log, qr/now agrees with the replayed WAL/,
+ 'convergence is reported once the states match again');
+
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+test_checksum_state($standby, 'off');
+
+# The local state survives both clean and immediate restarts.
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+# Still in scenario 1's divergence: take a base backup *from the
+# standby*. A backup taken on a standby uses the last restartpoint as
+# its starting checkpoint (do_pg_backup_start()), so the record at the
+# redo point was written by the upstream primary and carries the
+# primary's state. Recovery from such a backup must not adopt that
+# state: the files were copied from the standby, and the new node
+# would come up verifying checksums its files do not have.
+
+# Make sure the standby has written pages under its own "off" state, so
+# its files really do lack checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,50000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', "SELECT count(*) FROM t2;");
+
+$standby->backup('from_standby');
+my $newnode = PostgreSQL::Test::Cluster->new('newnode');
+$newnode->init_from_backup($standby, 'from_standby');
+
+# Stream from the primary, not from the standby the files were copied
+# from, so the new node replays the "on" primary's records.
+$newnode->enable_streaming($primary);
+$newnode->start;
+
+# The new node's files all came from a cluster running with checksums
+# off. Anything but "off" here means it adopted the primary's state
+# through the checkpoint record at the redo point.
+my ($rc, $stdout, $stderr) = $newnode->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'the node copied from the standby accepts connections')
+ or diag("stderr: $stderr");
+is($stdout, 'off', 'backup of an "off" standby comes up with checksums off');
+
+# A checkpoint record streamed from the "on" primary must not flip the
+# state either.
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($newnode);
+test_checksum_state($newnode, 'off');
+
+# And it must be able to read the pages the standby wrote without
+# checksums.
+(undef, $stdout, $stderr) =
+ $newnode->psql('postgres', "SELECT count(*) FROM t2;");
+is($stdout, '50000', 'pages copied from the standby are readable')
+ or diag("stderr: $stderr");
+
+my $newnode_log = PostgreSQL::Test::Utils::slurp_file($newnode->logfile);
+unlike(
+ $newnode_log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the node copied from the standby');
+
+$newnode->stop('immediate');
+
+# Converge the cluster: enable offline on the standby too.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+$primary->wait_for_catchup($standby);
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after converging');
+
+# Scenario 2: disable offline on the standby only. The replayed
+# checkpoint records of the still-enabled primary must not override it.
+$standby->stop;
+$standby->checksum_disable_offline;
+$standby->start;
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'off');
+
+# Restartpoints must persist the local state, not the replayed copy.
+$standby->safe_psql('postgres', "CHECKPOINT;");
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable with checksums disabled locally');
+
+# Scenario 3: crash-restart right after an online transition, before the
+# next restartpoint. Replay then resumes from an older restartpoint whose
+# checkpoint records still carry the pre-transition state. Those must
+# still match the state seeded from the control file, and the transition
+# itself must be re-established by re-replaying the XLOG2_CHECKSUMS record,
+# without a spurious mismatch warning along the way.
+
+# Converge first: bring the standby back to "on" offline.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+disable_data_checksums($primary, wait => 'off');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+# Crash-restart the standby immediately, before any restartpoint has had a
+# chance to persist the new state to its control file.
+$logstart = -s $standby->logfile;
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001',
+ 'standby readable after crash-restart across an online transition');
+
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no spurious mismatch warning after crash-restart across an online transition'
+);
+
+# Scenario 4: a standby stopped while replaying an interrupted online
+# transition keeps the interrupted state in its own control file. A
+# primary is never caught this way, as its checksums launcher resolves
+# inprogress-on back to off from its exit cleanup. A standby has no
+# launcher; it carries forward whatever the last replayed record left
+# it in.
+
+# Block an online enable on the primary at inprogress-on with a
+# blocking temp table, same trick as in 004_offline.pl.
+my $bsession = $primary->background_psql('postgres');
+$bsession->query_safe('CREATE TEMPORARY TABLE tt (a integer);');
+enable_data_checksums($primary, wait => 'inprogress-on');
+
+# The standby picks up the in-progress state from the XLOG2_CHECKSUMS record.
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# Stop the standby cleanly; its restartpoint persists inprogress-on to
+# its own control file, since nothing on a standby resolves it away.
+$standby->stop;
+$standby->start;
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# The primary still sits at inprogress-on, so a base backup taken now
+# copies a control file and a redo point that both carry the
+# in-progress state; replay of the XLOG2_CHECKSUMS records completes
+# the transition on the new standby.
+$primary->backup('inprogress_backup');
+my $standby2 = PostgreSQL::Test::Cluster->new('standby2');
+$standby2->init_from_backup($primary, 'inprogress_backup',
+ has_streaming => 1);
+$standby2->start;
+
+# Backup label recovery adopts the redo point's state immediately,
+# before any further WAL is replayed.
+wait_for_checksum_state($standby2, 'inprogress-on');
+
+$bsession->quit;
+wait_for_checksum_state($primary, 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable once the transition completes');
+
+# Scenario 4 continued: the backup taken mid-transition must also
+# complete the transition and stay readable.
+$primary->wait_for_catchup($standby2);
+wait_for_checksum_state($standby2, 'on');
+
+is($standby2->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'mid-transition backup standby readable after the transition');
+
+$standby2->stop;
+
+# Nothing along this standby's replay can legitimately disagree: it
+# started out at the same inprogress-on state as the primary.
+my $standby2_log = PostgreSQL::Test::Utils::slurp_file($standby2->logfile);
+unlike(
+ $standby2_log,
+ qr/does not match the state/,
+ 'no mismatch warning on a backup taken mid-transition');
+
+# No extra CHECKPOINT needed: the transition's completion checkpoint
+# wrote every page with a checksum, and the shutdown restartpoint
+# flushes the rest.
+command_ok([ 'pg_checksums', '--check', '-D', $standby2->data_dir ],
+ 'checksums valid on the mid-transition backup standby');
+
+$standby->stop;
+$primary->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/013_rewind.pl b/src/test/modules/test_checksums/t/013_rewind.pl
new file mode 100644
index 00000000000..a791e24317d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/013_rewind.pl
@@ -0,0 +1,201 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test pg_rewind across an online data checksum enable.
+#
+# A clean switchover leaves the shutdown checkpoint of the old primary
+# as the last common checkpoint between the two nodes. When data
+# checksums are enabled online on the new primary before the old one is
+# rewound, pg_rewind installs the control file of the new primary, which
+# already claims checksums are fully enabled, while replay begins at the
+# shutdown checkpoint whose record still carries the old state. Replay
+# of the WAL stretch from before the enable must run with checksums off,
+# as recorded in the checkpoint, else it would verify pages which never
+# had checksums written and fail recovery.
+#
+# The new primary requests a checkpoint right after promotion, and
+# replaying its CHECKPOINT_REDO record would repair the state before
+# any interesting WAL is reached. Hold the checkpointer on the standby
+# in a restartpoint over the promotion, like in the recovery test
+# 041_checkpoint_at_promote, so that the post-promotion writes end up
+# in WAL before the first checkpoint of the new timeline.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Old primary. full_page_writes is off so that the updates done on the
+# promoted node do not carry full page images, forcing replay on the
+# rewound node to read the pages from disk. wal_log_hints is required
+# by pg_rewind on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+full_page_writes = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+log_checkpoints = on
+]);
+$node_a->start;
+
+if (!$node_a->check_extension('injection_points'))
+{
+ plan skip_all => 'Extension injection_points not installed';
+}
+
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+$node_a->safe_psql('postgres', "CREATE EXTENSION injection_points;");
+
+# Set the hint bits on t before taking the backup, so that reads on the
+# promoted node do not emit full page images for its pages later.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->safe_psql('postgres', "SELECT count(*) FROM t;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_a, 'off');
+test_checksum_state($node_b, 'off');
+
+# Hold the next restartpoint on the standby.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point', 'wait');");
+
+# Give the restartpoint a checkpoint record to work on, then start it
+# in a background session; it will block on the injection point with
+# the checkpointer busy until released.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+my $bg_psql = $node_b->background_psql('postgres', on_error_stop => 0);
+$bg_psql->query_until(
+ qr/starting_restartpoint/, q(
+ \echo starting_restartpoint
+ CHECKPOINT;
+));
+$node_b->wait_for_event('checkpointer', 'create-restart-point');
+
+# Clean switchover: the shutdown checkpoint of A streams to B and
+# becomes the last common checkpoint.
+$node_a->stop('fast');
+
+my ($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+$stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+my $shutdown_ckpt = $1;
+$node_b->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_ckpt'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+
+my $logstart = -s $node_b->logfile;
+$node_b->promote;
+
+# Accidental restart of the old primary, diverging its timeline.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->stop('fast');
+
+# Updates on the new primary before its first checkpoint; replay of
+# these on the rewound node has to read the pages from disk.
+$node_b->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 25 = 0;");
+ok( !$node_b->log_contains("checkpoint complete", $logstart),
+ "no checkpoint on the new timeline before the updates");
+
+# Release the checkpointer; the queued post-promotion checkpoint runs
+# after the updates.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+$bg_psql->quit;
+
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+# pg_rewind refuses to run with full_page_writes disabled on the
+# source; the updates it was disabled for are already in WAL.
+$node_b->safe_psql('postgres', "ALTER SYSTEM SET full_page_writes = on;");
+$node_b->safe_psql('postgres', "SELECT pg_reload_conf();");
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the new primary');
+
+# Replay on the rewound node must start at the shutdown checkpoint of
+# the switchover, with the control file of the new primary.
+my $backup_label = slurp_file($node_a->data_dir . '/backup_label');
+$backup_label =~ /^CHECKPOINT LOCATION: ([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from backup_label";
+is($1, $shutdown_ckpt, 'replay starts at the switchover checkpoint');
+
+($stdout, $stderr) = run_command(
+ [
+ 'pg_waldump',
+ '-p' => $node_a->data_dir . '/pg_wal',
+ '-t' => 1,
+ '-s' => $shutdown_ckpt,
+ '-n' => 1,
+ ]);
+like($stdout, qr/CHECKPOINT_SHUTDOWN/,
+ 'last common checkpoint is a shutdown checkpoint');
+
+# pg_rewind keeps the target's own checksum state in the control file it
+# writes; the online enable reaches the rewound node through WAL replay
+# below, not through the copied control file.
+($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+like(
+ $stdout,
+ qr/^Data page checksum version:\s*0$/m,
+ 'rewound node keeps its own checksum state in the control file');
+
+# Start the rewound node as a standby of the new primary. Replay runs
+# through the pre-enable WAL stretch and the online enable.
+#
+# The rewind replaced the configuration files with those of the new
+# primary, so put the port back.
+my $connstr = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_b->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/014_lockstep.pl b/src/test/modules/test_checksums/t/014_lockstep.pl
new file mode 100644
index 00000000000..713f37fb22c
--- /dev/null
+++ b/src/test/modules/test_checksums/t/014_lockstep.pl
@@ -0,0 +1,182 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# The lockstep procedure for offline checksum changes in a replication
+# setup: stop all nodes, run pg_checksums on all of them, restart.
+# Replay of checkpoint records written before the change must not
+# revert the state of the standby.
+#
+# The same pair then runs the procedure in the disable direction, and
+# finally across a divergence: after a failover both nodes get their
+# checksums enabled offline, while the last common checkpoint still
+# carries "off". pg_rewind must keep the offline "on" on the rewound
+# node, since no record in the replayed WAL could ever restore it.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# wal_log_hints keeps the pair eligible for pg_rewind without checksums.
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Part 1: the lockstep procedure in the enable direction. Stop the
+# standby first: the WAL written after this point is replayed only
+# after the offline switch, and every checkpoint record in it still
+# carries the old state.
+$standby->stop;
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 0;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 1;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+# The lockstep procedure.
+$primary->checksum_enable_offline;
+$standby->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to off.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "off" in the replayed WAL/,
+ 0);
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep enable');
+
+# Crash the standby and replay the same stretch again.
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'on');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after crash restart');
+
+# Once a post-switch checkpoint has been replayed the states match and
+# no warning may be logged.
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no mismatch warning once the states match');
+
+# Every page the standby wrote in this window must carry a checksum.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$standby->stop;
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'checksums valid on the standby');
+
+# Part 2: the lockstep procedure in the other direction. The standby
+# is already stopped; give it a pre-switch WAL stretch to replay, with
+# every checkpoint record in it still carrying "on" -- an explicit
+# checkpoint plus the shutdown checkpoint written by stop().
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 2;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+$primary->checksum_disable_offline;
+$standby->checksum_disable_offline;
+
+my $disable_logstart = -s $standby->logfile;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to on.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $disable_logstart);
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep disable');
+
+# Part 3: failover, divergence and pg_rewind across offline enables.
+# The roles swap from here on: the standby becomes the new primary and
+# the old primary is rewound to follow it, but the variables keep the
+# names they had above.
+#
+# Failover: promote the standby, then diverge the old primary.
+$standby->promote;
+$standby->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$primary->stop;
+
+# The lockstep procedure across the divergence: both nodes get their
+# checksums enabled offline.
+$primary->checksum_enable_offline;
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+
+# The states match, so the rewind proceeds without complaint.
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $primary->data_dir,
+ '--source-server' => $standby->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled offline on both nodes');
+
+# The rewind replaced the configuration files with those of the source,
+# so put the port back. The copied file also carries the source's own
+# stale primary_conninfo; enable_streaming() below appends ours, which
+# wins as the later entry.
+$primary->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$primary->port]}
+]);
+$primary->enable_streaming($standby);
+
+my $rewind_logstart = -s $primary->logfile;
+$primary->start;
+$standby->wait_for_catchup($primary);
+
+is($primary->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+
+# The offline enable survives: replay from the common checkpoint must
+# not resurrect the pre-divergence "off".
+test_checksum_state($primary, 'on');
+
+$log =
+ PostgreSQL::Test::Utils::slurp_file($primary->logfile, $rewind_logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no divergence reported between the rewound server and its source');
+
+$primary->stop;
+$standby->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
new file mode 100644
index 00000000000..b9a79174546
--- /dev/null
+++ b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
@@ -0,0 +1,137 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that a standby crashing after a replayed online disable, but before
+# its next restartpoint, restarts cleanly.
+#
+# Pages dirtied before the XLOG2_CHECKSUMS record and evicted after it are
+# written out under the new "off" state, without checksums. The control
+# file must follow the record immediately in this direction: a crash-restart
+# initializes checksum verification from the control file, and one still
+# saying "on" would fail verification on exactly those pages while replaying
+# records older than the state change.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$primary->start;
+
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,1000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+
+# Small buffer pool so that replaying a bulk load evicts the pages we care
+# about, and no restartpoints so the control file stays where it is.
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 1MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_delay = 10000
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'on');
+
+# No full page images, so that replay has to read the pages of "t" back from
+# disk instead of overwriting them from the WAL.
+$primary->append_conf('postgresql.conf', 'full_page_writes = off');
+$primary->reload;
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" on the standby, *before* the state change. They are
+# not flushed: the standby has no restartpoint from here on.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+# Online disable. The standby replays it and switches to "off", but its
+# control file is not updated.
+$primary->safe_psql('postgres', 'SELECT pg_disable_data_checksums();');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+my ($ctl_before) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($ctl_state) = $ctl_before =~ /Data page checksum version:\s+(\d+)/;
+note(
+ "standby control file data_checksum_version after the disable: "
+ . "$ctl_state (0 = off, 1 = on)");
+
+# Force the standby to evict the dirty pages of "t" now that it is "off":
+# they are written back without checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE filler AS SELECT generate_series(1,300000) AS a;");
+$primary->wait_for_catchup($standby);
+
+# Crash the standby before it gets a chance to run a restartpoint.
+$standby->stop('immediate');
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'standby restarts after crashing between a checksum state '
+ . 'change and the next restartpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'the pages written while "off" are readable')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("standby log tail:\n" . join("\n", @lines[ -12 .. -1 ]))
+ if @lines >= 12;
+ fail('the pages written while "off" are readable');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/016_promote_enable_crash.pl b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
new file mode 100644
index 00000000000..13c983136ef
--- /dev/null
+++ b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
@@ -0,0 +1,146 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a standby promoted right after replaying an online enable, before any
+# restartpoint has flushed the rewritten pages.
+#
+# The end-of-recovery record persists the checksum state without flushing the
+# buffer pool, while the control file still points at a restartpoint older than
+# the transition. Persisting "on" there would make a crash before the
+# post-promotion checkpoint verify checksums over the pages that were flushed
+# while the state was still "off"; promotion must take the full checkpoint
+# instead.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Dirty the pages of "t" without full page images, then push them out to disk
+# on the standby while checksums are still off.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Online enable on the primary; the standby replays it into its own buffers,
+# where the rewritten pages stay dirty (no restartpoint, no bgwriter).
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+note("standby control file before the promotion: $before");
+
+# Promote. The end-of-recovery record persists the live state without any
+# buffer flush; the checkpoint requested afterwards is not immediate.
+$standby->promote;
+$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery();')
+ or die 'timed out waiting for the promotion';
+$standby->stop('immediate');
+
+($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($ckpt) = $ctl =~ /Latest checkpoint location:\s+(\S+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note(
+ "standby control file after the promotion: $after, checkpoint $ckpt, redo $redo"
+);
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'promoted node restarts after crashing before its first checkpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/017_restartpoint_race.pl b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
new file mode 100644
index 00000000000..10919c9741c
--- /dev/null
+++ b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
@@ -0,0 +1,151 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a restartpoint whose flush races the replay of an online enable.
+#
+# Replay keeps running while CheckPointGuts() writes out the buffer pool, so
+# the state at the end of the flush can be newer than the one the flush ran
+# under. Persisting it would claim checksums for pages the flush wrote out
+# before the transition, against a redo pointer that predates it.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# C1: the checkpoint the pending restartpoint will target.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" without full page images and push them to disk on
+# the standby while it is still "off".
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Start a restartpoint for C1 and hold it right after CheckPointGuts().
+$standby->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point','wait');");
+my $bg = $standby->background_psql('postgres');
+$bg->query_until(qr//, "\\echo restartpoint\nCHECKPOINT;\n");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint injection point';
+
+# The startup process keeps replaying while the checkpointer is held: the
+# whole online enable lands, and the rewritten pages stay dirty.
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+# Let the restartpoint finish; it persists the state it samples now.
+$standby->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint to finish';
+$standby->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+
+$standby->stop('immediate');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note("standby control file after the restartpoint: $after, redo $redo");
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started, 'standby restarts after a restartpoint that raced the enable');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/018_enable_crash_windows.pl b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
new file mode 100644
index 00000000000..4dac590e819
--- /dev/null
+++ b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
@@ -0,0 +1,508 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Crash and checkpoint windows inside an online data checksum enable on
+# a single node. Each scenario starts from "off" on the same cluster,
+# runs the enable into a held injection point, and checks what a crash
+# or a concurrent checkpoint in that window may and may not do to the
+# control file and the pages on disk.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+use IPC::Run;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum_extended' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum_extended\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Bring the node back to "off" with no leftover table between
+# scenarios. The crash restarts already resolve an interrupted enable
+# to "off", so only disable when something is left to disable. Those
+# restarts also clear any attached injection points. The checkpoint
+# bounds the WAL the next scenario's crash recovery has to replay.
+sub reset_scenario
+{
+ my ($node) = @_;
+
+ BAIL_OUT('node is not running after the previous scenario')
+ unless $node->is_alive;
+
+ my $state = $node->safe_psql('postgres', 'SHOW data_checksums;');
+ disable_data_checksums($node, wait => 'off') if $state ne 'off';
+ $node->safe_psql('postgres', 'DROP TABLE IF EXISTS t;');
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+ test_checksum_state($node, 'off');
+}
+
+my $node = PostgreSQL::Test::Cluster->new('crash_windows');
+$node->init(no_data_checksums => 1);
+
+# Scenario 1 needs the pages of "t" to reach disk without a full page
+# image, hence full_page_writes and wal_log_hints off. Scenario 2
+# needs them to stay dirty in shared buffers until a checkpoint, hence
+# no bgwriter and a large shared_buffers. wal_level is pinned because
+# init() writes "minimal" for a non-streaming node. The rest merely
+# tolerate the settings.
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+wal_level = replica
+max_wal_size = 10GB
+full_page_writes = off
+wal_log_hints = off
+shared_buffers = 512MB
+bgwriter_lru_maxpages = 0
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$node->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+
+test_checksum_state($node, 'off');
+
+# Scenario 1: a primary crashes inside SetDataChecksumsOn(), just before the
+# forced checkpoint that flushes the rewritten pages, with crash recovery
+# resuming from a checkpoint older than the transition. The control file
+# must still say "inprogress-on" there: replay does not adopt the state of
+# the checkpoint record it resumes from, so an "on" written before the flush
+# would stay in effect while replay reads pages whose rewrite never reached
+# disk. Here the pages went out through the rewriting worker's ring buffer,
+# so only the control file state discriminates; see the next scenario for
+# the case where they did not.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Establish the checkpoint that crash recovery will resume from.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty the pages of "t" without full page images, then push them out
+ # to disk while checksums are still off.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+ my $evicted = $node->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+ note("evict_relation on t: $evicted, relpath $relpath");
+
+ # Stop the enable right after the control file has been updated to
+ # "on" but before the checkpoint that flushes the rewritten pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+ $node->safe_psql('postgres', 'SELECT pg_enable_data_checksums();');
+
+ $node->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'datachecksums-on-before-checkpoint';"
+ ) or die 'timed out waiting for the injection point';
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ $node->stop('immediate');
+
+ # Show the on-disk checksum field of the first page of "t".
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 2: the same crash as above, but with the rewritten pages resident
+# in shared buffers instead of gone out through the ring. The rewriting
+# worker reads through a BAS_VACUUM ring, which writes pages back as the
+# ring recycles, but a page already resident in shared buffers is not read
+# through the ring: ReadBufferExtended() hands back the existing buffer, and
+# it stays dirty until a checkpoint. The control file may therefore not
+# say "on" before that checkpoint has run, or crash recovery would resume
+# from a checkpoint older than the transition with verification already
+# enabled, and replay records that read those still-unchecksummed pages.
+# full_page_writes is off so that the records do not simply overwrite the
+# pages with a full page image.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,100000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Establish the checkpoint that crash recovery will resume from, with the
+ # pages of "t" written out while checksums are still off.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty those pages again without emitting full page images, and leave
+ # them in shared buffers. Replay of these records has to read the
+ # pages from disk.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+
+ my $dirty_before = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_before, '>', 0,
+ 'pages of t are resident and dirty before enabling checksums');
+
+ # Hold the enabling right before the checkpoint that flushes the rewritten
+ # pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ # The rewritten pages must still be sitting dirty in shared buffers, or
+ # the window this test is about does not exist.
+ my $dirty_after = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_after, '>', 0,
+ 'rewritten pages of t are still dirty before the checkpoint');
+
+ $node->stop('immediate');
+
+ # The on-disk copy is the one written before the transition, without a
+ # checksum.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ is($pd_checksum, 0, 'block 0 of t on disk carries no checksum');
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 3: a checkpoint that runs between the XLOG2_CHECKSUMS("on")
+# record and the control file write at the end of SetDataChecksumsOn() must
+# not make a crash throw the completed transition away. SetDataChecksumsOn()
+# writes the record, flips shared memory to "on", emits the barrier and only
+# then requests the checkpoint that flushes the rewritten pages; the control
+# file is written after that checkpoint returns. A crash in that window is
+# harmless only as long as recovery still starts before the record. Any
+# checkpoint completing in the window moves the redo point past the record,
+# so recovery would never see it, would come up with the control file's
+# "inprogress-on" and StartupXLOG() would demote that to "off", even though
+# every page on disk carries a checksum by then. CreateCheckPoint()
+# therefore persists the state the checkpoint ran under, the same way
+# CreateRestartPoint() does. Here the window is made deterministic by
+# holding the launcher at the datachecksums-on-before-checkpoint injection
+# point and checkpointing from another session.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Hold the launcher after the record, the shared memory flip and the
+ # barrier, but before the checkpoint SetDataChecksumsOn() requests
+ # itself.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ # Every backend already sees "on" and writes checksums.
+ test_checksum_state($node, 'on');
+
+ # ... while the control file still says "inprogress-on".
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($before, '3', 'control file says "inprogress-on" inside the window');
+
+ # A concurrent checkpoint. It flushes the rewritten pages and moves
+ # the redo point past the XLOG2_CHECKSUMS("on") record, so it has to
+ # record the "on" state in the control file as well.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($after, '1',
+ 'the concurrent checkpoint records "on" in the control file');
+
+ $node->stop('immediate');
+
+ # The checkpoint flushed the rewritten pages, so they carry a checksum on
+ # disk: the transition really did complete.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ isnt($pd_checksum, 0, 'block 0 of t on disk carries a checksum');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition is complete on disk, so the cluster has to come back
+ # "on".
+ test_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+}
+reset_scenario($node);
+
+# Scenario 4: a crash between the checkpoint an online enable requests and
+# the control file write that follows it must not lose the transition. The
+# last steps of SetDataChecksumsOn() are
+#
+# WAL record -> shmem -> barrier -> checkpoint -> persist
+#
+# The checkpoint is what makes "on" safe to persist, but it also moves the
+# redo point above the XLOG2_CHECKSUMS record that carries the new state.
+# Crash recovery started from that checkpoint therefore never replays the
+# record, so without the state the checkpoint itself persists, a crash in
+# the remaining window would bring the cluster back at "inprogress-on",
+# which StartupXLOG() resolves to "off", discarding a transition whose
+# pages are all on disk with a checksum. The previous scenario exercises
+# the same window through a checkpoint requested by another session; this
+# one closes it from the other side: the transition's own checkpoint is the
+# one that moves the redo point, so persisting the state from
+# CreateCheckPoint() is what has to save it, not the write below the
+# injection point.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Hold the launcher after the checkpoint that licenses "on" has
+ # completed and before the state reaches the control file.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-after-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-after-checkpoint');
+
+ # The transition is complete as far as the running cluster is concerned.
+ test_checksum_state($node, 'on');
+
+ # The checkpoint has already recorded it, so the pending write below the
+ # injection point has nothing left to do.
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($version) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($version, '1',
+ 'the requested checkpoint recorded "on" in the control file');
+
+ # Crash before the control file write that follows the checkpoint.
+ $node->stop('immediate');
+
+ $node->start;
+
+ # Every page on disk carries a checksum and the checkpoint that
+ # flushed them completed, so the cluster has to come back verifying
+ # them.
+ test_checksum_state($node, 'on');
+
+ is($node->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '10000', 'relation readable after the crash');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ # The data directory must be one the offline tools accept.
+ $node->stop;
+ $node->command_ok([ 'pg_checksums', '--check', '-D', $node->data_dir ],
+ 'pg_checksums accepts the data directory');
+
+ # Leave the node running for the reset.
+ $node->start;
+}
+reset_scenario($node);
+
+# Scenario 5: a checkpoint racing SetDataChecksumsOn() between the insertion
+# of the XLOG2_CHECKSUMS("on") record and the shared memory update must not
+# insert an XLOG_CHECKPOINT_REDO record that follows the transition in WAL
+# order while still carrying "inprogress-on". Recovery resuming from such a
+# redo point never replays the preceding transition record, comes up in
+# "inprogress-on" and resolves the finished transition as interrupted.
+# XLogChecksums() closes the window by inserting the record and publishing
+# the new state under DataChecksumTransitionLock, which CreateCheckPoint()
+# takes around sampling the state and inserting the redo record. Here the
+# launcher is held between the two steps at the
+# datachecksums-on-before-publish injection point, and a concurrent
+# CHECKPOINT has to block on the lock instead of completing inside the
+# window.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # The datachecksums-on-before-publish point fires inside a critical
+ # section, where the wait machinery must not allocate. Waiting once
+ # at the datachecksums-enable-checksums-delay point, which the
+ # launcher runs outside the critical section, initializes it; see
+ # 050_redo_segment_missing.pl for the same recipe around
+ # create-checkpoint-run.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-enable-checksums-delay','wait');"
+ );
+
+ # Hold the launcher after the XLOG2_CHECKSUMS("on") record is in WAL but
+ # before the new state is published in shared memory.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-publish','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-enable-checksums-delay');
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-enable-checksums-delay');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-enable-checksums-delay');");
+
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-publish');
+
+ # The record is in WAL, the published state is still the old one.
+ test_checksum_state($node, 'inprogress-on');
+
+ # A concurrent checkpoint. It must block on DataChecksumTransitionLock
+ # before inserting its redo record rather than complete inside the window.
+ my $checkpointer = IPC::Run::start(
+ [ 'psql', '-XAtq', '-d', $node->connstr('postgres'), '-c', 'CHECKPOINT;' ],
+ '<' => '/dev/null',
+ '>' => '/dev/null',
+ '2>' => '/dev/null',
+ IPC::Run::timer($PostgreSQL::Test::Utils::timeout_default));
+
+ ok( $node->poll_query_until(
+ 'postgres',
+ "SELECT wait_event = 'DataChecksumTransition' "
+ . "FROM pg_stat_activity WHERE backend_type = 'checkpointer';"),
+ 'concurrent checkpoint blocks on DataChecksumTransitionLock');
+
+ # Release the launcher; the checkpoint then samples the published "on" and
+ # its redo record follows the transition record.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-publish');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-publish');");
+
+ $checkpointer->finish;
+
+ # Crash while the launcher's own checkpoint may still be in flight.
+ # Recovery resumes from the concurrent checkpoint's redo point, which
+ # now lies above the transition record and carries "on".
+ $node->stop('immediate');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition completed, so the cluster has to come back "on".
+ wait_for_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ $node->stop;
+}
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
new file mode 100644
index 00000000000..abcc234ab22
--- /dev/null
+++ b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
@@ -0,0 +1,138 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A standby stopped after replaying an online enable, but before the
+# checkpoint record that follows it on the primary.
+#
+# The shutdown restartpoint has no new checkpoint record to work from
+# and is skipped, so the control file would keep the in-progress state
+# the transition passed through, even though replay left the node at
+# "on" and rewrote every page. pg_checksums reads that field and
+# refuses to run on an in-progress state, so the shutdown has to flush
+# and catch it up instead.
+#
+# The caught-up control file then carries the watermark of the "on"
+# record, which the second half of this test relies on: an offline
+# disable made while the standby is down must survive the re-replay of
+# that record on the next startup, or the change would be silently
+# reverted.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+wal_keep_size = 1GB
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Anchor the standby's last restartpoint here, so that nothing replayed
+# from now on gives the shutdown restartpoint a newer checkpoint record
+# to use.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Hold the enable right after the state change record, before the
+# checkpoint it requests once the transition is complete.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+enable_data_checksums($primary);
+$primary->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+wait_for_checksum_state($primary, 'on');
+
+# The state change record is already flushed and streams on its own;
+# no checkpoint record follows while the launcher is held.
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+$standby->stop('fast');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($state, '1',
+ 'shutdown catches the control file up with the replayed state');
+
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'pg_checksums verifies the stopped standby');
+
+# Disable checksums offline while the standby is down. The restart
+# below resumes replay under the last restartpoint, re-reading the "on"
+# record; the watermark persisted with the catch-up above marks it as
+# already applied, so the offline change must survive.
+$standby->checksum_disable_offline;
+
+# Let the enable finish on the primary.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-checkpoint');");
+$primary->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-checkpoint');");
+
+# The launcher exits only after the checkpoint it requests completes,
+# so a checkpoint record now follows the transition in WAL.
+$primary->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_catalog.pg_stat_activity "
+ . "WHERE backend_type = 'datachecksums launcher';");
+
+my $log_offset = -s $standby->logfile;
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+
+# The nodes legitimately diverged, which replayed checkpoints report.
+$standby->wait_for_log(
+ qr/data checksum state "off" of this node does not match the state "on" in the replayed WAL/,
+ $log_offset);
+
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/020_cascade_divergence.pl b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
new file mode 100644
index 00000000000..8f2816ac405
--- /dev/null
+++ b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
@@ -0,0 +1,127 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# An offline data checksum state change made on one node of a cascading setup
+# is reported all the way down the chain.
+#
+# CheckReplayedDataChecksumState() is reached from xlog_redo() when a
+# checkpoint record (XLOG_CHECKPOINT_SHUTDOWN or XLOG_CHECKPOINT_REDO) is
+# replayed after consistency has been reached. Those records are written by
+# the root primary only and are relayed verbatim by every intermediate
+# standby, so a cascaded standby that was changed offline learns about the
+# divergence too. An intermediate standby's own restartpoints are not
+# WAL-logged and cannot surface it.
+#
+# Note that this makes detection checkpoint-driven: an idle or read-only
+# primary surfaces the divergence no sooner than its next checkpoint, so up to
+# checkpoint_timeout may pass before the operator sees anything. Closing that
+# window would require comparing the state when a walreceiver connects.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+# checkpoint_timeout is deliberately long so that the only checkpoint record in
+# this test is the one requested explicitly below.
+my $conf = qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+);
+
+my $primary = PostgreSQL::Test::Cluster->new('cascade_primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', $conf);
+$primary->start;
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,1000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('cascade_standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf('postgresql.conf', $conf);
+$standby->start;
+
+$standby->backup('backup');
+my $cascade = PostgreSQL::Test::Cluster->new('cascade_cascade');
+$cascade->init_from_backup($standby, 'backup', has_streaming => 1);
+$cascade->append_conf('postgresql.conf', $conf);
+$cascade->start;
+
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+test_checksum_state($cascade, 'off');
+
+# Change the cascaded standby offline, and only it.
+$cascade->stop;
+$cascade->command_ok([ 'pg_checksums', '--enable', '-D', $cascade->data_dir ],
+ 'pg_checksums enables checksums on the cascaded standby only');
+
+my $logstart = -s $cascade->logfile;
+$cascade->start;
+
+test_checksum_state($cascade, 'on');
+
+# Ordinary WAL traffic carries no state, so the cascaded standby keeps
+# streaming from a chain whose state is "off" without noticing.
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (1);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+is($cascade->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '1001', 'cascaded standby keeps streaming from a divergent chain');
+
+# Restartpoints on the intermediate standby are not WAL-logged, so this does
+# not surface anything on the cascaded standby either.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (2);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+my $log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+unlike(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'an intermediate restartpoint does not surface the divergence');
+
+# A checkpoint on the root primary does: the record is relayed down the whole
+# chain, so both the direct and the cascaded standby compare it against their
+# own state. wait_for_catchup() below waits for replay, not just receipt, so
+# the record has been through xlog_redo() on both nodes once it returns.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+$log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+like(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'the cascaded standby reports the divergence on the primary checkpoint');
+
+$cascade->stop;
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
new file mode 100644
index 00000000000..3f660623bd3
--- /dev/null
+++ b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
@@ -0,0 +1,205 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind across online data checksum transitions on both sides of a
+# divergence, with the two nodes trading roles between the scenarios.
+#
+# Scenario 1: after a switchover the new primary enables checksums
+# online, while the old primary restarts on its old timeline, advances
+# its WAL beyond the enable records and runs an online enable/disable
+# cycle of its own. The old primary ends "off" with a checksum
+# watermark numerically above every checksum record the new primary has
+# written. pg_rewind clamps the watermark it keeps to the divergence
+# point; without the clamp, replay on the rewound node would skip the
+# source's enable as already applied and stay "off" under an "on"
+# primary.
+#
+# Scenario 2: checksums are disabled again, and after another
+# switchover both nodes enable them online independently, so the
+# divergence checkpoint carries "off" while both control files say
+# "on". The rewind is allowed, the target keeps its own "on" state,
+# and replay re-walks the source's enable onto the already enabled
+# node.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum_extended' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum_extended\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+sub controldata_watermark
+{
+ my ($node) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $node->data_dir ]);
+ $stdout =~ /^Data checksum watermark:\s*([0-9A-F]+)\/([0-9A-F]+)$/m
+ or die "watermark missing from pg_controldata output";
+ return (hex($1) << 32) + hex($2);
+}
+
+# Wait until the standby has replayed the shutdown checkpoint of the
+# stopped primary, so that a subsequent promotion diverges after it and
+# the shutdown checkpoint becomes the last common checkpoint.
+sub wait_for_shutdown_checkpoint_replay
+{
+ my ($primary, $standby) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $primary->data_dir ]);
+ $stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+ my $shutdown_checkpoint = $1;
+
+ $standby->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_checkpoint'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+}
+
+# Old primary, checksums off. wal_log_hints is required by pg_rewind
+# on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+# Clean switchover to B; enable checksums online on it.
+$node_a->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_a, $node_b);
+$node_b->promote;
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+$node_b->stop('fast');
+my $watermark_b = controldata_watermark($node_b);
+$node_b->start;
+
+# Accidental restart of the old primary on the old timeline. Advance
+# its WAL beyond the enable watermark of B, then run an online enable
+# and disable cycle: the node ends "off" with a watermark above every
+# checksum record B has written.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t_pad AS SELECT generate_series(1,200000) AS a;");
+enable_data_checksums($node_a, wait => 'on');
+disable_data_checksums($node_a, wait => 'off');
+test_checksum_state($node_a, 'off');
+$node_a->stop('fast');
+
+my $watermark_a = controldata_watermark($node_a);
+die "test broken: target watermark not above the source's enable"
+ unless $watermark_a > $watermark_b;
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the promoted node');
+
+# Start the rewound node as a standby of B. Replay from the last
+# common checkpoint runs through B's online enable, which the target
+# never saw, so the rewound node must converge to "on".
+my $connstr_b = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr_b application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+# Scenario 2, reusing the pair with the roles reversed. Disable
+# checksums online so the next divergence point carries "off", and let
+# A replay the change.
+disable_data_checksums($node_b, wait => 'off');
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'off');
+
+# Clean switchover back to A; enable checksums online on it.
+$node_b->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_b, $node_a);
+$node_a->promote;
+enable_data_checksums($node_a, wait => 'on');
+test_checksum_state($node_a, 'on');
+my $source_enable_watermark = controldata_watermark($node_a);
+
+# The old primary restarts on its old timeline and enables checksums
+# online independently: both control files say "on", the divergence
+# checkpoint says "off".
+$node_b->start;
+$node_b->safe_psql('postgres', "INSERT INTO t_div VALUES (2);");
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+$node_b->stop('fast');
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_b->data_dir,
+ '--source-server' => $node_a->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled on both nodes');
+
+my $connstr_a = $node_a->connstr;
+$node_b->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_b->port]}
+primary_conninfo = '$connstr_a application_name=@{[$node_b->name]}'
+]);
+$node_b->set_standby_mode;
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_b, 'on');
+
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_b->stop('fast');
+is(controldata_watermark($node_b), $source_enable_watermark,
+ 'rewound node replayed the source checksum transition');
+command_ok([ 'pg_checksums', '--check', '-D', $node_b->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the twice-rewound node');
+
+done_testing();
--
2.39.3 (Apple Git-146)
=
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-10 10:05 Daniel Gustafsson <daniel@yesql.se>
parent: Daniel Gustafsson <daniel@yesql.se>
0 siblings, 1 reply; 43+ messages in thread
From: Daniel Gustafsson @ 2026-09-10 10:05 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-hackers@lists.postgresql.org
The attached v15 combines the v14 posted yesterday (again, no code changes
here) with the small cleanups from the other thread [1] with LLM review. I
added another one to the list after PGT-6 identified that page zeroing was
logged when ignore_checksum_failures was set, which avoids the zeroing. This
is what I have running in CI and locally, and will push today to close the open
item unless I hear any objections.
--
Daniel Gustafsson
[1] https://postgr.es/m/E15AC050-C4B5-488D-BB2D-3C7AC9F89EA8@yesql.se <mailto:E15AC050-C4B5-488D-BB2D-3C7AC9F89EA8@yesql.se>
Attachments:
[application/octet-stream] v15-0009-Only-log-page-zeroing-when-checksums-aren-t-igno.patch (1.4K, ../../4BCE2F21-B669-46F3-B781-010F71B0B065@yesql.se/2-v15-0009-Only-log-page-zeroing-when-checksums-aren-t-igno.patch)
download | inline diff:
From 3532513a54fc155c38c1fc8bfd150b5824f996d9 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 10 Sep 2026 10:39:36 +0200
Subject: [PATCH v15 9/9] Only log page zeroing when checksums aren't ignored
When checksum failures are ignored, the page won't be zeroed on error
even if zero_damaged_pages is set. Found via review by GPT-6.
Author: Daniel Gustafsson <daniel@yesql.se>
Reviewed-by: tbd
Discussion: https://postgr.es/m/..
Backpatch-through: 19
---
src/backend/storage/page/bufpage.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/src/backend/storage/page/bufpage.c b/src/backend/storage/page/bufpage.c
index 8df74618fcd..e112782202c 100644
--- a/src/backend/storage/page/bufpage.c
+++ b/src/backend/storage/page/bufpage.c
@@ -160,7 +160,7 @@ PageIsVerified(PageData *page, BlockNumber blkno, int flags, bool *checksum_fail
if ((flags & (PIV_LOG_WARNING | PIV_LOG_LOG)) != 0)
ereport(flags & PIV_LOG_WARNING ? WARNING : LOG,
(errcode(ERRCODE_DATA_CORRUPTED),
- (flags & PIV_ZERO_BUFFERS_ON_ERROR) ?
+ ((flags & PIV_ZERO_BUFFERS_ON_ERROR) && (flags & ~PIV_IGNORE_CHECKSUM_FAILURE)) ?
errmsg("page verification failed, calculated checksum %u but expected %u, buffer will be zeroed",
checksum, p->pd_checksum) :
errmsg("page verification failed, calculated checksum %u but expected %u",
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v15-0008-Only-reset-launcher-state-in-exit-handler.patch (1.3K, ../../4BCE2F21-B669-46F3-B781-010F71B0B065@yesql.se/3-v15-0008-Only-reset-launcher-state-in-exit-handler.patch)
download | inline diff:
From 398487e6ac2033ee5b238e3581ce47c70ef91dd9 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Tue, 8 Sep 2026 22:42:21 +0200
Subject: [PATCH v15 8/9] Only reset launcher state in exit handler
The launcher_running flag was reset to false during shutdown, which
left a small window where a new launcher could be started when the
exit handler was still running. Fix by only updating the running
flag during the exit handler.
Found by Noah using AI assisted review with Claude. Backpatch to
v19 where the online checksums feature was introduced.
Reported-by: Noah Misch <noah@leadboat.com>
Discussion: https://postgr.es/m/...
Backpatch-through: 19
---
src/backend/postmaster/datachecksum_state.c | 2 --
1 file changed, 2 deletions(-)
diff --git a/src/backend/postmaster/datachecksum_state.c b/src/backend/postmaster/datachecksum_state.c
index 099a6b4fe2e..86e378d8951 100644
--- a/src/backend/postmaster/datachecksum_state.c
+++ b/src/backend/postmaster/datachecksum_state.c
@@ -1395,8 +1395,6 @@ done:
/* Shut down progress reporting as we are done */
pgstat_progress_end_command();
- launcher_running = false;
- DataChecksumState->launcher_running = false;
LWLockRelease(DataChecksumsWorkerLock);
}
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v15-0007-Improve-error-message-for-checksum-state-in-pg_u.patch (1.6K, ../../4BCE2F21-B669-46F3-B781-010F71B0B065@yesql.se/4-v15-0007-Improve-error-message-for-checksum-state-in-pg_u.patch)
download | inline diff:
From b45de7ceaee3a986439a7ca11a3ae770324ec43a Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Mon, 7 Sep 2026 12:22:01 +0200
Subject: [PATCH v15 7/9] Improve error message for checksum state in
pg_upgrade
When attempting to upgrade a cluster which is an inprogress state, the
same error message was used regardless of which state it was. Fix by
using different error messages for the different inprogress states.
Found by Noah using AI assisted review with Claude. Backpatch to v19
where the online checksums feature was introduced.
Reported-by: Noah Misch <noah@leadboat.com>
Discussion: https://postgr.es/m/...
Backpatch-through: 19
---
src/bin/pg_upgrade/controldata.c | 4 +++-
1 file changed, 3 insertions(+), 1 deletion(-)
diff --git a/src/bin/pg_upgrade/controldata.c b/src/bin/pg_upgrade/controldata.c
index 759bd4ef77c..e83f2d963cd 100644
--- a/src/bin/pg_upgrade/controldata.c
+++ b/src/bin/pg_upgrade/controldata.c
@@ -658,8 +658,10 @@ check_control_data(ControlData *oldctrl,
* upgrade. The user should either let the process finish, or turn off
* data checksums, before retrying.
*/
- if (oldctrl->data_checksum_version > PG_DATA_CHECKSUM_VERSION)
+ if (oldctrl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON)
pg_fatal("data checksums are being enabled in the old cluster");
+ if (oldctrl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ pg_fatal("data checksums are being disabled in the old cluster");
/*
* We might eventually allow upgrades from checksum to no-checksum
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v15-0006-Fix-enum-value-visibility-for-data_checksums.patch (1.6K, ../../4BCE2F21-B669-46F3-B781-010F71B0B065@yesql.se/5-v15-0006-Fix-enum-value-visibility-for-data_checksums.patch)
download | inline diff:
From 11fe7f7d7ddf653ece03be486878741b68e3c611 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Sat, 5 Sep 2026 20:20:45 +0200
Subject: [PATCH v15 6/9] Fix enum value visibility for data_checksums
When the data_checksums GUC was changed into an enum the enum values
were all incorrectly marked as hidden, which made pg_settings report
an empty array. Fix by setting all as visible as they should be.
Found by Noah using AI assisted review with GPT. Backpatch to v19
where the GUC was changed in the online checksums feature.
Reported-by: Noah Misch <noah@leadboat.com>
Discussion: https://postgr.es/m/...
Backpatch-through: 19
---
src/backend/utils/misc/guc_tables.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index c6d9b2a6f89..342aaeef59a 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -513,10 +513,10 @@ static const struct config_enum_entry file_extend_method_options[] = {
};
static const struct config_enum_entry data_checksums_options[] = {
- {"on", PG_DATA_CHECKSUM_VERSION, true},
- {"off", PG_DATA_CHECKSUM_OFF, true},
- {"inprogress-on", PG_DATA_CHECKSUM_INPROGRESS_ON, true},
- {"inprogress-off", PG_DATA_CHECKSUM_INPROGRESS_OFF, true},
+ {"on", PG_DATA_CHECKSUM_VERSION, false},
+ {"off", PG_DATA_CHECKSUM_OFF, false},
+ {"inprogress-on", PG_DATA_CHECKSUM_INPROGRESS_ON, false},
+ {"inprogress-off", PG_DATA_CHECKSUM_INPROGRESS_OFF, false},
{NULL, 0, false}
};
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v15-0005-doc-Documentation-updates-for-online-checksums.patch (22.1K, ../../4BCE2F21-B669-46F3-B781-010F71B0B065@yesql.se/6-v15-0005-doc-Documentation-updates-for-online-checksums.patch)
download | inline diff:
From 1c65e0f756b07ae033778358e98000c93521f7c5 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Wed, 9 Sep 2026 23:38:11 +0200
Subject: [PATCH v15 5/9] doc: Documentation updates for online checksums
A set of documentation updates for the online checksums work, found via
manual as well as LLM-guided review.
* The state diagram was re-done to be improve readability and be
more in line with the look and feel of other diagrams.
* Paragraph about mismatched states moved to its own sect2.
* The documentation for the functions to enable/ disable checksums
failed to mention that they are superuser only. The ACL in the
procedure definition was also documenting a less strict model
than the code so it was updated to match the docs and code.
* The documentation for the checksums progress reporting had
accidentally omitted the work "fork" and for th blocks_* columns.
The values are per fork but the documentation made it seem they
were per relation.
Reported-by: Heikki Linnakangas <hlinnaka@iki.fi>
Reported-by: Noah Misch <noah@leadboat.com>
Discussion: https://postgr.es/m/5b08b4d1-0982-4b77-bac5-3bdffc6583f5@iki.fi
---
doc/src/sgml/func/func-admin.sgml | 13 +-
doc/src/sgml/images/datachecksums.gv | 61 ++++++++--
doc/src/sgml/images/datachecksums.svg | 163 ++++++++++++++++----------
doc/src/sgml/images/meson.build | 1 +
doc/src/sgml/monitoring.sgml | 4 +-
doc/src/sgml/ref/pg_checksums.sgml | 2 +-
doc/src/sgml/wal.sgml | 60 ++++++----
src/include/catalog/pg_proc.dat | 2 +-
8 files changed, 204 insertions(+), 102 deletions(-)
diff --git a/doc/src/sgml/func/func-admin.sgml b/doc/src/sgml/func/func-admin.sgml
index 54eeb42e5bc..8bfd290ce15 100644
--- a/doc/src/sgml/func/func-admin.sgml
+++ b/doc/src/sgml/func/func-admin.sgml
@@ -3128,7 +3128,8 @@ SELECT convert_from(pg_read_binary_file('file_in_utf8.txt'), 'UTF8');
<para>
The functions shown in <xref linkend="functions-checksums-table" /> can
- be used to enable or disable data checksums in a running cluster.
+ be used to enable or disable data checksums in a running cluster. Use of
+ these functions is restricted to superusers.
</para>
<para>
Changing data checksums can be done in a cluster with concurrent activity
@@ -3158,7 +3159,7 @@ SELECT convert_from(pg_read_binary_file('file_in_utf8.txt'), 'UTF8');
<indexterm>
<primary>pg_enable_data_checksums</primary>
</indexterm>
- <function>pg_enable_data_checksums</function> ( <optional><parameter>cost_delay</parameter> <type>int</type>, <parameter>cost_limit</parameter> <type>int</type></optional> )
+ <function>pg_enable_data_checksums</function> ( <optional><parameter>cost_delay</parameter> <type>int</type> <optional>, <parameter>cost_limit</parameter> <type>int</type></optional></optional> )
<returnvalue>void</returnvalue>
</para>
<para>
@@ -3174,6 +3175,11 @@ SELECT convert_from(pg_read_binary_file('file_in_utf8.txt'), 'UTF8');
If <parameter>cost_delay</parameter> and <parameter>cost_limit</parameter> are
specified, the process is throttled using the same principles as
<link linkend="runtime-config-resource-vacuum-cost">Cost-based Vacuum Delay</link>.
+ <parameter>cost_delay</parameter> defaults to 0,
+ <parameter>cost_limit</parameter> defaults to 100.
+ </para>
+ <para>
+ This function is restricted to superusers.
</para>
</entry>
</row>
@@ -3193,6 +3199,9 @@ SELECT convert_from(pg_read_binary_file('file_in_utf8.txt'), 'UTF8');
stopped validating data checksums, the data checksum state will be
set to <literal>off</literal>.
</para>
+ <para>
+ This function is restricted to superusers.
+ </para>
</entry>
</row>
</tbody>
diff --git a/doc/src/sgml/images/datachecksums.gv b/doc/src/sgml/images/datachecksums.gv
index dff3ff7340a..f032555db3e 100644
--- a/doc/src/sgml/images/datachecksums.gv
+++ b/doc/src/sgml/images/datachecksums.gv
@@ -1,14 +1,49 @@
-digraph G {
- A -> B [label="SELECT pg_enable_data_checksums()"];
- B -> C;
- D -> A;
- C -> D [label="SELECT pg_disable_data_checksums()"];
- E -> A [label=" --no-data-checksums"];
- E -> C [label=" --data-checksums"];
-
- A [label="off"];
- B [label="inprogress-on"];
- C [label="on"];
- D [label="inprogress-off"];
- E [label="initdb"];
+digraph "datachecksums_states" {
+ layout=dot;
+ node [label="", shape=box, style=filled, fillcolor=gray, width=1.0, fontname="sans-serif"];
+
+ m1 [label="initdb", shape=Mdiamond];
+
+ subgraph cluster01 {
+ label="Online Checksums";
+ subgraph agroup1 {
+ rank=same;
+ a1;
+ a3;
+ }
+
+ subgraph agroup2 {
+ rank=same;
+ a2;
+ a4;
+ }
+
+ a1 -> a2;
+ a2 -> a3;
+ a3 -> a4;
+ a4 -> a1;
+
+ a1 [fillcolor=lightblue, label="off"];
+ a2 [fillcolor=lightgreen, label="inprogress-on"];
+ a3 [fillcolor=lightblue, label="on"];
+ a4 [fillcolor=lightgreen, label="inprogress-off"];
+ }
+
+ subgraph cluster05 {
+ label="Offline Checksums";
+ subgraph bgroup2 {
+ rank=same;
+ b1 -> b2;
+ }
+
+ b2 -> b1;
+
+ b1 [fillcolor=lightblue, label="off"];
+ b2 [fillcolor=lightblue, label="on"];
+ }
+
+ m1 -> a1;
+ m1 -> a3;
+ m1 -> b1;
+ m1 -> b2;
}
diff --git a/doc/src/sgml/images/datachecksums.svg b/doc/src/sgml/images/datachecksums.svg
index 8c58f42922e..eee57c393c0 100644
--- a/doc/src/sgml/images/datachecksums.svg
+++ b/doc/src/sgml/images/datachecksums.svg
@@ -1,81 +1,126 @@
-<?xml version="1.0" encoding="UTF-8" standalone="no"?>
+<?xml version="1.0"?>
<!-- Generated by graphviz version 14.0.5 (20251129.0259)
-->
-<!-- Title: G Pages: 1 -->
-<svg width="409pt" height="383pt"
- viewBox="0.00 0.00 409.00 383.00" xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink">
-<g id="graph0" class="graph" transform="scale(1 1) rotate(0) translate(4 378.5)">
-<title>G</title>
-<polygon fill="white" stroke="none" points="-4,4 -4,-378.5 404.74,-378.5 404.74,4 -4,4"/>
-<!-- A -->
+<!-- Title: datachecksums_states Pages: 1 -->
+<svg xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" width="441pt" height="209pt" viewBox="0.00 0.00 441.00 209.00">
+<g id="graph0" class="graph" transform="scale(1 1) rotate(0) translate(4 204.5)">
+<title>datachecksums_states</title>
+<polygon fill="white" stroke="none" points="-4,4 -4,-204.5 437,-204.5 437,4 -4,4"/>
+<g id="clust1" class="cluster">
+<title>cluster01</title>
+<polygon fill="none" stroke="black" points="8,-8 8,-156.5 239,-156.5 239,-8 8,-8"/>
+<text xml:space="preserve" text-anchor="middle" x="123.5" y="-139.2" font-family="Times,serif" font-size="14.00">Online Checksums</text>
+</g>
+<g id="clust4" class="cluster">
+<title>cluster05</title>
+<polygon fill="none" stroke="black" points="247,-80 247,-156.5 425,-156.5 425,-80 247,-80"/>
+<text xml:space="preserve" text-anchor="middle" x="336" y="-139.2" font-family="Times,serif" font-size="14.00">Offline Checksums</text>
+</g>
+<!-- m1 -->
<g id="node1" class="node">
-<title>A</title>
-<ellipse fill="none" stroke="black" cx="80.12" cy="-268" rx="27" ry="18"/>
-<text xml:space="preserve" text-anchor="middle" x="80.12" y="-262.95" font-family="Times,serif" font-size="14.00">off</text>
+<title>m1</title>
+<polygon fill="gray" stroke="black" points="242,-200.5 197.65,-182.5 242,-164.5 286.35,-182.5 242,-200.5"/>
+<polyline fill="none" stroke="black" points="208.77,-187.01 208.77,-177.99"/>
+<polyline fill="none" stroke="black" points="230.88,-169.01 253.12,-169.01"/>
+<polyline fill="none" stroke="black" points="275.23,-177.99 275.23,-187.01"/>
+<polyline fill="none" stroke="black" points="253.12,-195.99 230.88,-195.99"/>
+<text xml:space="preserve" text-anchor="middle" x="242" y="-176.7" font-family="sans-serif" font-size="14.00">initdb</text>
</g>
-<!-- B -->
+<!-- a1 -->
<g id="node2" class="node">
-<title>B</title>
-<ellipse fill="none" stroke="black" cx="137.12" cy="-179.5" rx="61.59" ry="18"/>
-<text xml:space="preserve" text-anchor="middle" x="137.12" y="-174.45" font-family="Times,serif" font-size="14.00">inprogress-on</text>
+<title>a1</title>
+<polygon fill="lightblue" stroke="black" points="134,-124 62,-124 62,-88 134,-88 134,-124"/>
+<text xml:space="preserve" text-anchor="middle" x="98" y="-100.2" font-family="sans-serif" font-size="14.00">off</text>
</g>
-<!-- A->B -->
-<g id="edge1" class="edge">
-<title>A->B</title>
-<path fill="none" stroke="black" d="M76.5,-249.68C75.22,-239.14 75.3,-225.77 81.12,-215.5 84.2,-210.08 88.49,-205.38 93.35,-201.34"/>
-<polygon fill="black" stroke="black" points="95.22,-204.31 101.33,-195.66 91.16,-198.61 95.22,-204.31"/>
-<text xml:space="preserve" text-anchor="middle" x="187.62" y="-218.7" font-family="Times,serif" font-size="14.00">SELECT pg_enable_data_checksums()</text>
+<!-- m1->a1 -->
+<g id="edge7" class="edge">
+<title>m1->a1</title>
+<path fill="none" stroke="black" d="M208.02,-177.91C187.98,-174.61 162.76,-168.35 143,-156.5 132.95,-150.47 123.82,-141.5 116.46,-132.85"/>
+<polygon fill="black" stroke="black" points="119.38,-130.89 110.41,-125.26 113.91,-135.26 119.38,-130.89"/>
</g>
-<!-- C -->
+<!-- a3 -->
<g id="node3" class="node">
-<title>C</title>
-<ellipse fill="none" stroke="black" cx="137.12" cy="-106.5" rx="27" ry="18"/>
-<text xml:space="preserve" text-anchor="middle" x="137.12" y="-101.45" font-family="Times,serif" font-size="14.00">on</text>
+<title>a3</title>
+<polygon fill="lightblue" stroke="black" points="224,-124 152,-124 152,-88 224,-88 224,-124"/>
+<text xml:space="preserve" text-anchor="middle" x="188" y="-100.2" font-family="sans-serif" font-size="14.00">on</text>
</g>
-<!-- B->C -->
-<g id="edge2" class="edge">
-<title>B->C</title>
-<path fill="none" stroke="black" d="M137.12,-161.31C137.12,-153.73 137.12,-144.6 137.12,-136.04"/>
-<polygon fill="black" stroke="black" points="140.62,-136.04 137.12,-126.04 133.62,-136.04 140.62,-136.04"/>
+<!-- m1->a3 -->
+<g id="edge8" class="edge">
+<title>m1->a3</title>
+<path fill="none" stroke="black" d="M232.35,-168.18C225.36,-158.54 215.68,-145.18 207.14,-133.41"/>
+<polygon fill="black" stroke="black" points="210.21,-131.68 201.51,-125.64 204.54,-135.79 210.21,-131.68"/>
+</g>
+<!-- b1 -->
+<g id="node6" class="node">
+<title>b1</title>
+<polygon fill="lightblue" stroke="black" points="327,-124 255,-124 255,-88 327,-88 327,-124"/>
+<text xml:space="preserve" text-anchor="middle" x="291" y="-100.2" font-family="sans-serif" font-size="14.00">off</text>
+</g>
+<!-- m1->b1 -->
+<g id="edge9" class="edge">
+<title>m1->b1</title>
+<path fill="none" stroke="black" d="M250.99,-167.84C257.29,-158.25 265.92,-145.13 273.55,-133.53"/>
+<polygon fill="black" stroke="black" points="276.25,-135.79 278.82,-125.51 270.4,-131.95 276.25,-135.79"/>
+</g>
+<!-- b2 -->
+<g id="node7" class="node">
+<title>b2</title>
+<polygon fill="lightblue" stroke="black" points="417,-124 345,-124 345,-88 417,-88 417,-124"/>
+<text xml:space="preserve" text-anchor="middle" x="381" y="-100.2" font-family="sans-serif" font-size="14.00">on</text>
+</g>
+<!-- m1->b2 -->
+<g id="edge10" class="edge">
+<title>m1->b2</title>
+<path fill="none" stroke="black" d="M275.15,-177.42C294.03,-173.98 317.53,-167.73 336,-156.5 346.02,-150.41 355.13,-141.42 362.49,-132.78"/>
+<polygon fill="black" stroke="black" points="365.04,-135.2 368.56,-125.2 359.57,-130.82 365.04,-135.2"/>
</g>
-<!-- D -->
+<!-- a2 -->
<g id="node4" class="node">
-<title>D</title>
-<ellipse fill="none" stroke="black" cx="63.12" cy="-18" rx="63.12" ry="18"/>
-<text xml:space="preserve" text-anchor="middle" x="63.12" y="-12.95" font-family="Times,serif" font-size="14.00">inprogress-off</text>
+<title>a2</title>
+<polygon fill="lightgreen" stroke="black" points="114.25,-52 15.75,-52 15.75,-16 114.25,-16 114.25,-52"/>
+<text xml:space="preserve" text-anchor="middle" x="65" y="-28.2" font-family="sans-serif" font-size="14.00">inprogress-on</text>
</g>
-<!-- C->D -->
-<g id="edge4" class="edge">
-<title>C->D</title>
-<path fill="none" stroke="black" d="M124.23,-90.43C113.36,-77.73 97.58,-59.28 84.77,-44.31"/>
-<polygon fill="black" stroke="black" points="87.78,-42.44 78.62,-37.12 82.46,-46.99 87.78,-42.44"/>
-<text xml:space="preserve" text-anchor="middle" x="214.75" y="-57.2" font-family="Times,serif" font-size="14.00">SELECT pg_disable_data_checksums()</text>
+<!-- a1->a2 -->
+<g id="edge1" class="edge">
+<title>a1->a2</title>
+<path fill="none" stroke="black" d="M89.84,-87.7C86.25,-80.07 81.93,-70.92 77.92,-62.4"/>
+<polygon fill="black" stroke="black" points="81.14,-61.03 73.71,-53.47 74.81,-64.01 81.14,-61.03"/>
+</g>
+<!-- a4 -->
+<g id="node5" class="node">
+<title>a4</title>
+<polygon fill="lightgreen" stroke="black" points="231.25,-52 132.75,-52 132.75,-16 231.25,-16 231.25,-52"/>
+<text xml:space="preserve" text-anchor="middle" x="182" y="-28.2" font-family="sans-serif" font-size="14.00">inprogress-off</text>
</g>
-<!-- D->A -->
+<!-- a3->a4 -->
<g id="edge3" class="edge">
-<title>D->A</title>
-<path fill="none" stroke="black" d="M62.52,-36.28C61.62,-68.21 60.54,-138.57 66.12,-197.5 67.43,-211.24 70.27,-226.28 73.06,-238.85"/>
-<polygon fill="black" stroke="black" points="69.64,-239.59 75.32,-248.54 76.46,-238 69.64,-239.59"/>
+<title>a3->a4</title>
+<path fill="none" stroke="black" d="M186.52,-87.7C185.89,-80.41 185.15,-71.73 184.45,-63.54"/>
+<polygon fill="black" stroke="black" points="187.94,-63.28 183.6,-53.61 180.96,-63.87 187.94,-63.28"/>
</g>
-<!-- E -->
-<g id="node5" class="node">
-<title>E</title>
-<ellipse fill="none" stroke="black" cx="198.12" cy="-356.5" rx="32.41" ry="18"/>
-<text xml:space="preserve" text-anchor="middle" x="198.12" y="-351.45" font-family="Times,serif" font-size="14.00">initdb</text>
+<!-- a2->a3 -->
+<g id="edge2" class="edge">
+<title>a2->a3</title>
+<path fill="none" stroke="black" d="M95.33,-52.26C111.09,-61.23 130.56,-72.31 147.59,-82"/>
+<polygon fill="black" stroke="black" points="145.54,-84.86 155.96,-86.77 149,-78.78 145.54,-84.86"/>
+</g>
+<!-- a4->a1 -->
+<g id="edge4" class="edge">
+<title>a4->a1</title>
+<path fill="none" stroke="black" d="M161.18,-52.35C150.98,-60.85 138.51,-71.24 127.36,-80.53"/>
+<polygon fill="black" stroke="black" points="125.37,-77.64 119.93,-86.73 129.85,-83.01 125.37,-77.64"/>
</g>
-<!-- E->A -->
+<!-- b1->b2 -->
<g id="edge5" class="edge">
-<title>E->A</title>
-<path fill="none" stroke="black" d="M179.16,-341.6C159.64,-327.29 129.05,-304.86 107.03,-288.72"/>
-<polygon fill="black" stroke="black" points="109.23,-286 99.1,-282.91 105.09,-291.64 109.23,-286"/>
-<text xml:space="preserve" text-anchor="middle" x="208.57" y="-307.2" font-family="Times,serif" font-size="14.00"> --no-data-checksums</text>
+<title>b1->b2</title>
+<path fill="none" stroke="black" d="M327.21,-93.01C329.33,-92.77 331.44,-92.61 333.55,-92.54"/>
+<polygon fill="black" stroke="black" points="333.21,-96.03 343.35,-92.96 333.51,-89.03 333.21,-96.03"/>
</g>
-<!-- E->C -->
+<!-- b2->b1 -->
<g id="edge6" class="edge">
-<title>E->C</title>
-<path fill="none" stroke="black" d="M227.13,-348.04C242.29,-342.72 259.95,-334.06 271.12,-320.5 301.5,-283.62 316.36,-257.78 294.12,-215.5 268.41,-166.6 209.42,-135.53 171.52,-119.85"/>
-<polygon fill="black" stroke="black" points="172.96,-116.65 162.37,-116.21 170.37,-123.16 172.96,-116.65"/>
-<text xml:space="preserve" text-anchor="middle" x="350.87" y="-218.7" font-family="Times,serif" font-size="14.00"> --data-checksums</text>
+<title>b2->b1</title>
+<path fill="none" stroke="black" d="M344.86,-118.98C342.75,-119.23 340.63,-119.39 338.52,-119.46"/>
+<polygon fill="black" stroke="black" points="338.86,-115.97 328.72,-119.05 338.57,-122.96 338.86,-115.97"/>
</g>
</g>
</svg>
diff --git a/doc/src/sgml/images/meson.build b/doc/src/sgml/images/meson.build
index 220e3eaafb8..54e80fb402c 100644
--- a/doc/src/sgml/images/meson.build
+++ b/doc/src/sgml/images/meson.build
@@ -11,6 +11,7 @@ image_targets = []
fixup_svg_xsl = files('fixup-svg.xsl')
all_files = [
+ 'datachecksums.gv',
'genetic-algorithm.gv',
'gin.gv',
'pagelayout.txt',
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index 6a4cb9bb144..8614338fcab 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -8339,7 +8339,7 @@ FROM pg_stat_get_backend_idset() AS backendid;
<structfield>blocks_total</structfield> <type>bigint</type>
</para>
<para>
- The number of blocks in the current relation which will be processed,
+ The number of blocks in the current relation fork which will be processed,
or <literal>NULL</literal> if the worker process hasn't
calculated the number of blocks yet. The launcher process has
this set to <literal>NULL</literal>.
@@ -8353,7 +8353,7 @@ FROM pg_stat_get_backend_idset() AS backendid;
<structfield>blocks_done</structfield> <type>bigint</type>
</para>
<para>
- The number of blocks in the current relation which have been processed.
+ The number of blocks in the current relation fork which have been processed.
The launcher process has this set to <literal>NULL</literal>.
</para>
</entry>
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index abc035e5409..5a0bda2eca2 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -288,7 +288,7 @@ PostgreSQL documentation
</procedure>
<para>
- The replay requirement in <xref linkend="shut-down-nodes"/>exists
+ The replay requirement in <xref linkend="shut-down-nodes"/> exists
because an offline change is recorded only in the control file and
has no defined ordering against WAL the node has not replayed yet; see
<xref linkend="checksums-offline-enable-disable"/>. A node stopped
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index dbbfdae5e4f..edd39ee437f 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -316,30 +316,6 @@
application can be used to enable or disable data checksums, as well as
verify checksums, on an offline cluster.
</para>
-
- <para>
- An offline change provides durability differently from an
- <link linkend="checksums-online-enable-disable">online change</link>.
- An online transition is WAL-logged: it is ordered against all other
- WAL records, it is replayed after a crash, and it propagates to
- standbys. An offline change is recorded only in the cluster's
- control file: it writes no WAL, it is invisible to replication, and
- it has no defined ordering against WAL the node has not replayed
- yet. When a node later replays WAL that contains an online checksum
- state change, that change takes effect on the node even if it was
- written before the offline change was made.
- </para>
-
- <para>
- An offline change only affects the data directory it is run on; the
- new state does not propagate over replication. In a replication setup
- the same change must be applied to all nodes while all of them are
- stopped, as described in <xref linkend="app-pgchecksums"/>. A standby
- whose state diverges logs a warning but keeps its local setting. The
- mismatch persists until the states are brought together again, with
- the offline procedure or with an online transition; do this promptly.
- </para>
-
</sect2>
<sect2 id="checksums-online-enable-disable" xreflabel="Online Enabling of Checksums">
@@ -456,7 +432,43 @@
still required.
</para>
</sect3>
+ </sect2>
+
+ <sect2 id="checksums-mismatched-states">
+ <title>Mismatched Data Checksums States in a Replicated Cluster</title>
+ <para>
+ The primary and secondaries can end up with different data checksums
+ states due to the different durability models between offline and online
+ checksum operations.
+ </para>
+ <para>
+ An online transition is WAL-logged, which means that it is replayed after a
+ crash, and the transition along with all checksum updates are propagated to
+ standbys. An offline change is recorded only in the cluster's control
+ file, it writes no WAL and is invisible to replication. The changes made
+ to the datafiles to write checksums are not WAL logged. This means that it
+ has no defined ordering against WAL the node has not replayed yet. When a
+ node later replays WAL that contains an online checksum state change, that
+ change takes effect on the node even if it was written before the offline
+ change was made.
+ </para>
+ <para>
+ An offline change only affects the data directory it is run on; the new
+ state does not propagate over replication. This means that all nodes can
+ be operated on in parallel, and no additional network traffic or WAL
+ archive traffic will occur. In a replication setup the same change must be
+ applied to all nodes while all of them are stopped, as described in
+ <xref linkend="app-pgchecksums"/>. A standby whose state diverges logs a
+ warning but keeps its local setting. The mismatch persists until the
+ states are brought together again, with the offline procedure or with an
+ online transition.
+ </para>
+
+ <para>
+ Running a replicated cluster with different data checksum states on the
+ different nodes is not supported or recommended.
+ </para>
</sect2>
</sect1>
diff --git a/src/include/catalog/pg_proc.dat b/src/include/catalog/pg_proc.dat
index c53ce68c717..674159b3a22 100644
--- a/src/include/catalog/pg_proc.dat
+++ b/src/include/catalog/pg_proc.dat
@@ -12477,7 +12477,7 @@
prorettype => 'void', proargtypes => 'int4 int4',
proallargtypes => '{int4,int4}', proargmodes => '{i,i}',
proargnames => '{cost_delay,cost_limit}', proargdefaults => '{0,100}',
- prosrc => 'enable_data_checksums', proacl => '{POSTGRES=X}' },
+ prosrc => 'enable_data_checksums' },
# collation management functions
{ oid => '3445', descr => 'import collations from operating system',
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v15-0004-pg_combinebackup-Refuse-mixed-data-checksum-stat.patch (10.0K, ../../4BCE2F21-B669-46F3-B781-010F71B0B065@yesql.se/7-v15-0004-pg_combinebackup-Refuse-mixed-data-checksum-stat.patch)
download | inline diff:
From cb0fa8e654947cf711d6c33c0e0507146fa1e8d6 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 3 Sep 2026 23:32:04 +0200
Subject: [PATCH v15 4/9] pg_combinebackup: Refuse mixed data checksum states
in a backup chain
check_control_files() detected a chain whose backups were taken under
different data checksum states, warned, and proceeded. The output
directory keeps the last backup's control file, so a full backup taken
with checksums off combined with an incremental taken after an offline
enable produces a cluster whose control file says "on" while most of
its blocks carry no checksums; it starts, and then every connection
dies on the first unchecksummed catalog page. An offline enable
between two backups of a chain is all it takes, since it rewrites
every page without logging anything, so the incremental backup does
not re-ship the pages.
Turn the warning into an error, matching what pg_rewind does for the
equivalent combinations. The check stays asymmetric on purpose: when
the last backup was taken with checksums off, stale checksums from an
earlier backup are never verified, and that chain remains usable.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Reviewed-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Discussion: https://postgr.es/m/anwm6UPxoVS41QA2@bdtpg
---
doc/src/sgml/ref/pg_combinebackup.sgml | 19 +--
src/bin/pg_combinebackup/pg_combinebackup.c | 14 +-
src/test/modules/test_checksums/meson.build | 1 +
.../t/024_combinebackup_mixed.pl | 146 ++++++++++++++++++
4 files changed, 167 insertions(+), 13 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
diff --git a/doc/src/sgml/ref/pg_combinebackup.sgml b/doc/src/sgml/ref/pg_combinebackup.sgml
index 9a6d201e0b8..f5d7e177b36 100644
--- a/doc/src/sgml/ref/pg_combinebackup.sgml
+++ b/doc/src/sgml/ref/pg_combinebackup.sgml
@@ -306,18 +306,19 @@ PostgreSQL documentation
<para>
<literal>pg_combinebackup</literal> does not recompute page checksums when
- writing the output directory. Therefore, if any of the backups used for
- reconstruction were taken with checksums disabled, but the final backup was
- taken with checksums enabled, the resulting directory may contain pages
- with invalid checksums.
+ writing the output directory. It therefore refuses a chain in which some
+ of the backups used for reconstruction were taken with checksums disabled
+ but the final backup was taken with checksums enabled: the resulting
+ directory would contain pages with invalid checksums, and the cluster
+ would fail checksum verification as soon as it reads them. Take a new
+ full backup after enabling data checksums with
+ <xref linkend="app-pgchecksums"/>.
</para>
<para>
- To avoid this problem, taking a new full backup after changing the checksum
- state of the cluster using <xref linkend="app-pgchecksums"/> is
- recommended. Otherwise, you can disable and then optionally reenable
- checksums on the directory produced by <literal>pg_combinebackup</literal>
- in order to correct the problem.
+ The reverse case is accepted: when the final backup was taken with
+ checksums disabled, stale checksums copied from an earlier backup are
+ never verified.
</para>
</refsect1>
diff --git a/src/bin/pg_combinebackup/pg_combinebackup.c b/src/bin/pg_combinebackup/pg_combinebackup.c
index 86e2ee37c40..254a27b125b 100644
--- a/src/bin/pg_combinebackup/pg_combinebackup.c
+++ b/src/bin/pg_combinebackup/pg_combinebackup.c
@@ -670,13 +670,19 @@ check_control_files(int n_backups, char **backup_dirs)
pg_log_debug("system identifier is %" PRIu64, system_identifier);
/*
- * Warn the user if not all backups are in the same state with regards to
- * checksums.
+ * Reject a chain whose backups were not all in the same state with
+ * regards to checksums. The loop above only flags that when the last
+ * backup has checksums enabled: the blocks taken from the older backups
+ * have no checksums, and pg_combinebackup does not recompute them, so the
+ * combined cluster would fail verification as soon as it reads them. The
+ * other direction is harmless, since stale checksums are never verified
+ * while checksums are disabled.
*/
if (data_checksum_mismatch)
{
- pg_log_warning("only some backups have checksums enabled");
- pg_log_warning_hint("Disable, and optionally reenable, checksums on the output directory to avoid failures.");
+ pg_log_error("only some backups have checksums enabled");
+ pg_log_error_hint("Take a new full backup after changing the data checksum state with pg_checksums.");
+ exit(1);
}
return system_identifier;
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index 7d07c757052..7eccd5156b9 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -47,6 +47,7 @@ tests += {
't/021_rewind_divergent_transitions.pl',
't/022_rewind_state.pl',
't/023_rewind_standby_target.pl',
+ 't/024_combinebackup_mixed.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
new file mode 100644
index 00000000000..2f51c7fd9dc
--- /dev/null
+++ b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
@@ -0,0 +1,146 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_combinebackup refuses a chain whose backups were taken under
+# different data checksum states when the final state has them enabled.
+#
+# The output directory keeps the last backup's control file, so a full
+# backup taken with checksums off plus an incremental taken after an
+# offline enable would produce a directory whose control file says "on"
+# while most of its blocks have no checksums. An offline enable between
+# the two backups is enough to get there: it rewrites every page but logs
+# nothing, so the incremental backup does not re-ship the pages.
+#
+# The reverse order stays allowed: with the final backup taken with
+# checksums off, stale checksums from an earlier backup are never verified.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('node');
+$node->init(
+ no_data_checksums => 1,
+ has_archiving => 1,
+ allows_streaming => 1);
+$node->append_conf('postgresql.conf', 'summarize_wal = on');
+$node->append_conf('postgresql.conf', 'autovacuum = off');
+$node->start;
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+test_checksum_state($node, 'off');
+
+$node->backup('full');
+
+# Offline enable: rewrites every page, logs nothing.
+$node->stop;
+system_or_bail('pg_checksums', '--enable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'on');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full/backup_manifest',
+ ],
+ 'incremental backup with checksums on');
+
+# Combining the chain must be refused: the result would say "on" while the
+# blocks inherited from the full backup have no checksums.
+command_fails_like(
+ [
+ 'pg_combinebackup',
+ $node->backup_dir . '/full',
+ $node->backup_dir . '/incr',
+ '--output' => $node->backup_dir . '/combined',
+ ],
+ qr/only some backups have checksums enabled/,
+ 'pg_combinebackup refuses a chain crossing an offline enable');
+
+# The other direction: full backup with checksums on, offline disable, then
+# an incremental. The last backup wins, so the combined cluster comes up off.
+$node->backup('full2');
+
+$node->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr2',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full2/backup_manifest',
+ ],
+ 'incremental backup with checksums off');
+
+my $restored = PostgreSQL::Test::Cluster->new('restored');
+$restored->init_from_backup(
+ $node, 'incr2',
+ combine_with_prior => ['full2'],
+ has_restoring => 1,
+ standby => 0);
+
+$restored->start;
+
+my ($rc, $stdout, $stderr) = $restored->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'combined backup accepts connections') or diag("stderr: $stderr");
+is($stdout, 'off', 'combined backup follows the last backup state');
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t3;');
+is($rc, 0, 'blocks from the incremental backup are readable')
+ or diag("stderr: $stderr");
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t1;');
+is($rc, 0, 'blocks inherited from the full backup are readable')
+ or diag("stderr: $stderr");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($restored->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures in the combined backup');
+
+$restored->stop('immediate');
+$node->stop;
+
+done_testing();
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v15-0003-pg_rewind-Check-the-data-checksum-states-of-sour.patch (22.5K, ../../4BCE2F21-B669-46F3-B781-010F71B0B065@yesql.se/8-v15-0003-pg_rewind-Check-the-data-checksum-states-of-sour.patch)
download | inline diff:
From 16d9b4ab0bd94df601cf2cbe2644536f3c2394f5 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 3 Sep 2026 23:31:52 +0200
Subject: [PATCH v15 3/9] pg_rewind: Check the data checksum states of source
and target
pg_rewind installs the source's control file on the target, but every
block it does not copy keeps the target's content. With checksums
enabled on the target and disabled on the source, the rewound server
claims enabled checksums while the blocks copied from the source have
none, and fails checksum verification as soon as it reads them; in the
worst case every connection attempt dies on an unverifiable catalog
page.
Comparing the control files is not enough. Replay on the rewound
server resumes from the last common checkpoint and adopts the data
checksum state recorded there, so an offline disable on the source
after the divergence leaves both control files saying "off" while the
rewound server still resumes with verification enabled. Check the
state carried by the divergence checkpoint as well, which pg_rewind
already reads.
Refuse both cases, and an interrupted online transition on either
side, same as pg_checksums does. The opposite mismatch, checksums on
the source only, is allowed with a warning: recovery under the backup
label adopts the divergence state, so the rewound server keeps
checksums disabled (or converges through the replayed WAL if the
source enabled them online), and no page can fail verification.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Reviewed-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Discussion: https://postgr.es/m/anwm6UPxoVS41QA2@bdtpg
---
doc/src/sgml/ref/pg_rewind.sgml | 14 ++
src/bin/pg_rewind/parsexlog.c | 4 +-
src/bin/pg_rewind/pg_rewind.c | 93 ++++++++++-
src/bin/pg_rewind/pg_rewind.h | 1 +
src/test/modules/test_checksums/meson.build | 2 +
.../test_checksums/t/022_rewind_state.pl | 145 +++++++++++++++++
.../t/023_rewind_standby_target.pl | 153 ++++++++++++++++++
7 files changed, 410 insertions(+), 2 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/022_rewind_state.pl
create mode 100644 src/test/modules/test_checksums/t/023_rewind_standby_target.pl
diff --git a/doc/src/sgml/ref/pg_rewind.sgml b/doc/src/sgml/ref/pg_rewind.sgml
index b95e3868e9d..ac3d0c9328f 100644
--- a/doc/src/sgml/ref/pg_rewind.sgml
+++ b/doc/src/sgml/ref/pg_rewind.sgml
@@ -124,6 +124,20 @@ PostgreSQL documentation
<literal>on</literal>, but is enabled by default.
</para>
+ <para>
+ The data checksum states of the source and target server must be
+ compatible. <application>pg_rewind</application> refuses to run if an
+ online checksum state transition was interrupted on either server, or
+ if data checksums are enabled on the target server, or were enabled at
+ the point of divergence, while the source server runs without them: the
+ rewound server would fail checksum verification on the blocks copied
+ from the source. The opposite combination is allowed with a warning;
+ the rewound server keeps data checksums disabled unless the source
+ server enabled them online after the point of divergence. See
+ <xref linkend="app-pgchecksums"/> for how to change the state
+ consistently across a replication setup.
+ </para>
+
<warning>
<title>Warning: Failures While Rewinding</title>
<para>
diff --git a/src/bin/pg_rewind/parsexlog.c b/src/bin/pg_rewind/parsexlog.c
index 023e23b063c..6e87b00f8c2 100644
--- a/src/bin/pg_rewind/parsexlog.c
+++ b/src/bin/pg_rewind/parsexlog.c
@@ -167,7 +167,8 @@ readOneRecord(const char *datadir, XLogRecPtr ptr, int tliIndex,
void
findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
- XLogRecPtr *lastchkptredo, const char *restoreCommand)
+ XLogRecPtr *lastchkptredo, uint32 *lastchkptdatachecksums,
+ const char *restoreCommand)
{
/* Walk backwards, starting from the given record */
XLogRecord *record;
@@ -255,6 +256,7 @@ findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
*lastchkptrec = searchptr;
*lastchkpttli = checkPoint.ThisTimeLineID;
*lastchkptredo = checkPoint.redo;
+ *lastchkptdatachecksums = checkPoint.dataChecksumState;
break;
}
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 0e4c2df3f4b..d2521dab333 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -145,6 +145,7 @@ main(int argc, char **argv)
XLogRecPtr chkptrec;
TimeLineID chkpttli;
XLogRecPtr chkptredo;
+ uint32 chkptdatachecksums;
TimeLineID source_tli;
TimeLineID target_tli;
XLogRecPtr target_wal_endrec;
@@ -472,10 +473,51 @@ main(int argc, char **argv)
keepwal_init();
findLastCheckpoint(datadir_target, divergerec, lastcommontliIndex,
- &chkptrec, &chkpttli, &chkptredo, restore_command);
+ &chkptrec, &chkpttli, &chkptredo, &chkptdatachecksums,
+ restore_command);
pg_log_info("rewinding from last common checkpoint at %X/%08X on timeline %u",
LSN_FORMAT_ARGS(chkptrec), chkpttli);
+ /*
+ * Replay on the rewound server resumes from the last common checkpoint
+ * and adopts the data checksum state recorded there, not the state in the
+ * control file installed from the source. If checksums were enabled at
+ * the divergence point but the source runs without them, the rewound
+ * server would verify checksums while the blocks copied from the source
+ * have none. The control files cannot reveal this: an offline disable on
+ * the source after the divergence leaves both of them saying "off".
+ *
+ * Only the fully enabled state needs checking. An in-progress state at
+ * the divergence point means a transition was still running there, and
+ * all of its page rewrites are logged after that point, so replay brings
+ * the rewound server to whatever state the copied WAL ends in.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the source server or recreate the target server from a base backup.");
+ exit(1);
+ }
+
+ /*
+ * The same comparison is needed against the target: every block the
+ * rewind does not copy keeps the target's content. The control files
+ * cannot reveal this case either, in the other direction: a standby's
+ * checkpoints are written by its upstream primary, so a standby whose
+ * checksums were disabled offline still has "on" checkpoints in its WAL,
+ * while both control files may agree.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_target.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the target server");
+ pg_log_error_detail("Blocks kept from the target would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the target server with pg_checksums or recreate the target server from a base backup.");
+ exit(1);
+ }
+
/* Initialize the hash table to track the status of each file */
filehash_init();
@@ -799,6 +841,55 @@ sanityChecks(void)
pg_fatal("target server needs to use either data checksums or \"wal_log_hints = on\"");
}
+ /*
+ * The rewound target keeps its control file fields from the source, but
+ * every block the rewind does not copy keeps the target's content. If
+ * checksums are enabled on the target and disabled on the source, the
+ * result would claim enabled checksums while the blocks copied from the
+ * source have none, and the target would fail checksum verification as
+ * soon as it reads them. Refuse that combination, and refuse an
+ * interrupted online transition on either side, same as pg_checksums.
+ *
+ * The opposite mismatch is allowed: recovery under the backup label
+ * written by pg_rewind adopts the checksum state as of the divergence
+ * point, so a target whose own checkpoints carry no checksums stays
+ * without them, or converges through the replayed WAL if the source
+ * enabled them online. Only warn about it, so that an offline change on
+ * the source is not overlooked.
+ *
+ * That reasoning does not hold when the target is a standby, whose
+ * checkpoints were written by its upstream primary; see the checks after
+ * findLastCheckpoint().
+ */
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted on the target server");
+ pg_log_error_hint("Start the server and shut it down cleanly to reset the state, then retry; on a standby, let replication complete the transition first.");
+ exit(1);
+ }
+ if (ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition is incomplete on the source server");
+ pg_log_error_hint("Let the transition complete, or reset the state with a clean restart, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_OFF)
+ {
+ pg_log_error("data checksums are enabled on the target server but disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, and the target would fail checksum verification after the rewind.");
+ pg_log_error_hint("Either disable data checksums on the target server or enable them on the source server, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_OFF &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_warning("data checksums are disabled on the target server but enabled on the source server");
+ pg_log_warning_detail("The rewound server will keep data checksums disabled unless the source server enabled them online after the point of divergence.");
+ }
+
/*
* Target cluster better not be running. This doesn't guard against
* someone starting the cluster concurrently. Also, this is probably more
diff --git a/src/bin/pg_rewind/pg_rewind.h b/src/bin/pg_rewind/pg_rewind.h
index 9a981f7f246..9187c3bf72f 100644
--- a/src/bin/pg_rewind/pg_rewind.h
+++ b/src/bin/pg_rewind/pg_rewind.h
@@ -39,6 +39,7 @@ extern void findLastCheckpoint(const char *datadir, XLogRecPtr forkptr,
int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
XLogRecPtr *lastchkptredo,
+ uint32 *lastchkptdatachecksums,
const char *restoreCommand);
extern XLogRecPtr readOneRecord(const char *datadir, XLogRecPtr ptr,
int tliIndex, const char *restoreCommand);
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index e5d38fafb7d..7d07c757052 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -45,6 +45,8 @@ tests += {
't/019_standby_shutdown_catchup.pl',
't/020_cascade_divergence.pl',
't/021_rewind_divergent_transitions.pl',
+ 't/022_rewind_state.pl',
+ 't/023_rewind_standby_target.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/022_rewind_state.pl b/src/test/modules/test_checksums/t/022_rewind_state.pl
new file mode 100644
index 00000000000..e6d9c645a0b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/022_rewind_state.pl
@@ -0,0 +1,145 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind checks the data checksum states of source and target.
+# Checksums enabled on the target, or at the point of divergence, with
+# a source running without them are refused: the rewound server would
+# verify checksums on blocks copied from a source that has none. The
+# opposite mismatch only warns; replay keeps the target's own state.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+# Scenarios 1 and 2: checksums on from initdb. wal_log_hints keeps the
+# target eligible for pg_rewind after checksums are disabled on it.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b);
+
+# Failover to B, and divergence on A.
+$node_b->promote;
+$node_b->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_a->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_a->stop('fast');
+
+# Offline disable on the source only: enabled target, disabled source.
+$node_b->stop;
+$node_b->checksum_disable_offline;
+$node_b->start;
+test_checksum_state($node_b, 'off');
+
+my @rewind_cmd = (
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'));
+
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums are enabled on the target server but disabled on the source server/,
+ 'refuses an enabled target with a disabled source');
+
+# The refusal happens before any modification: the target still starts
+# on its own timeline.
+$node_a->start;
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'target untouched by the refused rewind');
+$node_a->stop('fast');
+
+# Scenario 2: disabling the target too makes the control files match,
+# but checksums were still enabled at the point of divergence, which is
+# the state replay on the rewound server would resume with.
+$node_a->checksum_disable_offline;
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums were enabled at the point of divergence but are disabled on the source server/,
+ 'refuses when checksums were enabled at the point of divergence');
+
+# Scenario 3: fresh pair without checksums; offline enable on the
+# source only. Allowed with a warning, and the rewound server keeps
+# checksums disabled.
+my $node_c = PostgreSQL::Test::Cluster->new('node_c');
+$node_c->init(allows_streaming => 1, no_data_checksums => 1);
+$node_c->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_c->start;
+$node_c->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_c->backup('backup');
+my $node_d = PostgreSQL::Test::Cluster->new('node_d');
+$node_d->init_from_backup($node_c, 'backup', has_streaming => 1);
+$node_d->start;
+$node_c->wait_for_catchup($node_d);
+
+$node_d->promote;
+$node_d->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_c->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_c->stop('fast');
+
+$node_d->stop;
+$node_d->checksum_enable_offline;
+$node_d->start;
+test_checksum_state($node_d, 'on');
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_c->data_dir,
+ '--source-server' => $node_d->connstr('postgres'),
+ ]);
+like(
+ $stderr,
+ qr/data checksums are disabled on the target server but enabled on the source server/,
+ 'warns for a disabled target with an enabled source');
+like($stderr, qr/Done!/, 'rewind completed despite the warning');
+
+# The rewound server follows D and keeps its own state.
+$node_c->append_conf('postgresql.conf', 'port = ' . $node_c->port);
+$node_c->enable_streaming($node_d);
+$node_c->set_standby_mode;
+$node_c->start;
+$node_d->wait_for_catchup($node_c);
+is($node_c->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+test_checksum_state($node_c, 'off');
+
+$node_c->stop;
+$node_d->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/023_rewind_standby_target.pl b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
new file mode 100644
index 00000000000..9071114a9ae
--- /dev/null
+++ b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
@@ -0,0 +1,153 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_rewind compares the data checksum state at the point of
+# divergence against the target as well as the source.
+#
+# A standby's checkpoints are written by its upstream primary, so a standby
+# whose checksums were disabled offline still has "on" checkpoints in its
+# WAL. Rewinding it onto a promoted sibling passes every control file
+# check (target off + source on only warns), yet recovery after the rewind
+# adopts the divergence checkpoint's state and the rewound server would
+# verify checksums over the pages it wrote while it was locally "off".
+# pg_rewind must refuse; after an offline enable on the target, which
+# rewrites all of its pages, the rewind goes through.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum_extended' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum_extended\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t0 AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+
+my $lagging = PostgreSQL::Test::Cluster->new('lagging');
+$lagging->init_from_backup($primary, 'backup', has_streaming => 1);
+$lagging->start;
+
+my $failover = PostgreSQL::Test::Cluster->new('failover');
+$failover->init_from_backup($primary, 'backup', has_streaming => 1);
+$failover->start;
+
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+
+test_checksum_state($primary, 'on');
+test_checksum_state($lagging, 'on');
+test_checksum_state($failover, 'on');
+
+# Offline-disable checksums on the lagging standby only. This is a
+# divergence 012_offline_standby.pl declares survivable: the standby keeps
+# its own state, warns, and stays readable.
+$lagging->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $lagging->data_dir);
+$lagging->start;
+test_checksum_state($lagging, 'off');
+
+# Give the lagging standby plenty of pages to write while it is locally
+# "off", and force a restartpoint so they reach disk without checksums.
+# The other standby replays the same WAL with checksums on.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+$lagging->safe_psql('postgres', "CHECKPOINT;");
+
+# Take the "failover" standby out of the picture. It is promoted from
+# this position later, so everything replayed from here on is the part of
+# the WAL that diverges and that pg_rewind will roll back. t1 is already
+# replayed everywhere, so its blocks are *not* rolled back.
+$failover->stop;
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$primary->wait_for_catchup($lagging);
+
+# Failover: promote the other standby, which forks the timeline behind the
+# replay position of the lagging standby.
+$primary->stop('fast');
+$failover->start;
+$failover->promote;
+$failover->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$failover->safe_psql('postgres', "CHECKPOINT;");
+
+# Re-attaching the lagging standby with pg_rewind must be refused: the
+# divergence checkpoint says "on" while the target is "off", so the rewound
+# server would verify checksums over the blocks it keeps.
+$lagging->stop;
+
+command_fails_like(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ],
+ qr/data checksums were enabled at the point of divergence but are disabled on the target server/,
+ 'pg_rewind refuses a target whose divergence checkpoint has checksums');
+
+# Enabling checksums offline on the target rewrites all of its pages, after
+# which the rewind is safe.
+system_or_bail('pg_checksums', '--enable', '--pgdata', $lagging->data_dir);
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ]);
+like($stderr, qr/Done!/, 'pg_rewind completes after the offline enable');
+
+$lagging->append_conf('postgresql.conf', 'port = ' . $lagging->port);
+$lagging->enable_streaming($failover);
+$lagging->set_standby_mode;
+$lagging->start;
+
+my ($rc, $out, $err) =
+ $lagging->psql('postgres',
+ "SELECT setting FROM pg_settings " . "WHERE name = 'data_checksums';");
+is($rc, 0, 'rewound standby accepts connections') or diag("stderr: $err");
+is($out, 'on', 'rewound standby resumes with checksums on');
+
+($rc, $out, $err) = $lagging->psql('postgres', "SELECT count(*) FROM t1;");
+is($rc, 0, 'blocks kept from the target stay readable')
+ or diag("stderr: $err");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($lagging->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the rewound standby');
+
+$lagging->stop('immediate');
+$failover->stop;
+done_testing();
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v15-0002-pg_checksums-Refuse-interrupted-transitions-note.patch (7.9K, ../../4BCE2F21-B669-46F3-B781-010F71B0B065@yesql.se/9-v15-0002-pg_checksums-Refuse-interrupted-transitions-note.patch)
download | inline diff:
From 782c04b1c65e992fe1769d0f371837aad7555698 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 3 Sep 2026 23:31:33 +0200
Subject: [PATCH v15 2/9] pg_checksums: Refuse interrupted transitions, note
change is local
An inprogress-on or inprogress-off control file means an online
transition was cut short; running the offline tool on top of it mixes
two procedures. Refuse it in every mode, with a hint pointing at the
start-stop cycle (or, on a standby, at letting replication finish)
that resets the state. Also print that the change applies to one
data directory only, with standby-aware wording, since in a
replication setup the same change must be made on every node.
On a primary, inprogress-on cannot survive a graceful stop: the
checksums launcher resolves it from its exit cleanup, and a crashed
primary is already rejected by the existing "cluster must be shut
down" check. inprogress-off can, however: it is set by the backend
running pg_disable_data_checksums(), and a fast shutdown arriving
between its two barriers leaves it behind in a cleanly shut down
control file, where the next start-stop cycle resolves it at end of
recovery. A standby can be stopped with either state, since it has
no launcher and only carries forward whatever state the last replayed
WAL record left it in, with a restartpoint persisting that as-is.
The standby is also the deterministic way to reach the new guard,
which is why the test coverage uses one; the primary window would
need an injection point between the two barriers.
The pg_checksums page said the tool still processes all relation files
regardless of an interrupted online transition; describe the refusal
instead.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Reviewed-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Discussion: https://postgr.es/m/anwm6UPxoVS41QA2@bdtpg
---
doc/src/sgml/ref/pg_checksums.sgml | 11 ++++---
src/bin/pg_checksums/pg_checksums.c | 21 +++++++++++++
.../test_checksums/t/012_offline_standby.pl | 31 ++++++++++++++-----
3 files changed, 51 insertions(+), 12 deletions(-)
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index 2d2057a8c19..abc035e5409 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -49,11 +49,12 @@ PostgreSQL documentation
</para>
<para>
- When enabling checksums with <application>pg_checksums</application>, if
- checksums were in the process of being enabled using
- <xref linkend="checksums-online-enable-disable"/> when the cluster was shut
- down, <application>pg_checksums</application> will still process all
- relation files regardless of the progress of online checksum processing.
+ If checksums were in the process of being enabled or disabled using
+ <xref linkend="checksums-online-enable-disable"/> when the cluster was
+ shut down, the control file still records that interrupted state, and
+ <application>pg_checksums</application> refuses to run in any mode.
+ Start the cluster and shut it down cleanly to reset the state, then
+ retry; on a standby, let replication complete the transition first.
</para>
<para>
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 5d6ea318784..3b58c6ca608 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -586,6 +586,21 @@ main(int argc, char *argv[])
ControlFile->state != DB_SHUTDOWNED_IN_RECOVERY)
pg_fatal("cluster must be shut down");
+ /*
+ * An inprogress state means an online transition was cut short. A
+ * standby stopped mid-transition carries either state; a cleanly shut
+ * down primary can still carry inprogress-off, which a fast shutdown
+ * during pg_disable_data_checksums() leaves behind, while inprogress-on
+ * is always resolved by the launcher's exit cleanup.
+ */
+ if (ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted");
+ pg_log_error_hint("Start and cleanly shut down the cluster once to reset the data checksum state, then retry. On a standby, let replication complete the transition first.");
+ exit(1);
+ }
+
if (ControlFile->data_checksum_version != PG_DATA_CHECKSUM_VERSION &&
mode == PG_MODE_CHECK)
pg_fatal("data checksums are not enabled in cluster");
@@ -673,6 +688,12 @@ main(int argc, char *argv[])
printf(_("Checksums enabled in cluster\n"));
else
printf(_("Checksums disabled in cluster\n"));
+
+ printf(_("This change applies to this data directory only.\n"));
+ if (ControlFile->state == DB_SHUTDOWNED_IN_RECOVERY)
+ printf(_("This node appears to be a standby; apply the same change to the primary and all other standbys.\n"));
+ else
+ printf(_("In a replication setup, apply the same change to every node while all are stopped, before restarting any of them.\n"));
}
return 0;
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
index c0ec374c8be..ba1f14e0399 100644
--- a/src/test/modules/test_checksums/t/012_offline_standby.pl
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -164,7 +164,10 @@ $newnode->stop('immediate');
# Converge the cluster: enable offline on the standby too.
$standby->stop;
-$standby->checksum_enable_offline;
+command_checks_all(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ 0, [qr/appears to be a standby/],
+ [], 'standby-role notice on offline enable');
$standby->start;
test_checksum_state($standby, 'on');
$primary->wait_for_catchup($standby);
@@ -232,11 +235,11 @@ unlike(
);
# Scenario 4: a standby stopped while replaying an interrupted online
-# transition keeps the interrupted state in its own control file. A
-# primary is never caught this way, as its checksums launcher resolves
-# inprogress-on back to off from its exit cleanup. A standby has no
-# launcher; it carries forward whatever the last replayed record left
-# it in.
+# transition keeps the interrupted state in its own control file, and
+# pg_checksums must refuse to touch it. A primary is never caught this
+# way, as its checksums launcher resolves inprogress-on back to off from
+# its exit cleanup. A standby has no launcher; it carries forward
+# whatever the last replayed record left it in.
# Block an online enable on the primary at inprogress-on with a
# blocking temp table, same trick as in 004_offline.pl.
@@ -250,6 +253,20 @@ wait_for_checksum_state($standby, 'inprogress-on');
# Stop the standby cleanly; its restartpoint persists inprogress-on to
# its own control file, since nothing on a standby resolves it away.
$standby->stop;
+
+command_fails_like(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --enable refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --check refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--disable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --disable refuses a standby stopped mid-transition');
+
$standby->start;
wait_for_checksum_state($standby, 'inprogress-on');
@@ -272,7 +289,7 @@ wait_for_checksum_state($primary, 'on');
$primary->wait_for_catchup($standby);
wait_for_checksum_state($standby, 'on');
-is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
'10001', 'standby readable once the transition completes');
# Scenario 4 continued: the backup taken mid-transition must also
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v15-0001-Do-not-adopt-data-checksum-state-from-another-no.patch (137.0K, ../../4BCE2F21-B669-46F3-B781-010F71B0B065@yesql.se/10-v15-0001-Do-not-adopt-data-checksum-state-from-another-no.patch)
download | inline diff:
From 7cbd6adadbc3360a3f191755ffc35a9ef7fb3cd8 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 3 Sep 2026 23:29:17 +0200
Subject: [PATCH v15 1/9] Do not adopt data checksum state from another node
during replay
Offline pg_checksums changes are local to one node, but replay adopted
the data checksum state carried by checkpoint records unconditionally.
After an offline enable on the primary, a standby whose pages were
never rewritten started verifying checksums it does not have; after an
offline change on a standby, the next replayed checkpoint silently
reverted it.
To fix, make the control file track this node's state alone, and have
replay cross-check the replayed state against it instead of adopting
it, warning once per divergent value and reporting when the states
agree again. pg_control gains a watermark, normally the end LSN of
the newest XLOG2_CHECKSUMS record the node has written or applied, so
that replay can skip transition records whose effect the control file
already contains, and a flag marking a state last written by
pg_checksums, which recovery must never overwrite with a replayed one.
Since the control file may only claim "on" once every page on disk
carries a checksum, persisting the state is tied to flushes:
XLOG2_CHECKSUMS replay persists every state but "on" immediately and
leaves "on" to the next restartpoint, checkpoints and restartpoints
only persist a state their flush ran under from beginning to end, and
transitions publish their state under the new
DataChecksumTransitionLock so that the states carried by WAL records
match their WAL order. The comments in xlog.c spell out the
individual rules.
Recovery from a base backup may be an exception to not adopting: its
control file was copied at an arbitrary moment, so the state carried
by the starting checkpoint is the one the WAL from there on was
written under. That holds for a backup taken from a primary, and only
when the copied state is not node local and its watermark is below the
starting checkpoint; a backup taken from a standby keeps the copied
state. pg_rewind keeps the target's own state and clamps the
watermark to the divergence point, since a watermark set on the
target's abandoned fork could numerically cover transition records the
source wrote after the divergence.
Bump PG_CONTROL_VERSION.
Document the offline procedure for replication setups, the lockstep
one: stop all nodes, run pg_checksums on each of them, and only then
restart. An offline change writes no WAL and has no ordering against
WAL a node has not replayed yet, so a node must not stop before
replaying all WAL of its upstream, or the change is overridden on
restart.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Author: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reported-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Discussion: https://postgr.es/m/anwm6UPxoVS41QA2@bdtpg
---
doc/src/sgml/ref/pg_checksums.sgml | 87 ++-
doc/src/sgml/wal.sgml | 23 +
src/backend/access/transam/xlog.c | 571 ++++++++++++++++--
src/backend/postmaster/datachecksum_state.c | 26 +-
.../utils/activity/wait_event_names.txt | 1 +
src/bin/pg_checksums/pg_checksums.c | 10 +
src/bin/pg_controldata/pg_controldata.c | 4 +
src/bin/pg_resetwal/pg_resetwal.c | 7 +
src/bin/pg_rewind/pg_rewind.c | 35 +-
src/bin/pg_upgrade/controldata.c | 2 +-
src/include/catalog/pg_control.h | 29 +-
src/include/storage/lwlocklist.h | 1 +
src/test/modules/test_checksums/Makefile | 2 +-
src/test/modules/test_checksums/meson.build | 10 +
.../test_checksums/t/012_offline_standby.pl | 304 ++++++++++
.../modules/test_checksums/t/013_rewind.pl | 201 ++++++
.../modules/test_checksums/t/014_lockstep.pl | 182 ++++++
.../t/015_standby_crash_after_disable.pl | 136 +++++
.../t/016_promote_enable_crash.pl | 146 +++++
.../test_checksums/t/017_restartpoint_race.pl | 151 +++++
.../t/018_enable_crash_windows.pl | 508 ++++++++++++++++
.../t/019_standby_shutdown_catchup.pl | 138 +++++
.../t/020_cascade_divergence.pl | 127 ++++
.../t/021_rewind_divergent_transitions.pl | 205 +++++++
24 files changed, 2822 insertions(+), 84 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/012_offline_standby.pl
create mode 100644 src/test/modules/test_checksums/t/013_rewind.pl
create mode 100644 src/test/modules/test_checksums/t/014_lockstep.pl
create mode 100644 src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
create mode 100644 src/test/modules/test_checksums/t/016_promote_enable_crash.pl
create mode 100644 src/test/modules/test_checksums/t/017_restartpoint_race.pl
create mode 100644 src/test/modules/test_checksums/t/018_enable_crash_windows.pl
create mode 100644 src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
create mode 100644 src/test/modules/test_checksums/t/020_cascade_divergence.pl
create mode 100644 src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index ae66fad3f0f..2d2057a8c19 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -242,15 +242,79 @@ PostgreSQL documentation
data directory must not be started or else data loss may occur.
</para>
<para>
- When using a replication setup with tools which perform direct copies
- of relation file blocks (for example <xref linkend="app-pgrewind"/>),
- enabling or disabling checksums can lead to page corruptions in the
- shape of incorrect checksums if the operation is not done consistently
- across all nodes. When enabling or disabling checksums in a replication
- setup, it is thus recommended to stop all the clusters before switching
- them all consistently. Destroying all standbys, performing the operation
- on the primary and finally recreating the standbys from scratch is also
- safe.
+ Enabling or disabling checksums with
+ <application>pg_checksums</application> changes only the local data
+ directory; the new state is not replicated to any other node. In a
+ replication setup the same change must be applied to every node:
+ </para>
+
+ <procedure>
+ <step id="shut-down-nodes">
+ <title>Shut down all nodes</title>
+ <para>
+ All nodes participating in the replication must be stopped with a
+ clean shutdown; <application>pg_checksums</application> refuses to run
+ on a data directory left behind by an immediate shutdown. Before
+ stopping a standby, make sure it has replayed all WAL of the primary:
+ stop the primary first, read its <quote>Latest checkpoint
+ location</quote> with <xref linkend="app-pgcontroldata"/>, and check
+ that <function>pg_last_wal_replay_lsn()</function> on the standby has
+ advanced past it. Comparing the replay position with
+ <function>pg_last_wal_receive_lsn()</function> is not enough, as it
+ only shows that the WAL the standby received has been replayed.
+ </para>
+ </step>
+
+ <step>
+ <title>Enable or disable data checksums on each node</title>
+ <para>
+ Run <application>pg_checksums</application> on the data directory of
+ each node in the replication setup. Nodes can be processed in
+ parallel while they are shut down. Processing must complete
+ successfully on all nodes before continuing.
+ </para>
+ </step>
+
+ <step>
+ <title>Restart all nodes</title>
+ <para>
+ Start the nodes normally, verify that
+ <xref linkend="guc-data-checksums"/> matches on all of them, and
+ monitor the logs of the standbys for data checksum state mismatch
+ warnings.
+ </para>
+ </step>
+ </procedure>
+
+ <para>
+ The replay requirement in <xref linkend="shut-down-nodes"/>exists
+ because an offline change is recorded only in the control file and
+ has no defined ordering against WAL the node has not replayed yet; see
+ <xref linkend="checksums-offline-enable-disable"/>. A node stopped
+ before replaying an online checksum state change applies that change
+ when it is restarted, overriding the offline change, and the states of
+ the nodes silently diverge until a later checkpoint record triggers
+ the warning described below. Because of this it is best not to mix
+ the two mechanisms: change the state of a replication setup either
+ with the offline procedure above or with an online transition, and
+ make sure the previous change has reached every node before starting
+ the next one.
+ </para>
+ <para>
+ If the change is applied inconsistently, each node keeps its own state,
+ and a standby logs a warning when the state recorded in the replayed WAL
+ differs from its own. The same warning can appear transiently while a
+ standby catches up over WAL written before a consistent change; it stops
+ once a checkpoint record carrying the new state has been replayed.
+ </para>
+ <para>
+ A standby whose data directory was never checksummed must not have
+ checksums enabled by catching up this way. Converge the cluster by
+ running <application>pg_checksums</application> on it while stopped as
+ outlined above, by enabling checksums online, or by recreating it from
+ a base backup. Note
+ that an online enable only starts from a primary whose checksums are off,
+ so if they are already enabled there, disable them online first.
</para>
<para>
If <application>pg_checksums</application> is aborted or killed while
@@ -258,6 +322,11 @@ PostgreSQL documentation
remains unchanged, and <application>pg_checksums</application> can be
re-run to perform the same operation.
</para>
+ <para>
+ Tools that copy relation file blocks directly between nodes, such as
+ <xref linkend="app-pgrewind"/>, require all nodes to be in
+ the same data checksum state, else there is risk for data corruption.
+ </para>
<para>
The target cluster must have the same major version as
<application>pg_checksums</application>.
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index ec62d17fbcc..dbbfdae5e4f 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -317,6 +317,29 @@
verify checksums, on an offline cluster.
</para>
+ <para>
+ An offline change provides durability differently from an
+ <link linkend="checksums-online-enable-disable">online change</link>.
+ An online transition is WAL-logged: it is ordered against all other
+ WAL records, it is replayed after a crash, and it propagates to
+ standbys. An offline change is recorded only in the cluster's
+ control file: it writes no WAL, it is invisible to replication, and
+ it has no defined ordering against WAL the node has not replayed
+ yet. When a node later replays WAL that contains an online checksum
+ state change, that change takes effect on the node even if it was
+ written before the offline change was made.
+ </para>
+
+ <para>
+ An offline change only affects the data directory it is run on; the
+ new state does not propagate over replication. In a replication setup
+ the same change must be applied to all nodes while all of them are
+ stopped, as described in <xref linkend="app-pgchecksums"/>. A standby
+ whose state diverges logs a warning but keeps its local setting. The
+ mismatch persists until the states are brought together again, with
+ the offline procedure or with an online transition; do this promptly.
+ </para>
+
</sect2>
<sect2 id="checksums-online-enable-disable" xreflabel="Online Enabling of Checksums">
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index c7a8d64c1f7..d1a8a16bd2d 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -556,14 +556,22 @@ typedef struct XLogCtlData
*/
XLogRecPtr lastFpwDisableRecPtr;
- /* last data_checksum_version we've seen */
+ /* current data checksum state of this node */
uint32 data_checksum_version;
+ /*
+ * Copies of control file fields with the same names, see pg_control.h for
+ * an in-depth description of these fields. Must be updated together with
+ * data_checksum_version under info_lck.
+ */
+ XLogRecPtr data_checksum_lsn;
+ bool data_checksum_is_local;
+
slock_t info_lck; /* locks shared variables shown above */
/*
* lastChecksumChangeRecPtr points to the end of the last XLOG2_CHECKSUMS
- * record inserted or replayed, i.e. the last change of
+ * record inserted or replayed which corresponds to the last change of
* data_checksum_version. InvalidXLogRecPtr if the state hasn't changed
* since the server started.
*/
@@ -690,6 +698,20 @@ static ChecksumStateType LocalDataChecksumState = 0;
*/
int data_checksums = 0;
+/*
+ * Whether replay of the next checkpoint-family record must adopt the data
+ * checksum state it carries. Set when recovery starts from a base backup,
+ * where the state at the redo point takes precedence over the control file
+ * copied with the backup later.
+ */
+static bool adoptChecksumStateFromNextCheckpoint = false;
+
+/*
+ * Sentinel for the data checksum mismatch warning tracking in
+ * CheckReplayedDataChecksumState(): no warning is outstanding.
+ */
+#define NO_WARNING_ISSUED PG_UINT32_MAX
+
/* For WALInsertLockAcquire/Release functions */
static int MyLockNo = 0;
static bool holdingAllLocks = false;
@@ -730,6 +752,8 @@ static void ValidateXLOGDirectoryStructure(void);
static void CleanupBackupHistory(void);
static void UpdateMinRecoveryPoint(XLogRecPtr lsn, bool force);
static bool PerformRecoveryXLogAction(void);
+static void CheckReplayedDataChecksumState(uint32 replayed_version);
+static void AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn);
static void InitControlFile(uint64 sysidentifier, uint32 data_checksum_version);
static void WriteControlFile(void);
static void ReadControlFile(void);
@@ -757,7 +781,7 @@ static void WALInsertLockAcquireExclusive(void);
static void WALInsertLockRelease(void);
static void WALInsertLockUpdateInsertingAt(XLogRecPtr insertingAt);
-static void XLogChecksums(uint32 new_type);
+static XLogRecPtr XLogChecksums(uint32 new_type);
/*
* Insert an XLOG record represented by an already-constructed chain of data
@@ -4292,6 +4316,8 @@ InitControlFile(uint64 sysidentifier, uint32 data_checksum_version)
ControlFile->track_commit_timestamp = track_commit_timestamp;
ControlFile->data_checksum_version = data_checksum_version;
ControlFile->data_checksum_version_init = data_checksum_version;
+ ControlFile->data_checksum_is_local = false;
+ ControlFile->data_checksum_lsn = InvalidXLogRecPtr;
/*
* Set the data_checksum_version value into XLogCtl, which is where all
@@ -4774,6 +4800,7 @@ void
SetDataChecksumsOnInProgress(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
/*
* The state transition is performed in a critical section with
@@ -4782,14 +4809,12 @@ SetDataChecksumsOnInProgress(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4827,6 +4852,8 @@ void
SetDataChecksumsOn(void)
{
uint64 barrier;
+ bool persist;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4847,23 +4874,11 @@ SetDataChecksumsOn(void)
SpinLockRelease(&XLogCtl->info_lck);
INJECTION_POINT("datachecksums-enable-checksums-delay", NULL);
+ INJECTION_POINT_LOAD("datachecksums-on-before-publish");
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_VERSION);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- SpinLockRelease(&XLogCtl->info_lck);
-
- /*
- * Update the controlfile before waiting since if we have an immediate
- * shutdown while waiting we want to come back up with checksums enabled.
- */
- LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- UpdateControlFile();
- LWLockRelease(ControlFileLock);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_VERSION);
barrier = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_CHECKSUM_ON);
@@ -4873,6 +4888,41 @@ SetDataChecksumsOn(void)
INJECTION_POINT("datachecksums-on-before-checkpoint", NULL);
RequestCheckpoint(CHECKPOINT_FORCE | CHECKPOINT_WAIT | CHECKPOINT_FAST);
+
+ INJECTION_POINT("datachecksums-on-after-checkpoint", NULL);
+
+ /*
+ * Persist "on" only now that the checkpoint has flushed the pages the
+ * transition rewrote. Crash recovery initializes verification from the
+ * control file but resumes from a checkpoint that can predate the
+ * transition, so an "on" persisted earlier would verify pages during
+ * replay whose rewrite never reached the disk: the pages found by the
+ * data checksums worker in shared buffers are not written back by its
+ * ring buffer.
+ *
+ * The checkpoint above normally persists the state itself; this covers
+ * the case where it started before the record written above and left the
+ * field alone. Crashing before this point is safe, as replay then
+ * re-establishes "on" from the full page images of the rewrite. Skip the
+ * write if the state moved on meanwhile, since whatever moved it persists
+ * its own. Compare the watermark rather than the state: a state
+ * comparison could not tell our transition from a later round trip back
+ * to "on".
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ persist = (XLogCtl->data_checksum_lsn == recptr);
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (persist)
+ {
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+ }
+
WaitForProcSignalBarrier(barrier);
}
@@ -4893,6 +4943,7 @@ void
SetDataChecksumsOff(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4918,14 +4969,12 @@ SetDataChecksumsOff(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4956,14 +5005,12 @@ SetDataChecksumsOff(void)
/* Ensure that we don't incur a checkpoint during disabling checksums */
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -5001,6 +5048,138 @@ SetLocalDataChecksumState(uint32 data_checksum_version)
data_checksums = data_checksum_version;
}
+/*
+ * CheckReplayedDataChecksumState
+ * Cross-check the data checksum state carried by a replayed checkpoint
+ * record against the state of this node.
+ *
+ * The state in the record belongs to whichever node wrote the WAL, and must
+ * not be adopted: an offline state change made with pg_checksums on one node
+ * of a replication set generates no WAL, and adopting would leak it into the
+ * other nodes through replay. Backup label recovery is the one exception,
+ * see AdoptReplayedDataChecksumState(). XLOG_CHECKPOINT_ONLINE needs no call
+ * here, as an online checkpoint's state already traveled in the preceding
+ * XLOG_CHECKPOINT_REDO record.
+ *
+ * Only archive recovery can see a lasting mismatch, as only there can the WAL
+ * and the control file come from different nodes or different times. In
+ * crash recovery a mismatch means replay resumed from a restartpoint
+ * predating an already-applied XLOG2_CHECKSUMS record, and replaying forward
+ * re-establishes the same state.
+ */
+static void
+CheckReplayedDataChecksumState(uint32 replayed_version)
+{
+ /*
+ * Warn once per remote value, so a lasting mismatch does not flood the
+ * log. Matching states re-arm the warning. Backend-local state is
+ * enough: replay only runs in the startup process, and restarting it
+ * re-arms as well.
+ */
+ static uint32 last_warned_version = NO_WARNING_ISSUED;
+ uint32 local_version;
+
+ if (!ArchiveRecoveryRequested)
+ return;
+
+ /*
+ * Re-replayed WAL below the consistency point was already cross-checked
+ * before minRecoveryPoint was last persisted, and the persisted state can
+ * legitimately be newer than what checkpoint records there carry:
+ * XLOG2_CHECKSUMS replay persists most states ahead of the restartpoint
+ * horizon. In particular the checkpoint record recovery restarts from is
+ * such a re-replay.
+ */
+ if (!reachedConsistency)
+ return;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ local_version = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (replayed_version == local_version)
+ {
+ /*
+ * Report convergence if this process warned before. Nothing else
+ * tells the operator that running pg_checksums on the other nodes, or
+ * a rebuild, took effect.
+ */
+ if (last_warned_version != NO_WARNING_ISSUED)
+ ereport(LOG,
+ errmsg("data checksum state \"%s\" of this node now agrees with the replayed WAL",
+ get_checksum_state_string(local_version)));
+
+ last_warned_version = NO_WARNING_ISSUED;
+ return;
+ }
+
+ /* the nodes legitimately differ while an online transition runs */
+ if (replayed_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ replayed_version == PG_DATA_CHECKSUM_INPROGRESS_OFF ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ return;
+
+ if (replayed_version == last_warned_version)
+ return;
+ last_warned_version = replayed_version;
+
+ ereport(WARNING,
+ errmsg("data checksum state \"%s\" of this node does not match the state \"%s\" in the replayed WAL",
+ get_checksum_state_string(local_version),
+ get_checksum_state_string(replayed_version)),
+ errdetail("The data checksum state may have been changed with pg_checksums on another node."),
+ errhint("Apply the same change with pg_checksums on the primary and all standby servers, or rebuild this server from a base backup."));
+}
+
+/*
+ * AdoptReplayedDataChecksumState
+ * Adopt the data checksum state at the redo point of backup label
+ * recovery.
+ *
+ * The state is persisted immediately so that a crash before the first
+ * restartpoint does not resurrect the state copied with the backup; a crash
+ * at this point restarts from the same redo point, so the control file does
+ * not run ahead of the replay position. If the value is unchanged the
+ * control file already carries it, so both the barrier and the persist are
+ * skipped.
+ *
+ * lsn is the location of the checkpoint-family record the state was taken
+ * from and becomes the new watermark: the adopted state covers everything
+ * below the redo point, which replay never revisits. For the same reason
+ * the old watermark can stay when the value is unchanged: the records
+ * between the two positions are never replayed, so nothing depends on
+ * which one is recorded.
+ */
+static void
+AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn)
+{
+ bool changed = false;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (XLogCtl->data_checksum_version != new_version)
+ {
+ XLogCtl->data_checksum_version = new_version;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(new_version);
+ changed = true;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (!changed)
+ return;
+
+ EmitAndWaitDataChecksumsBarrier(new_version);
+
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = new_version;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+}
+
/* guc hook */
const char *
show_data_checksums(void)
@@ -5454,6 +5633,8 @@ XLOGShmemInit(void *arg)
/* Use the checksum info from control file */
XLogCtl->data_checksum_version = ControlFile->data_checksum_version;
+ XLogCtl->data_checksum_lsn = ControlFile->data_checksum_lsn;
+ XLogCtl->data_checksum_is_local = ControlFile->data_checksum_is_local;
SetLocalDataChecksumState(XLogCtl->data_checksum_version);
SpinLockInit(&XLogCtl->Insert.insertpos_lck);
@@ -6031,6 +6212,51 @@ StartupXLOG(void)
SetCommitTsLimit(checkPoint.oldestCommitTsXid,
checkPoint.newestCommitTsXid);
+ /*
+ * When recovery starts from a base backup, the control file was copied at
+ * an arbitrary moment and its data checksum state may differ from the
+ * state at the redo point, which is what the WAL from there on was
+ * written under. Adopt the state of the starting checkpoint: a shutdown
+ * checkpoint is not replayed, so take it from the record read above; the
+ * redo point of an online checkpoint is its CHECKPOINT_REDO record, so
+ * let the replay of that record adopt it. Check backupStartPoint in
+ * addition to the label: on a crash restart during backup recovery the
+ * label file is already renamed away, but the start point persists until
+ * the backup end record.
+ *
+ * Not for a base backup taken from a standby, though. Its starting
+ * checkpoint is the standby's last restartpoint, a record written by the
+ * upstream primary, whose state is not the one the copied files were
+ * written under. The copied control file is already correct: a standby
+ * persists its state only at restartpoint horizons and never claims more
+ * than what reached disk. Such backups are recognized by backupEndPoint
+ * together with backupEndRequired; backupEndPoint is only set for "BACKUP
+ * FROM: standby" labels and persists across a crash restart. pg_rewind
+ * writes a standby label as well, but no backupEndPoint, and its recovery
+ * keeps adopting: the control file it installs carries the target's own
+ * checksum state, which can lag the redo point of the last common
+ * checkpoint the same way a restartpoint horizon can.
+ *
+ * Never adopt over a state the control file's watermark or local flag
+ * marks as newer than the starting checkpoint. A pg_checksums change is
+ * local to the node and generates no WAL, so nothing in the replayed WAL
+ * could ever restore it once overwritten; and a watermark above the redo
+ * point means the control file already contains the effect of every
+ * transition record up to there.
+ */
+ if ((haveBackupLabel || XLogRecPtrIsValid(ControlFile->backupStartPoint)) &&
+ !(XLogRecPtrIsValid(ControlFile->backupEndPoint) &&
+ ControlFile->backupEndRequired) &&
+ !ControlFile->data_checksum_is_local &&
+ checkPoint.redo > ControlFile->data_checksum_lsn)
+ {
+ if (wasShutdown)
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ checkPoint.redo);
+ else
+ adoptChecksumStateFromNextCheckpoint = true;
+ }
+
/*
* Clear out any old relcache cache files. This is *necessary* if we do
* any WAL replay, since that would probably result in the cache files
@@ -6634,11 +6860,7 @@ StartupXLOG(void)
if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
ereport(WARNING,
@@ -6655,11 +6877,7 @@ StartupXLOG(void)
else if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
}
@@ -6684,6 +6902,8 @@ StartupXLOG(void)
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
XLogCtl->SharedRecoveryState = RECOVERY_STATE_DONE;
SpinLockRelease(&XLogCtl->info_lck);
@@ -6815,6 +7035,23 @@ static bool
PerformRecoveryXLogAction(void)
{
bool promoted = false;
+ bool flushForChecksums;
+ uint32 checksum_state;
+
+ /*
+ * The end-of-recovery record persists the data checksum state without
+ * flushing the buffer pool, but the control file may only claim "on" once
+ * every page on disk carries a checksum. If replay entered that state
+ * without a restartpoint following it, the pages rewritten by the
+ * transition are still only in the buffer pool, so take the full
+ * checkpoint below instead of the lightweight record.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ flushForChecksums = (checksum_state == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile->data_checksum_version != checksum_state);
/*
* Perform a checkpoint to update all our recovery activity to disk.
@@ -6830,7 +7067,7 @@ PerformRecoveryXLogAction(void)
* fully out of recovery mode and already accepting queries.
*/
if (ArchiveRecoveryRequested && IsUnderPostmaster &&
- PromoteIsTriggered())
+ PromoteIsTriggered() && !flushForChecksums)
{
promoted = true;
@@ -7437,6 +7674,7 @@ CreateCheckPoint(int flags)
uint32 freespace;
XLogRecPtr PriorRedoPtr;
XLogRecPtr last_important_lsn;
+ XLogRecPtr checksumLsn;
VirtualTransactionId *vxids;
int nvxids;
int oldXLogAllowed = 0;
@@ -7549,11 +7787,14 @@ CreateCheckPoint(int flags)
checkPoint.wal_level = wal_level;
/*
- * Get the current data_checksum_version value from xlogctl, valid at the
- * time of the checkpoint.
+ * Get the current data_checksum_version value from xlogctl. This is
+ * final only for a shutdown checkpoint, where no concurrent transition is
+ * possible; an online checkpoint resamples it together with the redo
+ * record below.
*/
SpinLockAcquire(&XLogCtl->info_lck);
checkPoint.dataChecksumState = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
if (shutdown)
@@ -7610,10 +7851,21 @@ CreateCheckPoint(int flags)
{
xl_checkpoint_redo redo_rec;
+ /*
+ * Sample the data checksum state and insert the redo record under
+ * DataChecksumTransitionLock, so that a concurrent transition cannot
+ * insert its XLOG2_CHECKSUMS record between the sampling and the
+ * insertion below. Without this, the redo record could follow the
+ * transition record in WAL while carrying the pre-transition state,
+ * and recovery resuming here would never learn about the transition.
+ * See XLogChecksums().
+ */
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
WALInsertLockAcquire();
redo_rec.wal_level = wal_level;
SpinLockAcquire(&XLogCtl->info_lck);
redo_rec.data_checksum_version = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
WALInsertLockRelease();
@@ -7621,6 +7873,14 @@ CreateCheckPoint(int flags)
XLogBeginInsert();
XLogRegisterData(&redo_rec, sizeof(xl_checkpoint_redo));
(void) XLogInsert(RM_XLOG_ID, XLOG_CHECKPOINT_REDO);
+ LWLockRelease(DataChecksumTransitionLock);
+
+ /*
+ * The checkpoint record must carry the same state as the redo record
+ * just inserted: the sample taken before redo determination can be
+ * stale by now, and the pair would otherwise disagree.
+ */
+ checkPoint.dataChecksumState = redo_rec.data_checksum_version;
/*
* XLogInsertRecord will have updated XLogCtl->Insert.RedoRecPtr in
@@ -7824,6 +8084,41 @@ CreateCheckPoint(int flags)
ControlFile->minRecoveryPoint = InvalidXLogRecPtr;
ControlFile->minRecoveryPointTLI = 0;
+ /*
+ * Persist the data checksum state this node runs under into the control
+ * file. Only the top-level field tracks this node, checkPointCopy is a
+ * historical record used to resume replay.
+ *
+ * checkPoint.dataChecksumState was sampled while holding the
+ * DataChecksumTransitionLock together with the redo record, so it is the
+ * state in effect at the redo point. If it was "on", the XLOG2_CHECKSUMS
+ * record announcing that precedes the redo point and every page the
+ * transition rewrote was dirtied before it, so CheckPointGuts() has just
+ * written all of them out. Recording the state here is what keeps a
+ * finished transition from being resolved as interrupted when this
+ * checkpoint is the one crash recovery resumes from: replay never sees
+ * the record announcing it.
+ *
+ * Persist it only if the state did not change while the flush was in
+ * progress. If it changed in between, the pages written out straddle two
+ * states, and the newer one could claim checksums that pages already on
+ * disk do not carry; leave the field to the next checkpoint then, the
+ * transition itself has already persisted every state that is safe
+ * without a flush.
+ *
+ * Compare the watermark rather than the state: record positions are
+ * unique, so a full round trip back to the sampled state cannot alias,
+ * while its flushed pages straddle the intermediate states all the same.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksumLsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checkPoint.dataChecksumState;
+ ControlFile->data_checksum_lsn = checksumLsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
/*
* Persist unloggedLSN value. It's reset on crash recovery, so this goes
* unused on non-shutdown checkpoints, but seems useful to store it always
@@ -7968,9 +8263,11 @@ CreateEndOfRecoveryRecord(void)
ControlFile->minRecoveryPoint = recptr;
ControlFile->minRecoveryPointTLI = xlrec.ThisTimeLineID;
- /* start with the latest checksum version (as of the end of recovery) */
+ /* persist the data checksum state this node ended recovery with */
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
@@ -8175,6 +8472,9 @@ CreateRestartPoint(int flags)
XLogRecPtr endptr;
XLogSegNo _logSegNo;
TimestampTz xtime;
+ uint32 checksum_state;
+ XLogRecPtr checksum_lsn;
+ bool checksum_is_local;
/* Concurrent checkpoint/restartpoint cannot happen */
Assert(!IsUnderPostmaster || MyBackendType == B_CHECKPOINTER);
@@ -8221,8 +8521,47 @@ CreateRestartPoint(int flags)
UpdateMinRecoveryPoint(InvalidXLogRecPtr, true);
if (flags & CHECKPOINT_IS_SHUTDOWN)
{
+ bool catchUpChecksums;
+
+ /*
+ * There is no new restartpoint to persist the data checksum state
+ * with, but a cleanly stopped node should not leave the control
+ * file behind the state replay reached: pg_checksums and
+ * pg_rewind read it, and an in-progress state there makes them
+ * refuse to run. Catching it up needs the same guarantee a
+ * restartpoint gives: that every page on disk carries a checksum,
+ * so flush the buffer pool first. Only a transition to "on" that
+ * no restartpoint followed can get here; the other states are
+ * already persisted by XLOG2_CHECKSUMS replay. Replay has ended
+ * by now, so the state cannot change under us.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ checksum_is_local = XLogCtl->data_checksum_is_local;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockAcquire(ControlFileLock, LW_SHARED);
+ catchUpChecksums =
+ (checksum_lsn != ControlFile->data_checksum_lsn &&
+ XLogRecPtrIsValid(lastCheckPointRecPtr));
+ LWLockRelease(ControlFileLock);
+
+ if (catchUpChecksums)
+ {
+ MemSet(&CheckpointStats, 0, sizeof(CheckpointStats));
+ CheckpointStats.ckpt_start_t = GetCurrentTimestamp();
+ CheckPointGuts(lastCheckPoint.redo, flags);
+ }
+
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
+ if (catchUpChecksums)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = checksum_is_local;
+ }
UpdateControlFile();
LWLockRelease(ControlFileLock);
}
@@ -8263,6 +8602,17 @@ CreateRestartPoint(int flags)
/* Update the process title */
update_checkpoint_display(flags, true, false);
+ /*
+ * Note the data checksum state the flush below starts under. Replay runs
+ * concurrently and can change the state while the flush is in progress,
+ * in which case the flush covers pages written under both states; see
+ * where the state is persisted further down.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
CheckPointGuts(lastCheckPoint.redo, flags);
/*
@@ -8322,8 +8672,26 @@ CreateRestartPoint(int flags)
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
}
- /* we shall start with the latest checksum version */
- ControlFile->data_checksum_version = lastCheckPoint.dataChecksumState;
+ /*
+ * Persist the data checksum state of this node. Not the state of the
+ * replayed checkpoint: that one belongs to the node that wrote it and
+ * may differ after an offline change on either side.
+ * ControlFile->checkPointCopy above keeps the replayed value on
+ * purpose, being a historical record used to resume replay rather
+ * than a tracker of node state.
+ *
+ * Persist only if the flush above ran under one state throughout; see
+ * CreateCheckPoint() for why, including why this compares the
+ * watermark and not the state.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksum_lsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
}
@@ -8764,9 +9132,22 @@ XLogReportParameters(void)
}
/*
- * Log the new state of checksums
+ * XLogChecksums
+ * Log and publish the new state of checksums
+ *
+ * Inserting the record and publishing the new state must be atomic with
+ * respect to a checkpoint sampling the state for its XLOG_CHECKPOINT_REDO
+ * record: without that, a checkpoint could read the old state after the
+ * record is already in WAL and insert a redo record that both follows the
+ * transition in WAL order and carries the pre-transition state. Recovery
+ * resuming from such a redo point would never replay the transition record
+ * and resolve the finished transition as interrupted.
+ * DataChecksumTransitionLock serializes the two; see CreateCheckPoint().
+ *
+ * Returns the end LSN of the inserted record, which the caller persists
+ * together with the new state as the data checksum watermark.
*/
-static void
+static XLogRecPtr
XLogChecksums(uint32 new_type)
{
xl_checksum_state xlrec;
@@ -8774,12 +9155,28 @@ XLogChecksums(uint32 new_type)
xlrec.new_checksum_state = new_type;
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
+
XLogBeginInsert();
XLogRegisterData((char *) &xlrec, sizeof(xl_checksum_state));
recptr = XLogInsert(RM_XLOG2_ID, XLOG2_CHECKSUMS);
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, recptr);
+
+ /* only loaded by SetDataChecksumsOn(), a no-op for the other callers */
+ INJECTION_POINT_CACHED("datachecksums-on-before-publish", NULL);
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ XLogCtl->data_checksum_version = new_type;
+ XLogCtl->data_checksum_lsn = recptr;
+ XLogCtl->data_checksum_is_local = false;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockRelease(DataChecksumTransitionLock);
+
XLogFlush(recptr);
+
+ return recptr;
}
/*
@@ -8968,11 +9365,19 @@ xlog_redo(XLogReaderState *record)
/* ControlFile->checkPointCopy always tracks the latest ckpt XID */
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->checkPointCopy.nextXid = checkPoint.nextXid;
- ControlFile->data_checksum_version = checkPoint.dataChecksumState;
UpdateControlFile();
LWLockRelease(ControlFileLock);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(checkPoint.dataChecksumState);
+
/*
* We should've already switched to the new TLI before replaying this
* record.
@@ -9207,19 +9612,17 @@ xlog_redo(XLogReaderState *record)
else if (info == XLOG_CHECKPOINT_REDO)
{
xl_checkpoint_redo redo_rec;
- bool new_state = false;
memcpy(&redo_rec, XLogRecGetData(record), sizeof(xl_checkpoint_redo));
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = redo_rec.data_checksum_version;
- SetLocalDataChecksumState(redo_rec.data_checksum_version);
- if (redo_rec.data_checksum_version != ControlFile->data_checksum_version)
- new_state = true;
- SpinLockRelease(&XLogCtl->info_lck);
-
- if (new_state)
- EmitAndWaitDataChecksumsBarrier(redo_rec.data_checksum_version);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(redo_rec.data_checksum_version,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(redo_rec.data_checksum_version);
}
else if (info == XLOG_LOGICAL_DECODING_STATUS_CHANGE)
{
@@ -9281,25 +9684,63 @@ xlog2_redo(XLogReaderState *record)
{
xl_checksum_state state;
XLogRecPtr lsn = record->EndRecPtr;
+ XLogRecPtr watermark;
memcpy(&state, XLogRecGetData(record), sizeof(xl_checksum_state));
+ SpinLockAcquire(&XLogCtl->info_lck);
+ watermark = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ /*
+ * Skip records this node has already applied. The control file
+ * carries the watermark, so this holds across restarts: recovery
+ * resuming below a record whose effect the control file already
+ * contains must not re-apply it, or it would revert a state change
+ * made with pg_checksums in between, which moves the state without
+ * writing any record of its own.
+ */
+ if (lsn <= watermark)
+ return;
+
/* advertise the location before the new state becomes visible */
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, lsn);
SpinLockAcquire(&XLogCtl->info_lck);
XLogCtl->data_checksum_version = state.new_checksum_state;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(state.new_checksum_state);
SpinLockRelease(&XLogCtl->info_lck);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = state.new_checksum_state;
+
+ /*
+ * Persist the new state, except when it is "on". Only "on" verifies
+ * checksums during reads, and between the last restartpoint and this
+ * record there may be pages on disk flushed under the old state; a
+ * crash-restart initializes verification from the control file and
+ * replay reads those pages back, so the control file may only say
+ * "on" once everything written under the transition has been flushed,
+ * as restartpoints and the end of recovery do. The opposite direction
+ * cannot wait for the restartpoint: once this record is replayed,
+ * evicted pages are written without checksums, and a control file
+ * still saying "on" would fail verification on exactly those pages
+ * after a crash.
+ */
+ if (state.new_checksum_state != PG_DATA_CHECKSUM_VERSION)
+ {
+ ControlFile->data_checksum_version = state.new_checksum_state;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ }
/*
* Update minRecoveryPoint to ensure that if recovery is aborted, we
* recover back up to this point before allowing hot standby again.
- * The new state is durable in pg_control while its location is only
- * tracked in shared memory; a standby becoming consistent below this
- * record would let base backups resume checksum verification with the
+ * The change location is only tracked in shared memory and is lost
+ * over a restart; a standby becoming consistent below this record
+ * would let base backups resume checksum verification with the
* location unknown. The local copies cannot be updated as long as
* crash recovery is happening and we expect all the WAL to be
* replayed.
diff --git a/src/backend/postmaster/datachecksum_state.c b/src/backend/postmaster/datachecksum_state.c
index f69258bc33d..099a6b4fe2e 100644
--- a/src/backend/postmaster/datachecksum_state.c
+++ b/src/backend/postmaster/datachecksum_state.c
@@ -94,9 +94,9 @@
*
* If processing is started in an online cluster then all backends are in Bd.
* If processing was halted by the cluster shutting down (due to a crash or
- * intentional restart), the controlfile state "inprogress-on" will be observed
- * on system startup and all backends will be placed in Bd. The controlfile
- * state will also be set to "off".
+ * intentional restart), the control file state "inprogress-on" will be
+ * observed on system startup and all backends will be placed in Bd. The
+ * control file state will also be set to "off".
*
* Backends transition Bd -> Bi via a procsignalbarrier which is emitted by the
* DataChecksumsWorkerLauncherMain. When all backends have acknowledged the
@@ -146,7 +146,25 @@
* stop writing data checksums as no backend is enforcing data checksum
* validation any longer.
*
- * 4. Future opportunities for optimizations
+ * 4. Interaction with offline data checksum changes
+ * -------------------------------------------------
+ * Enabling or disabling checksums offline with pg_checksums uses none of the
+ * machinery in this file, but the two mechanisms share the state kept in the
+ * control file, so their interaction is documented here.
+ *
+ * pg_checksums writes the new state to the control file and sets
+ * data_checksum_is_local, marking a state that no WAL record accounts for.
+ * Recovery then does not adopt the state carried by a replayed checkpoint
+ * record over it. The control file also carries a watermark, the WAL
+ * position through which data checksum transitions are covered. Replay skips
+ * transition records ending at or below the watermark, as their effect is
+ * already contained in the control file, and applies records above it as
+ * usual, whether they were written before or after an offline change. This
+ * is why an offline change in a replicated setup must be made on every node
+ * while all of them are stopped and caught up; see the pg_checksums
+ * documentation for the procedure.
+ *
+ * 5. Future opportunities for optimizations
* -----------------------------------------
* Below are some potential optimizations and improvements which were brought
* up during reviews of this feature, but which weren't implemented in the
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 0a70cffa081..3d366fd1114 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -371,6 +371,7 @@ WaitLSN "Waiting to read or update shared Wait-for-LSN state."
LogicalDecodingControl "Waiting to read or update logical decoding status information."
DataChecksumsWorker "Waiting for data checksums worker."
AioWorkerControl "Waiting to update AIO worker information."
+DataChecksumTransition "Waiting for a data checksum state transition to be written to WAL."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 3b3ae23f1a6..5d6ea318784 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -648,6 +648,16 @@ main(int argc, char *argv[])
ControlFile->data_checksum_version =
(mode == PG_MODE_ENABLE) ? PG_DATA_CHECKSUM_VERSION : PG_DATA_CHECKSUM_OFF;
+ /*
+ * Mark the state as changed locally, without a WAL record. Recovery
+ * then does not let a replayed checkpoint overwrite it, as no record
+ * could restore the change afterwards. The watermark is left alone:
+ * XLOG2_CHECKSUMS records at or below it stay covered, while records
+ * above it, which this node has not applied yet, still take effect on
+ * replay no matter when they were written.
+ */
+ ControlFile->data_checksum_is_local = true;
+
if (do_sync)
{
pg_log_info("syncing data directory");
diff --git a/src/bin/pg_controldata/pg_controldata.c b/src/bin/pg_controldata/pg_controldata.c
index b785f7f4070..33f0e9e4ea9 100644
--- a/src/bin/pg_controldata/pg_controldata.c
+++ b/src/bin/pg_controldata/pg_controldata.c
@@ -349,6 +349,10 @@ main(int argc, char *argv[])
(ControlFile->float8ByVal ? _("by value") : _("by reference")));
printf(_("Data page checksum version: %u\n"),
ControlFile->data_checksum_version);
+ printf(_("Data checksum watermark: %X/%08X\n"),
+ LSN_FORMAT_ARGS(ControlFile->data_checksum_lsn));
+ printf(_("Data checksum state is node-local: %s\n"),
+ (ControlFile->data_checksum_is_local ? _("yes") : _("no")));
printf(_("Default char data signedness: %s\n"),
(ControlFile->default_char_signedness ? _("signed") : _("unsigned")));
printf(_("Mock authentication nonce: %s\n"),
diff --git a/src/bin/pg_resetwal/pg_resetwal.c b/src/bin/pg_resetwal/pg_resetwal.c
index 63e4381e03f..b11f76a401e 100644
--- a/src/bin/pg_resetwal/pg_resetwal.c
+++ b/src/bin/pg_resetwal/pg_resetwal.c
@@ -923,6 +923,13 @@ RewriteControlFile(void)
ControlFile.backupEndPoint = InvalidXLogRecPtr;
ControlFile.backupEndRequired = false;
+ /*
+ * The old WAL is gone and the new position may lie below the old
+ * watermark, which would make replay ignore future checksum transition
+ * records. The state itself is kept.
+ */
+ ControlFile.data_checksum_lsn = InvalidXLogRecPtr;
+
/*
* Force the defaults for max_* settings. The values don't really matter
* as long as wal_level='minimal'; the postmaster will reset these fields
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 2e86fd158d0..0e4c2df3f4b 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -37,7 +37,8 @@ static void usage(const char *progname);
static void perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo);
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec);
static void createBackupLabel(XLogRecPtr startpoint, TimeLineID starttli,
XLogRecPtr checkpointloc);
@@ -531,7 +532,8 @@ main(int argc, char **argv)
* This is the point of no return. Once we start copying things, there is
* no turning back!
*/
- perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo);
+ perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo,
+ divergerec);
if (showprogress)
pg_log_info("syncing target data directory");
@@ -566,7 +568,8 @@ static void
perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo)
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec)
{
XLogRecPtr endrec;
TimeLineID endtli;
@@ -738,6 +741,32 @@ perform_rewind(filemap_t *filemap, rewind_source *source,
ControlFile_new.minRecoveryPoint = endrec;
ControlFile_new.minRecoveryPointTLI = endtli;
ControlFile_new.state = DB_IN_ARCHIVE_RECOVERY;
+
+ /*
+ * Keep the target's own data checksum state. Most of the data directory
+ * is still the target's: only blocks it changed since the divergence were
+ * copied from the source, so the source's state says nothing about the
+ * pages that stay. Replay from the last common checkpoint applies any
+ * WAL-logged transition the target has not seen (the watermark tells them
+ * apart), which converges the rewound server to the source's state
+ * whenever the WAL carries it.
+ */
+ ControlFile_new.data_checksum_version = ControlFile_target.data_checksum_version;
+ ControlFile_new.data_checksum_lsn = ControlFile_target.data_checksum_lsn;
+ ControlFile_new.data_checksum_is_local = ControlFile_target.data_checksum_is_local;
+
+ /*
+ * The watermark is only meaningful within the history the node replays.
+ * Records at or below the divergence point are common to both histories
+ * and stay covered, but a watermark above it was set by a transition
+ * record on the target's own abandoned fork: numerically it can cover
+ * transition records the source wrote after the divergence, and replay
+ * would skip them as already applied. Clamp it to the divergence point,
+ * so that every transition record on the source's history takes effect.
+ */
+ if (ControlFile_new.data_checksum_lsn > divergerec)
+ ControlFile_new.data_checksum_lsn = divergerec;
+
if (!dry_run)
update_controlfile(datadir_target, &ControlFile_new, do_sync);
}
diff --git a/src/bin/pg_upgrade/controldata.c b/src/bin/pg_upgrade/controldata.c
index f0fa2b1689f..759bd4ef77c 100644
--- a/src/bin/pg_upgrade/controldata.c
+++ b/src/bin/pg_upgrade/controldata.c
@@ -431,7 +431,7 @@ get_control_data(ClusterInfo *cluster)
cluster->controldata.date_is_int = strstr(p, "64-bit integers") != NULL;
got_date_is_int = true;
}
- else if ((p = strstr(bufin, "checksum")) != NULL)
+ else if ((p = strstr(bufin, "Data page checksum version:")) != NULL)
{
p = strchr(p, ':');
diff --git a/src/include/catalog/pg_control.h b/src/include/catalog/pg_control.h
index 89ab43dd4fc..c3c934d0012 100644
--- a/src/include/catalog/pg_control.h
+++ b/src/include/catalog/pg_control.h
@@ -22,7 +22,7 @@
/* Version identifier for this pg_control format */
-#define PG_CONTROL_VERSION 2000
+#define PG_CONTROL_VERSION 2001
/* Nonce key length, see below */
#define MOCK_AUTH_NONCE_LEN 32
@@ -237,6 +237,33 @@ typedef struct ControlFileData
/* Current data checksums state */
uint32 data_checksum_version;
+ /*
+ * WAL position through which data checksum transitions are covered.
+ * Replay ignores XLOG2_CHECKSUMS records ending at or below this point:
+ * their effect is already contained in data_checksum_version, or an
+ * offline pg_checksums change made after they were first applied
+ * supersedes them. Ordinarily this is the end of the newest such record
+ * this node has written or applied, but a tool may store any position
+ * that covers the same set of records. If the node has never written or
+ * applied such a record this field shall be set to InvalidXLogRecPtr.
+ *
+ * The comparison has no timeline context, so the value is only valid
+ * within the WAL history this node replays. A tool that moves the node
+ * to another history must clamp the watermark to the point where the
+ * histories fork, as pg_rewind does, or reset it, as pg_resetwal does.
+ */
+ XLogRecPtr data_checksum_lsn;
+
+ /*
+ * True when data_checksum_version was last set by pg_checksums in an
+ * offline operation rather than by an online, WAL-logged transition. Such
+ * a state is local to this node and not derived from WAL, so recovery
+ * must not replace it with a state taken from a checkpoint record;
+ * nothing in the WAL could restore the change once it is overwritten.
+ * Cleared by the next WAL-logged transition.
+ */
+ bool data_checksum_is_local;
+
/*
* True if the default signedness of char is "signed" on a platform where
* the cluster is initialized.
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index d7eb648bd27..8d858be9927 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -89,6 +89,7 @@ PG_LWLOCK(54, WaitLSN)
PG_LWLOCK(55, LogicalDecodingControl)
PG_LWLOCK(56, DataChecksumsWorker)
PG_LWLOCK(57, AioWorkerControl)
+PG_LWLOCK(58, DataChecksumTransition)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/test/modules/test_checksums/Makefile b/src/test/modules/test_checksums/Makefile
index 71455cd5577..80f54bf6d8c 100644
--- a/src/test/modules/test_checksums/Makefile
+++ b/src/test/modules/test_checksums/Makefile
@@ -9,7 +9,7 @@
#
#-------------------------------------------------------------------------
-EXTRA_INSTALL = src/test/modules/injection_points
+EXTRA_INSTALL = contrib/pg_buffercache src/test/modules/injection_points
export enable_injection_points
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index fb7129d796f..e5d38fafb7d 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -35,6 +35,16 @@ tests += {
't/009_fpi.pl',
't/010_backup_straddle.pl',
't/011_standby_straddle.pl',
+ 't/012_offline_standby.pl',
+ 't/013_rewind.pl',
+ 't/014_lockstep.pl',
+ 't/015_standby_crash_after_disable.pl',
+ 't/016_promote_enable_crash.pl',
+ 't/017_restartpoint_race.pl',
+ 't/018_enable_crash_windows.pl',
+ 't/019_standby_shutdown_catchup.pl',
+ 't/020_cascade_divergence.pl',
+ 't/021_rewind_divergent_transitions.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
new file mode 100644
index 00000000000..c0ec374c8be
--- /dev/null
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -0,0 +1,304 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Offline checksum changes with pg_checksums are local to one node. A
+# standby must neither adopt the state of the primary from replayed
+# checkpoint records, nor lose its own offline change to them.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Scenario 1: enable offline on the primary only. The standby must
+# stay off, warn about the mismatch, and remain readable.
+$standby->stop;
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+test_checksum_state($primary, 'on');
+test_checksum_state($standby, 'off');
+
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after offline enable on the primary');
+
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+my @warnings = $log =~ /(does not match the state)/g;
+is(scalar(@warnings), 1, 'mismatch warned once per remote value');
+
+# Matching states re-arm the warning: undo the divergence on the primary,
+# then diverge again to the same value, all without restarting the standby.
+$primary->stop;
+$primary->checksum_disable_offline;
+$logstart = -s $standby->logfile;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no warning while the states match again');
+like($log, qr/now agrees with the replayed WAL/,
+ 'convergence is reported once the states match again');
+
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+test_checksum_state($standby, 'off');
+
+# The local state survives both clean and immediate restarts.
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+# Still in scenario 1's divergence: take a base backup *from the
+# standby*. A backup taken on a standby uses the last restartpoint as
+# its starting checkpoint (do_pg_backup_start()), so the record at the
+# redo point was written by the upstream primary and carries the
+# primary's state. Recovery from such a backup must not adopt that
+# state: the files were copied from the standby, and the new node
+# would come up verifying checksums its files do not have.
+
+# Make sure the standby has written pages under its own "off" state, so
+# its files really do lack checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,50000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', "SELECT count(*) FROM t2;");
+
+$standby->backup('from_standby');
+my $newnode = PostgreSQL::Test::Cluster->new('newnode');
+$newnode->init_from_backup($standby, 'from_standby');
+
+# Stream from the primary, not from the standby the files were copied
+# from, so the new node replays the "on" primary's records.
+$newnode->enable_streaming($primary);
+$newnode->start;
+
+# The new node's files all came from a cluster running with checksums
+# off. Anything but "off" here means it adopted the primary's state
+# through the checkpoint record at the redo point.
+my ($rc, $stdout, $stderr) = $newnode->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'the node copied from the standby accepts connections')
+ or diag("stderr: $stderr");
+is($stdout, 'off', 'backup of an "off" standby comes up with checksums off');
+
+# A checkpoint record streamed from the "on" primary must not flip the
+# state either.
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($newnode);
+test_checksum_state($newnode, 'off');
+
+# And it must be able to read the pages the standby wrote without
+# checksums.
+(undef, $stdout, $stderr) =
+ $newnode->psql('postgres', "SELECT count(*) FROM t2;");
+is($stdout, '50000', 'pages copied from the standby are readable')
+ or diag("stderr: $stderr");
+
+my $newnode_log = PostgreSQL::Test::Utils::slurp_file($newnode->logfile);
+unlike(
+ $newnode_log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the node copied from the standby');
+
+$newnode->stop('immediate');
+
+# Converge the cluster: enable offline on the standby too.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+$primary->wait_for_catchup($standby);
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after converging');
+
+# Scenario 2: disable offline on the standby only. The replayed
+# checkpoint records of the still-enabled primary must not override it.
+$standby->stop;
+$standby->checksum_disable_offline;
+$standby->start;
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'off');
+
+# Restartpoints must persist the local state, not the replayed copy.
+$standby->safe_psql('postgres', "CHECKPOINT;");
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable with checksums disabled locally');
+
+# Scenario 3: crash-restart right after an online transition, before the
+# next restartpoint. Replay then resumes from an older restartpoint whose
+# checkpoint records still carry the pre-transition state. Those must
+# still match the state seeded from the control file, and the transition
+# itself must be re-established by re-replaying the XLOG2_CHECKSUMS record,
+# without a spurious mismatch warning along the way.
+
+# Converge first: bring the standby back to "on" offline.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+disable_data_checksums($primary, wait => 'off');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+# Crash-restart the standby immediately, before any restartpoint has had a
+# chance to persist the new state to its control file.
+$logstart = -s $standby->logfile;
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001',
+ 'standby readable after crash-restart across an online transition');
+
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no spurious mismatch warning after crash-restart across an online transition'
+);
+
+# Scenario 4: a standby stopped while replaying an interrupted online
+# transition keeps the interrupted state in its own control file. A
+# primary is never caught this way, as its checksums launcher resolves
+# inprogress-on back to off from its exit cleanup. A standby has no
+# launcher; it carries forward whatever the last replayed record left
+# it in.
+
+# Block an online enable on the primary at inprogress-on with a
+# blocking temp table, same trick as in 004_offline.pl.
+my $bsession = $primary->background_psql('postgres');
+$bsession->query_safe('CREATE TEMPORARY TABLE tt (a integer);');
+enable_data_checksums($primary, wait => 'inprogress-on');
+
+# The standby picks up the in-progress state from the XLOG2_CHECKSUMS record.
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# Stop the standby cleanly; its restartpoint persists inprogress-on to
+# its own control file, since nothing on a standby resolves it away.
+$standby->stop;
+$standby->start;
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# The primary still sits at inprogress-on, so a base backup taken now
+# copies a control file and a redo point that both carry the
+# in-progress state; replay of the XLOG2_CHECKSUMS records completes
+# the transition on the new standby.
+$primary->backup('inprogress_backup');
+my $standby2 = PostgreSQL::Test::Cluster->new('standby2');
+$standby2->init_from_backup($primary, 'inprogress_backup',
+ has_streaming => 1);
+$standby2->start;
+
+# Backup label recovery adopts the redo point's state immediately,
+# before any further WAL is replayed.
+wait_for_checksum_state($standby2, 'inprogress-on');
+
+$bsession->quit;
+wait_for_checksum_state($primary, 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable once the transition completes');
+
+# Scenario 4 continued: the backup taken mid-transition must also
+# complete the transition and stay readable.
+$primary->wait_for_catchup($standby2);
+wait_for_checksum_state($standby2, 'on');
+
+is($standby2->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'mid-transition backup standby readable after the transition');
+
+$standby2->stop;
+
+# Nothing along this standby's replay can legitimately disagree: it
+# started out at the same inprogress-on state as the primary.
+my $standby2_log = PostgreSQL::Test::Utils::slurp_file($standby2->logfile);
+unlike(
+ $standby2_log,
+ qr/does not match the state/,
+ 'no mismatch warning on a backup taken mid-transition');
+
+# No extra CHECKPOINT needed: the transition's completion checkpoint
+# wrote every page with a checksum, and the shutdown restartpoint
+# flushes the rest.
+command_ok([ 'pg_checksums', '--check', '-D', $standby2->data_dir ],
+ 'checksums valid on the mid-transition backup standby');
+
+$standby->stop;
+$primary->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/013_rewind.pl b/src/test/modules/test_checksums/t/013_rewind.pl
new file mode 100644
index 00000000000..a791e24317d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/013_rewind.pl
@@ -0,0 +1,201 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test pg_rewind across an online data checksum enable.
+#
+# A clean switchover leaves the shutdown checkpoint of the old primary
+# as the last common checkpoint between the two nodes. When data
+# checksums are enabled online on the new primary before the old one is
+# rewound, pg_rewind installs the control file of the new primary, which
+# already claims checksums are fully enabled, while replay begins at the
+# shutdown checkpoint whose record still carries the old state. Replay
+# of the WAL stretch from before the enable must run with checksums off,
+# as recorded in the checkpoint, else it would verify pages which never
+# had checksums written and fail recovery.
+#
+# The new primary requests a checkpoint right after promotion, and
+# replaying its CHECKPOINT_REDO record would repair the state before
+# any interesting WAL is reached. Hold the checkpointer on the standby
+# in a restartpoint over the promotion, like in the recovery test
+# 041_checkpoint_at_promote, so that the post-promotion writes end up
+# in WAL before the first checkpoint of the new timeline.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Old primary. full_page_writes is off so that the updates done on the
+# promoted node do not carry full page images, forcing replay on the
+# rewound node to read the pages from disk. wal_log_hints is required
+# by pg_rewind on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+full_page_writes = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+log_checkpoints = on
+]);
+$node_a->start;
+
+if (!$node_a->check_extension('injection_points'))
+{
+ plan skip_all => 'Extension injection_points not installed';
+}
+
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+$node_a->safe_psql('postgres', "CREATE EXTENSION injection_points;");
+
+# Set the hint bits on t before taking the backup, so that reads on the
+# promoted node do not emit full page images for its pages later.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->safe_psql('postgres', "SELECT count(*) FROM t;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_a, 'off');
+test_checksum_state($node_b, 'off');
+
+# Hold the next restartpoint on the standby.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point', 'wait');");
+
+# Give the restartpoint a checkpoint record to work on, then start it
+# in a background session; it will block on the injection point with
+# the checkpointer busy until released.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+my $bg_psql = $node_b->background_psql('postgres', on_error_stop => 0);
+$bg_psql->query_until(
+ qr/starting_restartpoint/, q(
+ \echo starting_restartpoint
+ CHECKPOINT;
+));
+$node_b->wait_for_event('checkpointer', 'create-restart-point');
+
+# Clean switchover: the shutdown checkpoint of A streams to B and
+# becomes the last common checkpoint.
+$node_a->stop('fast');
+
+my ($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+$stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+my $shutdown_ckpt = $1;
+$node_b->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_ckpt'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+
+my $logstart = -s $node_b->logfile;
+$node_b->promote;
+
+# Accidental restart of the old primary, diverging its timeline.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->stop('fast');
+
+# Updates on the new primary before its first checkpoint; replay of
+# these on the rewound node has to read the pages from disk.
+$node_b->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 25 = 0;");
+ok( !$node_b->log_contains("checkpoint complete", $logstart),
+ "no checkpoint on the new timeline before the updates");
+
+# Release the checkpointer; the queued post-promotion checkpoint runs
+# after the updates.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+$bg_psql->quit;
+
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+# pg_rewind refuses to run with full_page_writes disabled on the
+# source; the updates it was disabled for are already in WAL.
+$node_b->safe_psql('postgres', "ALTER SYSTEM SET full_page_writes = on;");
+$node_b->safe_psql('postgres', "SELECT pg_reload_conf();");
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the new primary');
+
+# Replay on the rewound node must start at the shutdown checkpoint of
+# the switchover, with the control file of the new primary.
+my $backup_label = slurp_file($node_a->data_dir . '/backup_label');
+$backup_label =~ /^CHECKPOINT LOCATION: ([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from backup_label";
+is($1, $shutdown_ckpt, 'replay starts at the switchover checkpoint');
+
+($stdout, $stderr) = run_command(
+ [
+ 'pg_waldump',
+ '-p' => $node_a->data_dir . '/pg_wal',
+ '-t' => 1,
+ '-s' => $shutdown_ckpt,
+ '-n' => 1,
+ ]);
+like($stdout, qr/CHECKPOINT_SHUTDOWN/,
+ 'last common checkpoint is a shutdown checkpoint');
+
+# pg_rewind keeps the target's own checksum state in the control file it
+# writes; the online enable reaches the rewound node through WAL replay
+# below, not through the copied control file.
+($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+like(
+ $stdout,
+ qr/^Data page checksum version:\s*0$/m,
+ 'rewound node keeps its own checksum state in the control file');
+
+# Start the rewound node as a standby of the new primary. Replay runs
+# through the pre-enable WAL stretch and the online enable.
+#
+# The rewind replaced the configuration files with those of the new
+# primary, so put the port back.
+my $connstr = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_b->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/014_lockstep.pl b/src/test/modules/test_checksums/t/014_lockstep.pl
new file mode 100644
index 00000000000..713f37fb22c
--- /dev/null
+++ b/src/test/modules/test_checksums/t/014_lockstep.pl
@@ -0,0 +1,182 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# The lockstep procedure for offline checksum changes in a replication
+# setup: stop all nodes, run pg_checksums on all of them, restart.
+# Replay of checkpoint records written before the change must not
+# revert the state of the standby.
+#
+# The same pair then runs the procedure in the disable direction, and
+# finally across a divergence: after a failover both nodes get their
+# checksums enabled offline, while the last common checkpoint still
+# carries "off". pg_rewind must keep the offline "on" on the rewound
+# node, since no record in the replayed WAL could ever restore it.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# wal_log_hints keeps the pair eligible for pg_rewind without checksums.
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Part 1: the lockstep procedure in the enable direction. Stop the
+# standby first: the WAL written after this point is replayed only
+# after the offline switch, and every checkpoint record in it still
+# carries the old state.
+$standby->stop;
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 0;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 1;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+# The lockstep procedure.
+$primary->checksum_enable_offline;
+$standby->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to off.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "off" in the replayed WAL/,
+ 0);
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep enable');
+
+# Crash the standby and replay the same stretch again.
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'on');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after crash restart');
+
+# Once a post-switch checkpoint has been replayed the states match and
+# no warning may be logged.
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no mismatch warning once the states match');
+
+# Every page the standby wrote in this window must carry a checksum.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$standby->stop;
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'checksums valid on the standby');
+
+# Part 2: the lockstep procedure in the other direction. The standby
+# is already stopped; give it a pre-switch WAL stretch to replay, with
+# every checkpoint record in it still carrying "on" -- an explicit
+# checkpoint plus the shutdown checkpoint written by stop().
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 2;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+$primary->checksum_disable_offline;
+$standby->checksum_disable_offline;
+
+my $disable_logstart = -s $standby->logfile;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to on.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $disable_logstart);
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep disable');
+
+# Part 3: failover, divergence and pg_rewind across offline enables.
+# The roles swap from here on: the standby becomes the new primary and
+# the old primary is rewound to follow it, but the variables keep the
+# names they had above.
+#
+# Failover: promote the standby, then diverge the old primary.
+$standby->promote;
+$standby->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$primary->stop;
+
+# The lockstep procedure across the divergence: both nodes get their
+# checksums enabled offline.
+$primary->checksum_enable_offline;
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+
+# The states match, so the rewind proceeds without complaint.
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $primary->data_dir,
+ '--source-server' => $standby->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled offline on both nodes');
+
+# The rewind replaced the configuration files with those of the source,
+# so put the port back. The copied file also carries the source's own
+# stale primary_conninfo; enable_streaming() below appends ours, which
+# wins as the later entry.
+$primary->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$primary->port]}
+]);
+$primary->enable_streaming($standby);
+
+my $rewind_logstart = -s $primary->logfile;
+$primary->start;
+$standby->wait_for_catchup($primary);
+
+is($primary->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+
+# The offline enable survives: replay from the common checkpoint must
+# not resurrect the pre-divergence "off".
+test_checksum_state($primary, 'on');
+
+$log =
+ PostgreSQL::Test::Utils::slurp_file($primary->logfile, $rewind_logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no divergence reported between the rewound server and its source');
+
+$primary->stop;
+$standby->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
new file mode 100644
index 00000000000..4ab51c34d8f
--- /dev/null
+++ b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
@@ -0,0 +1,136 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that a standby crashing after a replayed online disable, but before
+# its next restartpoint, restarts cleanly.
+#
+# Pages dirtied before the XLOG2_CHECKSUMS record and evicted after it are
+# written out under the new "off" state, without checksums. The control
+# file must follow the record immediately in this direction: a crash-restart
+# initializes checksum verification from the control file, and one still
+# saying "on" would fail verification on exactly those pages while replaying
+# records older than the state change.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$primary->start;
+
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,1000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+
+# Small buffer pool so that replaying a bulk load evicts the pages we care
+# about, and no restartpoints so the control file stays where it is.
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 1MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_delay = 10000
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'on');
+
+# No full page images, so that replay has to read the pages of "t" back from
+# disk instead of overwriting them from the WAL.
+$primary->append_conf('postgresql.conf', 'full_page_writes = off');
+$primary->reload;
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" on the standby, *before* the state change. They are
+# not flushed: the standby has no restartpoint from here on.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+# Online disable. The standby replays it and switches to "off".
+$primary->safe_psql('postgres', 'SELECT pg_disable_data_checksums();');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+my ($ctl_before) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($ctl_state) = $ctl_before =~ /Data page checksum version:\s+(\d+)/;
+note(
+ "standby control file data_checksum_version after the disable: "
+ . "$ctl_state (0 = off, 1 = on)");
+
+# Force the standby to evict the dirty pages of "t" now that it is "off":
+# they are written back without checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE filler AS SELECT generate_series(1,300000) AS a;");
+$primary->wait_for_catchup($standby);
+
+# Crash the standby before it gets a chance to run a restartpoint.
+$standby->stop('immediate');
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'standby restarts after crashing between a checksum state '
+ . 'change and the next restartpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'the pages written while "off" are readable')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("standby log tail:\n" . join("\n", @lines[ -12 .. -1 ]))
+ if @lines >= 12;
+ fail('the pages written while "off" are readable');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/016_promote_enable_crash.pl b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
new file mode 100644
index 00000000000..13c983136ef
--- /dev/null
+++ b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
@@ -0,0 +1,146 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a standby promoted right after replaying an online enable, before any
+# restartpoint has flushed the rewritten pages.
+#
+# The end-of-recovery record persists the checksum state without flushing the
+# buffer pool, while the control file still points at a restartpoint older than
+# the transition. Persisting "on" there would make a crash before the
+# post-promotion checkpoint verify checksums over the pages that were flushed
+# while the state was still "off"; promotion must take the full checkpoint
+# instead.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Dirty the pages of "t" without full page images, then push them out to disk
+# on the standby while checksums are still off.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Online enable on the primary; the standby replays it into its own buffers,
+# where the rewritten pages stay dirty (no restartpoint, no bgwriter).
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+note("standby control file before the promotion: $before");
+
+# Promote. The end-of-recovery record persists the live state without any
+# buffer flush; the checkpoint requested afterwards is not immediate.
+$standby->promote;
+$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery();')
+ or die 'timed out waiting for the promotion';
+$standby->stop('immediate');
+
+($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($ckpt) = $ctl =~ /Latest checkpoint location:\s+(\S+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note(
+ "standby control file after the promotion: $after, checkpoint $ckpt, redo $redo"
+);
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'promoted node restarts after crashing before its first checkpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/017_restartpoint_race.pl b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
new file mode 100644
index 00000000000..10919c9741c
--- /dev/null
+++ b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
@@ -0,0 +1,151 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a restartpoint whose flush races the replay of an online enable.
+#
+# Replay keeps running while CheckPointGuts() writes out the buffer pool, so
+# the state at the end of the flush can be newer than the one the flush ran
+# under. Persisting it would claim checksums for pages the flush wrote out
+# before the transition, against a redo pointer that predates it.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# C1: the checkpoint the pending restartpoint will target.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" without full page images and push them to disk on
+# the standby while it is still "off".
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Start a restartpoint for C1 and hold it right after CheckPointGuts().
+$standby->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point','wait');");
+my $bg = $standby->background_psql('postgres');
+$bg->query_until(qr//, "\\echo restartpoint\nCHECKPOINT;\n");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint injection point';
+
+# The startup process keeps replaying while the checkpointer is held: the
+# whole online enable lands, and the rewritten pages stay dirty.
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+# Let the restartpoint finish; it persists the state it samples now.
+$standby->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint to finish';
+$standby->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+
+$standby->stop('immediate');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note("standby control file after the restartpoint: $after, redo $redo");
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started, 'standby restarts after a restartpoint that raced the enable');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/018_enable_crash_windows.pl b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
new file mode 100644
index 00000000000..4dac590e819
--- /dev/null
+++ b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
@@ -0,0 +1,508 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Crash and checkpoint windows inside an online data checksum enable on
+# a single node. Each scenario starts from "off" on the same cluster,
+# runs the enable into a held injection point, and checks what a crash
+# or a concurrent checkpoint in that window may and may not do to the
+# control file and the pages on disk.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+use IPC::Run;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum_extended' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum_extended\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Bring the node back to "off" with no leftover table between
+# scenarios. The crash restarts already resolve an interrupted enable
+# to "off", so only disable when something is left to disable. Those
+# restarts also clear any attached injection points. The checkpoint
+# bounds the WAL the next scenario's crash recovery has to replay.
+sub reset_scenario
+{
+ my ($node) = @_;
+
+ BAIL_OUT('node is not running after the previous scenario')
+ unless $node->is_alive;
+
+ my $state = $node->safe_psql('postgres', 'SHOW data_checksums;');
+ disable_data_checksums($node, wait => 'off') if $state ne 'off';
+ $node->safe_psql('postgres', 'DROP TABLE IF EXISTS t;');
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+ test_checksum_state($node, 'off');
+}
+
+my $node = PostgreSQL::Test::Cluster->new('crash_windows');
+$node->init(no_data_checksums => 1);
+
+# Scenario 1 needs the pages of "t" to reach disk without a full page
+# image, hence full_page_writes and wal_log_hints off. Scenario 2
+# needs them to stay dirty in shared buffers until a checkpoint, hence
+# no bgwriter and a large shared_buffers. wal_level is pinned because
+# init() writes "minimal" for a non-streaming node. The rest merely
+# tolerate the settings.
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+wal_level = replica
+max_wal_size = 10GB
+full_page_writes = off
+wal_log_hints = off
+shared_buffers = 512MB
+bgwriter_lru_maxpages = 0
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$node->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+
+test_checksum_state($node, 'off');
+
+# Scenario 1: a primary crashes inside SetDataChecksumsOn(), just before the
+# forced checkpoint that flushes the rewritten pages, with crash recovery
+# resuming from a checkpoint older than the transition. The control file
+# must still say "inprogress-on" there: replay does not adopt the state of
+# the checkpoint record it resumes from, so an "on" written before the flush
+# would stay in effect while replay reads pages whose rewrite never reached
+# disk. Here the pages went out through the rewriting worker's ring buffer,
+# so only the control file state discriminates; see the next scenario for
+# the case where they did not.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Establish the checkpoint that crash recovery will resume from.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty the pages of "t" without full page images, then push them out
+ # to disk while checksums are still off.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+ my $evicted = $node->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+ note("evict_relation on t: $evicted, relpath $relpath");
+
+ # Stop the enable right after the control file has been updated to
+ # "on" but before the checkpoint that flushes the rewritten pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+ $node->safe_psql('postgres', 'SELECT pg_enable_data_checksums();');
+
+ $node->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'datachecksums-on-before-checkpoint';"
+ ) or die 'timed out waiting for the injection point';
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ $node->stop('immediate');
+
+ # Show the on-disk checksum field of the first page of "t".
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 2: the same crash as above, but with the rewritten pages resident
+# in shared buffers instead of gone out through the ring. The rewriting
+# worker reads through a BAS_VACUUM ring, which writes pages back as the
+# ring recycles, but a page already resident in shared buffers is not read
+# through the ring: ReadBufferExtended() hands back the existing buffer, and
+# it stays dirty until a checkpoint. The control file may therefore not
+# say "on" before that checkpoint has run, or crash recovery would resume
+# from a checkpoint older than the transition with verification already
+# enabled, and replay records that read those still-unchecksummed pages.
+# full_page_writes is off so that the records do not simply overwrite the
+# pages with a full page image.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,100000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Establish the checkpoint that crash recovery will resume from, with the
+ # pages of "t" written out while checksums are still off.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty those pages again without emitting full page images, and leave
+ # them in shared buffers. Replay of these records has to read the
+ # pages from disk.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+
+ my $dirty_before = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_before, '>', 0,
+ 'pages of t are resident and dirty before enabling checksums');
+
+ # Hold the enabling right before the checkpoint that flushes the rewritten
+ # pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ # The rewritten pages must still be sitting dirty in shared buffers, or
+ # the window this test is about does not exist.
+ my $dirty_after = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_after, '>', 0,
+ 'rewritten pages of t are still dirty before the checkpoint');
+
+ $node->stop('immediate');
+
+ # The on-disk copy is the one written before the transition, without a
+ # checksum.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ is($pd_checksum, 0, 'block 0 of t on disk carries no checksum');
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 3: a checkpoint that runs between the XLOG2_CHECKSUMS("on")
+# record and the control file write at the end of SetDataChecksumsOn() must
+# not make a crash throw the completed transition away. SetDataChecksumsOn()
+# writes the record, flips shared memory to "on", emits the barrier and only
+# then requests the checkpoint that flushes the rewritten pages; the control
+# file is written after that checkpoint returns. A crash in that window is
+# harmless only as long as recovery still starts before the record. Any
+# checkpoint completing in the window moves the redo point past the record,
+# so recovery would never see it, would come up with the control file's
+# "inprogress-on" and StartupXLOG() would demote that to "off", even though
+# every page on disk carries a checksum by then. CreateCheckPoint()
+# therefore persists the state the checkpoint ran under, the same way
+# CreateRestartPoint() does. Here the window is made deterministic by
+# holding the launcher at the datachecksums-on-before-checkpoint injection
+# point and checkpointing from another session.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Hold the launcher after the record, the shared memory flip and the
+ # barrier, but before the checkpoint SetDataChecksumsOn() requests
+ # itself.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ # Every backend already sees "on" and writes checksums.
+ test_checksum_state($node, 'on');
+
+ # ... while the control file still says "inprogress-on".
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($before, '3', 'control file says "inprogress-on" inside the window');
+
+ # A concurrent checkpoint. It flushes the rewritten pages and moves
+ # the redo point past the XLOG2_CHECKSUMS("on") record, so it has to
+ # record the "on" state in the control file as well.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($after, '1',
+ 'the concurrent checkpoint records "on" in the control file');
+
+ $node->stop('immediate');
+
+ # The checkpoint flushed the rewritten pages, so they carry a checksum on
+ # disk: the transition really did complete.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ isnt($pd_checksum, 0, 'block 0 of t on disk carries a checksum');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition is complete on disk, so the cluster has to come back
+ # "on".
+ test_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+}
+reset_scenario($node);
+
+# Scenario 4: a crash between the checkpoint an online enable requests and
+# the control file write that follows it must not lose the transition. The
+# last steps of SetDataChecksumsOn() are
+#
+# WAL record -> shmem -> barrier -> checkpoint -> persist
+#
+# The checkpoint is what makes "on" safe to persist, but it also moves the
+# redo point above the XLOG2_CHECKSUMS record that carries the new state.
+# Crash recovery started from that checkpoint therefore never replays the
+# record, so without the state the checkpoint itself persists, a crash in
+# the remaining window would bring the cluster back at "inprogress-on",
+# which StartupXLOG() resolves to "off", discarding a transition whose
+# pages are all on disk with a checksum. The previous scenario exercises
+# the same window through a checkpoint requested by another session; this
+# one closes it from the other side: the transition's own checkpoint is the
+# one that moves the redo point, so persisting the state from
+# CreateCheckPoint() is what has to save it, not the write below the
+# injection point.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Hold the launcher after the checkpoint that licenses "on" has
+ # completed and before the state reaches the control file.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-after-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-after-checkpoint');
+
+ # The transition is complete as far as the running cluster is concerned.
+ test_checksum_state($node, 'on');
+
+ # The checkpoint has already recorded it, so the pending write below the
+ # injection point has nothing left to do.
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($version) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($version, '1',
+ 'the requested checkpoint recorded "on" in the control file');
+
+ # Crash before the control file write that follows the checkpoint.
+ $node->stop('immediate');
+
+ $node->start;
+
+ # Every page on disk carries a checksum and the checkpoint that
+ # flushed them completed, so the cluster has to come back verifying
+ # them.
+ test_checksum_state($node, 'on');
+
+ is($node->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '10000', 'relation readable after the crash');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ # The data directory must be one the offline tools accept.
+ $node->stop;
+ $node->command_ok([ 'pg_checksums', '--check', '-D', $node->data_dir ],
+ 'pg_checksums accepts the data directory');
+
+ # Leave the node running for the reset.
+ $node->start;
+}
+reset_scenario($node);
+
+# Scenario 5: a checkpoint racing SetDataChecksumsOn() between the insertion
+# of the XLOG2_CHECKSUMS("on") record and the shared memory update must not
+# insert an XLOG_CHECKPOINT_REDO record that follows the transition in WAL
+# order while still carrying "inprogress-on". Recovery resuming from such a
+# redo point never replays the preceding transition record, comes up in
+# "inprogress-on" and resolves the finished transition as interrupted.
+# XLogChecksums() closes the window by inserting the record and publishing
+# the new state under DataChecksumTransitionLock, which CreateCheckPoint()
+# takes around sampling the state and inserting the redo record. Here the
+# launcher is held between the two steps at the
+# datachecksums-on-before-publish injection point, and a concurrent
+# CHECKPOINT has to block on the lock instead of completing inside the
+# window.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # The datachecksums-on-before-publish point fires inside a critical
+ # section, where the wait machinery must not allocate. Waiting once
+ # at the datachecksums-enable-checksums-delay point, which the
+ # launcher runs outside the critical section, initializes it; see
+ # 050_redo_segment_missing.pl for the same recipe around
+ # create-checkpoint-run.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-enable-checksums-delay','wait');"
+ );
+
+ # Hold the launcher after the XLOG2_CHECKSUMS("on") record is in WAL but
+ # before the new state is published in shared memory.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-publish','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-enable-checksums-delay');
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-enable-checksums-delay');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-enable-checksums-delay');");
+
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-publish');
+
+ # The record is in WAL, the published state is still the old one.
+ test_checksum_state($node, 'inprogress-on');
+
+ # A concurrent checkpoint. It must block on DataChecksumTransitionLock
+ # before inserting its redo record rather than complete inside the window.
+ my $checkpointer = IPC::Run::start(
+ [ 'psql', '-XAtq', '-d', $node->connstr('postgres'), '-c', 'CHECKPOINT;' ],
+ '<' => '/dev/null',
+ '>' => '/dev/null',
+ '2>' => '/dev/null',
+ IPC::Run::timer($PostgreSQL::Test::Utils::timeout_default));
+
+ ok( $node->poll_query_until(
+ 'postgres',
+ "SELECT wait_event = 'DataChecksumTransition' "
+ . "FROM pg_stat_activity WHERE backend_type = 'checkpointer';"),
+ 'concurrent checkpoint blocks on DataChecksumTransitionLock');
+
+ # Release the launcher; the checkpoint then samples the published "on" and
+ # its redo record follows the transition record.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-publish');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-publish');");
+
+ $checkpointer->finish;
+
+ # Crash while the launcher's own checkpoint may still be in flight.
+ # Recovery resumes from the concurrent checkpoint's redo point, which
+ # now lies above the transition record and carries "on".
+ $node->stop('immediate');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition completed, so the cluster has to come back "on".
+ wait_for_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ $node->stop;
+}
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
new file mode 100644
index 00000000000..abcc234ab22
--- /dev/null
+++ b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
@@ -0,0 +1,138 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A standby stopped after replaying an online enable, but before the
+# checkpoint record that follows it on the primary.
+#
+# The shutdown restartpoint has no new checkpoint record to work from
+# and is skipped, so the control file would keep the in-progress state
+# the transition passed through, even though replay left the node at
+# "on" and rewrote every page. pg_checksums reads that field and
+# refuses to run on an in-progress state, so the shutdown has to flush
+# and catch it up instead.
+#
+# The caught-up control file then carries the watermark of the "on"
+# record, which the second half of this test relies on: an offline
+# disable made while the standby is down must survive the re-replay of
+# that record on the next startup, or the change would be silently
+# reverted.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+wal_keep_size = 1GB
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Anchor the standby's last restartpoint here, so that nothing replayed
+# from now on gives the shutdown restartpoint a newer checkpoint record
+# to use.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Hold the enable right after the state change record, before the
+# checkpoint it requests once the transition is complete.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+enable_data_checksums($primary);
+$primary->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+wait_for_checksum_state($primary, 'on');
+
+# The state change record is already flushed and streams on its own;
+# no checkpoint record follows while the launcher is held.
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+$standby->stop('fast');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($state, '1',
+ 'shutdown catches the control file up with the replayed state');
+
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'pg_checksums verifies the stopped standby');
+
+# Disable checksums offline while the standby is down. The restart
+# below resumes replay under the last restartpoint, re-reading the "on"
+# record; the watermark persisted with the catch-up above marks it as
+# already applied, so the offline change must survive.
+$standby->checksum_disable_offline;
+
+# Let the enable finish on the primary.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-checkpoint');");
+$primary->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-checkpoint');");
+
+# The launcher exits only after the checkpoint it requests completes,
+# so a checkpoint record now follows the transition in WAL.
+$primary->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_catalog.pg_stat_activity "
+ . "WHERE backend_type = 'datachecksums launcher';");
+
+my $log_offset = -s $standby->logfile;
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+
+# The nodes legitimately diverged, which replayed checkpoints report.
+$standby->wait_for_log(
+ qr/data checksum state "off" of this node does not match the state "on" in the replayed WAL/,
+ $log_offset);
+
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/020_cascade_divergence.pl b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
new file mode 100644
index 00000000000..8f2816ac405
--- /dev/null
+++ b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
@@ -0,0 +1,127 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# An offline data checksum state change made on one node of a cascading setup
+# is reported all the way down the chain.
+#
+# CheckReplayedDataChecksumState() is reached from xlog_redo() when a
+# checkpoint record (XLOG_CHECKPOINT_SHUTDOWN or XLOG_CHECKPOINT_REDO) is
+# replayed after consistency has been reached. Those records are written by
+# the root primary only and are relayed verbatim by every intermediate
+# standby, so a cascaded standby that was changed offline learns about the
+# divergence too. An intermediate standby's own restartpoints are not
+# WAL-logged and cannot surface it.
+#
+# Note that this makes detection checkpoint-driven: an idle or read-only
+# primary surfaces the divergence no sooner than its next checkpoint, so up to
+# checkpoint_timeout may pass before the operator sees anything. Closing that
+# window would require comparing the state when a walreceiver connects.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+# checkpoint_timeout is deliberately long so that the only checkpoint record in
+# this test is the one requested explicitly below.
+my $conf = qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+);
+
+my $primary = PostgreSQL::Test::Cluster->new('cascade_primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', $conf);
+$primary->start;
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,1000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('cascade_standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf('postgresql.conf', $conf);
+$standby->start;
+
+$standby->backup('backup');
+my $cascade = PostgreSQL::Test::Cluster->new('cascade_cascade');
+$cascade->init_from_backup($standby, 'backup', has_streaming => 1);
+$cascade->append_conf('postgresql.conf', $conf);
+$cascade->start;
+
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+test_checksum_state($cascade, 'off');
+
+# Change the cascaded standby offline, and only it.
+$cascade->stop;
+$cascade->command_ok([ 'pg_checksums', '--enable', '-D', $cascade->data_dir ],
+ 'pg_checksums enables checksums on the cascaded standby only');
+
+my $logstart = -s $cascade->logfile;
+$cascade->start;
+
+test_checksum_state($cascade, 'on');
+
+# Ordinary WAL traffic carries no state, so the cascaded standby keeps
+# streaming from a chain whose state is "off" without noticing.
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (1);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+is($cascade->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '1001', 'cascaded standby keeps streaming from a divergent chain');
+
+# Restartpoints on the intermediate standby are not WAL-logged, so this does
+# not surface anything on the cascaded standby either.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (2);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+my $log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+unlike(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'an intermediate restartpoint does not surface the divergence');
+
+# A checkpoint on the root primary does: the record is relayed down the whole
+# chain, so both the direct and the cascaded standby compare it against their
+# own state. wait_for_catchup() below waits for replay, not just receipt, so
+# the record has been through xlog_redo() on both nodes once it returns.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+$log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+like(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'the cascaded standby reports the divergence on the primary checkpoint');
+
+$cascade->stop;
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
new file mode 100644
index 00000000000..3f660623bd3
--- /dev/null
+++ b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
@@ -0,0 +1,205 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind across online data checksum transitions on both sides of a
+# divergence, with the two nodes trading roles between the scenarios.
+#
+# Scenario 1: after a switchover the new primary enables checksums
+# online, while the old primary restarts on its old timeline, advances
+# its WAL beyond the enable records and runs an online enable/disable
+# cycle of its own. The old primary ends "off" with a checksum
+# watermark numerically above every checksum record the new primary has
+# written. pg_rewind clamps the watermark it keeps to the divergence
+# point; without the clamp, replay on the rewound node would skip the
+# source's enable as already applied and stay "off" under an "on"
+# primary.
+#
+# Scenario 2: checksums are disabled again, and after another
+# switchover both nodes enable them online independently, so the
+# divergence checkpoint carries "off" while both control files say
+# "on". The rewind is allowed, the target keeps its own "on" state,
+# and replay re-walks the source's enable onto the already enabled
+# node.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum_extended' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum_extended\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+sub controldata_watermark
+{
+ my ($node) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $node->data_dir ]);
+ $stdout =~ /^Data checksum watermark:\s*([0-9A-F]+)\/([0-9A-F]+)$/m
+ or die "watermark missing from pg_controldata output";
+ return (hex($1) << 32) + hex($2);
+}
+
+# Wait until the standby has replayed the shutdown checkpoint of the
+# stopped primary, so that a subsequent promotion diverges after it and
+# the shutdown checkpoint becomes the last common checkpoint.
+sub wait_for_shutdown_checkpoint_replay
+{
+ my ($primary, $standby) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $primary->data_dir ]);
+ $stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+ my $shutdown_checkpoint = $1;
+
+ $standby->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_checkpoint'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+}
+
+# Old primary, checksums off. wal_log_hints is required by pg_rewind
+# on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+# Clean switchover to B; enable checksums online on it.
+$node_a->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_a, $node_b);
+$node_b->promote;
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+$node_b->stop('fast');
+my $watermark_b = controldata_watermark($node_b);
+$node_b->start;
+
+# Accidental restart of the old primary on the old timeline. Advance
+# its WAL beyond the enable watermark of B, then run an online enable
+# and disable cycle: the node ends "off" with a watermark above every
+# checksum record B has written.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t_pad AS SELECT generate_series(1,200000) AS a;");
+enable_data_checksums($node_a, wait => 'on');
+disable_data_checksums($node_a, wait => 'off');
+test_checksum_state($node_a, 'off');
+$node_a->stop('fast');
+
+my $watermark_a = controldata_watermark($node_a);
+die "test broken: target watermark not above the source's enable"
+ unless $watermark_a > $watermark_b;
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the promoted node');
+
+# Start the rewound node as a standby of B. Replay from the last
+# common checkpoint runs through B's online enable, which the target
+# never saw, so the rewound node must converge to "on".
+my $connstr_b = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr_b application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+# Scenario 2, reusing the pair with the roles reversed. Disable
+# checksums online so the next divergence point carries "off", and let
+# A replay the change.
+disable_data_checksums($node_b, wait => 'off');
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'off');
+
+# Clean switchover back to A; enable checksums online on it.
+$node_b->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_b, $node_a);
+$node_a->promote;
+enable_data_checksums($node_a, wait => 'on');
+test_checksum_state($node_a, 'on');
+my $source_enable_watermark = controldata_watermark($node_a);
+
+# The old primary restarts on its old timeline and enables checksums
+# online independently: both control files say "on", the divergence
+# checkpoint says "off".
+$node_b->start;
+$node_b->safe_psql('postgres', "INSERT INTO t_div VALUES (2);");
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+$node_b->stop('fast');
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_b->data_dir,
+ '--source-server' => $node_a->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled on both nodes');
+
+my $connstr_a = $node_a->connstr;
+$node_b->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_b->port]}
+primary_conninfo = '$connstr_a application_name=@{[$node_b->name]}'
+]);
+$node_b->set_standby_mode;
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_b, 'on');
+
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_b->stop('fast');
+is(controldata_watermark($node_b), $source_enable_watermark,
+ 'rewound node replayed the source checksum transition');
+command_ok([ 'pg_checksums', '--check', '-D', $node_b->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the twice-rewound node');
+
+done_testing();
--
2.39.3 (Apple Git-146)
=
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-10 13:18 Daniel Gustafsson <daniel@yesql.se>
parent: Daniel Gustafsson <daniel@yesql.se>
0 siblings, 1 reply; 43+ messages in thread
From: Daniel Gustafsson @ 2026-09-10 13:18 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-hackers@lists.postgresql.org
> On 10 Sep 2026, at 12:05, Daniel Gustafsson <daniel@yesql.se> wrote:
>
> The attached v15 ..
..and in case anyone need more evidence that doing things when stressed is bad, I
accidentally attached the wrong version.
--
Daniel Gustafsson
Attachments:
[application/octet-stream] v16-0009-Only-log-page-zeroing-when-checksums-aren-t-igno.patch (1.4K, ../../0E4734FA-4764-405C-9670-C45DF440A718@yesql.se/2-v16-0009-Only-log-page-zeroing-when-checksums-aren-t-igno.patch)
download | inline diff:
From 63a6951217fd1fcb4210c5fb77e5bfc0597d8195 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 10 Sep 2026 10:39:36 +0200
Subject: [PATCH v16 9/9] Only log page zeroing when checksums aren't ignored
When checksum failures are ignored, the page won't be zeroed on error
even if zero_damaged_pages is set. Found via review by GPT-6.
Author: Daniel Gustafsson <daniel@yesql.se>
Reviewed-by: tbd
Discussion: https://postgr.es/m/..
Backpatch-through: 19
---
src/backend/storage/page/bufpage.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/src/backend/storage/page/bufpage.c b/src/backend/storage/page/bufpage.c
index 8df74618fcd..a6d6e13f0b3 100644
--- a/src/backend/storage/page/bufpage.c
+++ b/src/backend/storage/page/bufpage.c
@@ -160,7 +160,7 @@ PageIsVerified(PageData *page, BlockNumber blkno, int flags, bool *checksum_fail
if ((flags & (PIV_LOG_WARNING | PIV_LOG_LOG)) != 0)
ereport(flags & PIV_LOG_WARNING ? WARNING : LOG,
(errcode(ERRCODE_DATA_CORRUPTED),
- (flags & PIV_ZERO_BUFFERS_ON_ERROR) ?
+ ((flags & PIV_ZERO_BUFFERS_ON_ERROR) && !(flags & PIV_IGNORE_CHECKSUM_FAILURE)) ?
errmsg("page verification failed, calculated checksum %u but expected %u, buffer will be zeroed",
checksum, p->pd_checksum) :
errmsg("page verification failed, calculated checksum %u but expected %u",
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v16-0008-Only-reset-launcher-state-in-exit-handler.patch (1.3K, ../../0E4734FA-4764-405C-9670-C45DF440A718@yesql.se/3-v16-0008-Only-reset-launcher-state-in-exit-handler.patch)
download | inline diff:
From 2cb38d553462ad515320821f09392ab9bc90e23e Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Tue, 8 Sep 2026 22:42:21 +0200
Subject: [PATCH v16 8/9] Only reset launcher state in exit handler
The launcher_running flag was reset to false during shutdown, which
left a small window where a new launcher could be started when the
exit handler was still running. Fix by only updating the running
flag during the exit handler.
Found by Noah using AI assisted review with Claude. Backpatch to
v19 where the online checksums feature was introduced.
Reported-by: Noah Misch <noah@leadboat.com>
Discussion: https://postgr.es/m/...
Backpatch-through: 19
---
src/backend/postmaster/datachecksum_state.c | 2 --
1 file changed, 2 deletions(-)
diff --git a/src/backend/postmaster/datachecksum_state.c b/src/backend/postmaster/datachecksum_state.c
index 099a6b4fe2e..86e378d8951 100644
--- a/src/backend/postmaster/datachecksum_state.c
+++ b/src/backend/postmaster/datachecksum_state.c
@@ -1395,8 +1395,6 @@ done:
/* Shut down progress reporting as we are done */
pgstat_progress_end_command();
- launcher_running = false;
- DataChecksumState->launcher_running = false;
LWLockRelease(DataChecksumsWorkerLock);
}
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v16-0007-Improve-error-message-for-checksum-state-in-pg_u.patch (1.6K, ../../0E4734FA-4764-405C-9670-C45DF440A718@yesql.se/4-v16-0007-Improve-error-message-for-checksum-state-in-pg_u.patch)
download | inline diff:
From 3f6504034c95292fba9ae721c857f71a23d1598c Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Mon, 7 Sep 2026 12:22:01 +0200
Subject: [PATCH v16 7/9] Improve error message for checksum state in
pg_upgrade
When attempting to upgrade a cluster which is an inprogress state, the
same error message was used regardless of which state it was. Fix by
using different error messages for the different inprogress states.
Found by Noah using AI assisted review with Claude. Backpatch to v19
where the online checksums feature was introduced.
Reported-by: Noah Misch <noah@leadboat.com>
Discussion: https://postgr.es/m/...
Backpatch-through: 19
---
src/bin/pg_upgrade/controldata.c | 4 +++-
1 file changed, 3 insertions(+), 1 deletion(-)
diff --git a/src/bin/pg_upgrade/controldata.c b/src/bin/pg_upgrade/controldata.c
index 759bd4ef77c..e83f2d963cd 100644
--- a/src/bin/pg_upgrade/controldata.c
+++ b/src/bin/pg_upgrade/controldata.c
@@ -658,8 +658,10 @@ check_control_data(ControlData *oldctrl,
* upgrade. The user should either let the process finish, or turn off
* data checksums, before retrying.
*/
- if (oldctrl->data_checksum_version > PG_DATA_CHECKSUM_VERSION)
+ if (oldctrl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON)
pg_fatal("data checksums are being enabled in the old cluster");
+ if (oldctrl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ pg_fatal("data checksums are being disabled in the old cluster");
/*
* We might eventually allow upgrades from checksum to no-checksum
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v16-0006-Fix-enum-value-visibility-for-data_checksums.patch (1.6K, ../../0E4734FA-4764-405C-9670-C45DF440A718@yesql.se/5-v16-0006-Fix-enum-value-visibility-for-data_checksums.patch)
download | inline diff:
From a5ac10eef98549d9a360f8861afa9b2561d53e1e Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Sat, 5 Sep 2026 20:20:45 +0200
Subject: [PATCH v16 6/9] Fix enum value visibility for data_checksums
When the data_checksums GUC was changed into an enum the enum values
were all incorrectly marked as hidden, which made pg_settings report
an empty array. Fix by setting all as visible as they should be.
Found by Noah using AI assisted review with GPT. Backpatch to v19
where the GUC was changed in the online checksums feature.
Reported-by: Noah Misch <noah@leadboat.com>
Discussion: https://postgr.es/m/...
Backpatch-through: 19
---
src/backend/utils/misc/guc_tables.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
diff --git a/src/backend/utils/misc/guc_tables.c b/src/backend/utils/misc/guc_tables.c
index c6d9b2a6f89..342aaeef59a 100644
--- a/src/backend/utils/misc/guc_tables.c
+++ b/src/backend/utils/misc/guc_tables.c
@@ -513,10 +513,10 @@ static const struct config_enum_entry file_extend_method_options[] = {
};
static const struct config_enum_entry data_checksums_options[] = {
- {"on", PG_DATA_CHECKSUM_VERSION, true},
- {"off", PG_DATA_CHECKSUM_OFF, true},
- {"inprogress-on", PG_DATA_CHECKSUM_INPROGRESS_ON, true},
- {"inprogress-off", PG_DATA_CHECKSUM_INPROGRESS_OFF, true},
+ {"on", PG_DATA_CHECKSUM_VERSION, false},
+ {"off", PG_DATA_CHECKSUM_OFF, false},
+ {"inprogress-on", PG_DATA_CHECKSUM_INPROGRESS_ON, false},
+ {"inprogress-off", PG_DATA_CHECKSUM_INPROGRESS_OFF, false},
{NULL, 0, false}
};
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v16-0005-doc-Documentation-updates-for-online-checksums.patch (21.3K, ../../0E4734FA-4764-405C-9670-C45DF440A718@yesql.se/6-v16-0005-doc-Documentation-updates-for-online-checksums.patch)
download | inline diff:
From 76aea64cdf1d27cce266eddc7419228e563520b3 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Wed, 9 Sep 2026 23:38:11 +0200
Subject: [PATCH v16 5/9] doc: Documentation updates for online checksums
A set of documentation updates for the online checksums work, found via
manual as well as LLM-guided review.
* The state diagram was re-done to be improve readability and be
more in line with the look and feel of other diagrams.
* Paragraph about mismatched states moved to its own sect2.
* The documentation for the functions to enable/ disable checksums
failed to mention that they are superuser only.
* The documentation for the checksums progress reporting had
accidentally omitted the work "fork" and for th blocks_* columns.
The values are per fork but the documentation made it seem they
were per relation.
Reported-by: Heikki Linnakangas <hlinnaka@iki.fi>
Reported-by: Noah Misch <noah@leadboat.com>
Discussion: https://postgr.es/m/5b08b4d1-0982-4b77-bac5-3bdffc6583f5@iki.fi
---
doc/src/sgml/func/func-admin.sgml | 13 +-
doc/src/sgml/images/datachecksums.gv | 61 ++++++++--
doc/src/sgml/images/datachecksums.svg | 163 ++++++++++++++++----------
doc/src/sgml/images/meson.build | 1 +
doc/src/sgml/monitoring.sgml | 4 +-
doc/src/sgml/ref/pg_checksums.sgml | 2 +-
doc/src/sgml/wal.sgml | 60 ++++++----
7 files changed, 203 insertions(+), 101 deletions(-)
diff --git a/doc/src/sgml/func/func-admin.sgml b/doc/src/sgml/func/func-admin.sgml
index 54eeb42e5bc..8bfd290ce15 100644
--- a/doc/src/sgml/func/func-admin.sgml
+++ b/doc/src/sgml/func/func-admin.sgml
@@ -3128,7 +3128,8 @@ SELECT convert_from(pg_read_binary_file('file_in_utf8.txt'), 'UTF8');
<para>
The functions shown in <xref linkend="functions-checksums-table" /> can
- be used to enable or disable data checksums in a running cluster.
+ be used to enable or disable data checksums in a running cluster. Use of
+ these functions is restricted to superusers.
</para>
<para>
Changing data checksums can be done in a cluster with concurrent activity
@@ -3158,7 +3159,7 @@ SELECT convert_from(pg_read_binary_file('file_in_utf8.txt'), 'UTF8');
<indexterm>
<primary>pg_enable_data_checksums</primary>
</indexterm>
- <function>pg_enable_data_checksums</function> ( <optional><parameter>cost_delay</parameter> <type>int</type>, <parameter>cost_limit</parameter> <type>int</type></optional> )
+ <function>pg_enable_data_checksums</function> ( <optional><parameter>cost_delay</parameter> <type>int</type> <optional>, <parameter>cost_limit</parameter> <type>int</type></optional></optional> )
<returnvalue>void</returnvalue>
</para>
<para>
@@ -3174,6 +3175,11 @@ SELECT convert_from(pg_read_binary_file('file_in_utf8.txt'), 'UTF8');
If <parameter>cost_delay</parameter> and <parameter>cost_limit</parameter> are
specified, the process is throttled using the same principles as
<link linkend="runtime-config-resource-vacuum-cost">Cost-based Vacuum Delay</link>.
+ <parameter>cost_delay</parameter> defaults to 0,
+ <parameter>cost_limit</parameter> defaults to 100.
+ </para>
+ <para>
+ This function is restricted to superusers.
</para>
</entry>
</row>
@@ -3193,6 +3199,9 @@ SELECT convert_from(pg_read_binary_file('file_in_utf8.txt'), 'UTF8');
stopped validating data checksums, the data checksum state will be
set to <literal>off</literal>.
</para>
+ <para>
+ This function is restricted to superusers.
+ </para>
</entry>
</row>
</tbody>
diff --git a/doc/src/sgml/images/datachecksums.gv b/doc/src/sgml/images/datachecksums.gv
index dff3ff7340a..f032555db3e 100644
--- a/doc/src/sgml/images/datachecksums.gv
+++ b/doc/src/sgml/images/datachecksums.gv
@@ -1,14 +1,49 @@
-digraph G {
- A -> B [label="SELECT pg_enable_data_checksums()"];
- B -> C;
- D -> A;
- C -> D [label="SELECT pg_disable_data_checksums()"];
- E -> A [label=" --no-data-checksums"];
- E -> C [label=" --data-checksums"];
-
- A [label="off"];
- B [label="inprogress-on"];
- C [label="on"];
- D [label="inprogress-off"];
- E [label="initdb"];
+digraph "datachecksums_states" {
+ layout=dot;
+ node [label="", shape=box, style=filled, fillcolor=gray, width=1.0, fontname="sans-serif"];
+
+ m1 [label="initdb", shape=Mdiamond];
+
+ subgraph cluster01 {
+ label="Online Checksums";
+ subgraph agroup1 {
+ rank=same;
+ a1;
+ a3;
+ }
+
+ subgraph agroup2 {
+ rank=same;
+ a2;
+ a4;
+ }
+
+ a1 -> a2;
+ a2 -> a3;
+ a3 -> a4;
+ a4 -> a1;
+
+ a1 [fillcolor=lightblue, label="off"];
+ a2 [fillcolor=lightgreen, label="inprogress-on"];
+ a3 [fillcolor=lightblue, label="on"];
+ a4 [fillcolor=lightgreen, label="inprogress-off"];
+ }
+
+ subgraph cluster05 {
+ label="Offline Checksums";
+ subgraph bgroup2 {
+ rank=same;
+ b1 -> b2;
+ }
+
+ b2 -> b1;
+
+ b1 [fillcolor=lightblue, label="off"];
+ b2 [fillcolor=lightblue, label="on"];
+ }
+
+ m1 -> a1;
+ m1 -> a3;
+ m1 -> b1;
+ m1 -> b2;
}
diff --git a/doc/src/sgml/images/datachecksums.svg b/doc/src/sgml/images/datachecksums.svg
index 8c58f42922e..eee57c393c0 100644
--- a/doc/src/sgml/images/datachecksums.svg
+++ b/doc/src/sgml/images/datachecksums.svg
@@ -1,81 +1,126 @@
-<?xml version="1.0" encoding="UTF-8" standalone="no"?>
+<?xml version="1.0"?>
<!-- Generated by graphviz version 14.0.5 (20251129.0259)
-->
-<!-- Title: G Pages: 1 -->
-<svg width="409pt" height="383pt"
- viewBox="0.00 0.00 409.00 383.00" xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink">
-<g id="graph0" class="graph" transform="scale(1 1) rotate(0) translate(4 378.5)">
-<title>G</title>
-<polygon fill="white" stroke="none" points="-4,4 -4,-378.5 404.74,-378.5 404.74,4 -4,4"/>
-<!-- A -->
+<!-- Title: datachecksums_states Pages: 1 -->
+<svg xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" width="441pt" height="209pt" viewBox="0.00 0.00 441.00 209.00">
+<g id="graph0" class="graph" transform="scale(1 1) rotate(0) translate(4 204.5)">
+<title>datachecksums_states</title>
+<polygon fill="white" stroke="none" points="-4,4 -4,-204.5 437,-204.5 437,4 -4,4"/>
+<g id="clust1" class="cluster">
+<title>cluster01</title>
+<polygon fill="none" stroke="black" points="8,-8 8,-156.5 239,-156.5 239,-8 8,-8"/>
+<text xml:space="preserve" text-anchor="middle" x="123.5" y="-139.2" font-family="Times,serif" font-size="14.00">Online Checksums</text>
+</g>
+<g id="clust4" class="cluster">
+<title>cluster05</title>
+<polygon fill="none" stroke="black" points="247,-80 247,-156.5 425,-156.5 425,-80 247,-80"/>
+<text xml:space="preserve" text-anchor="middle" x="336" y="-139.2" font-family="Times,serif" font-size="14.00">Offline Checksums</text>
+</g>
+<!-- m1 -->
<g id="node1" class="node">
-<title>A</title>
-<ellipse fill="none" stroke="black" cx="80.12" cy="-268" rx="27" ry="18"/>
-<text xml:space="preserve" text-anchor="middle" x="80.12" y="-262.95" font-family="Times,serif" font-size="14.00">off</text>
+<title>m1</title>
+<polygon fill="gray" stroke="black" points="242,-200.5 197.65,-182.5 242,-164.5 286.35,-182.5 242,-200.5"/>
+<polyline fill="none" stroke="black" points="208.77,-187.01 208.77,-177.99"/>
+<polyline fill="none" stroke="black" points="230.88,-169.01 253.12,-169.01"/>
+<polyline fill="none" stroke="black" points="275.23,-177.99 275.23,-187.01"/>
+<polyline fill="none" stroke="black" points="253.12,-195.99 230.88,-195.99"/>
+<text xml:space="preserve" text-anchor="middle" x="242" y="-176.7" font-family="sans-serif" font-size="14.00">initdb</text>
</g>
-<!-- B -->
+<!-- a1 -->
<g id="node2" class="node">
-<title>B</title>
-<ellipse fill="none" stroke="black" cx="137.12" cy="-179.5" rx="61.59" ry="18"/>
-<text xml:space="preserve" text-anchor="middle" x="137.12" y="-174.45" font-family="Times,serif" font-size="14.00">inprogress-on</text>
+<title>a1</title>
+<polygon fill="lightblue" stroke="black" points="134,-124 62,-124 62,-88 134,-88 134,-124"/>
+<text xml:space="preserve" text-anchor="middle" x="98" y="-100.2" font-family="sans-serif" font-size="14.00">off</text>
</g>
-<!-- A->B -->
-<g id="edge1" class="edge">
-<title>A->B</title>
-<path fill="none" stroke="black" d="M76.5,-249.68C75.22,-239.14 75.3,-225.77 81.12,-215.5 84.2,-210.08 88.49,-205.38 93.35,-201.34"/>
-<polygon fill="black" stroke="black" points="95.22,-204.31 101.33,-195.66 91.16,-198.61 95.22,-204.31"/>
-<text xml:space="preserve" text-anchor="middle" x="187.62" y="-218.7" font-family="Times,serif" font-size="14.00">SELECT pg_enable_data_checksums()</text>
+<!-- m1->a1 -->
+<g id="edge7" class="edge">
+<title>m1->a1</title>
+<path fill="none" stroke="black" d="M208.02,-177.91C187.98,-174.61 162.76,-168.35 143,-156.5 132.95,-150.47 123.82,-141.5 116.46,-132.85"/>
+<polygon fill="black" stroke="black" points="119.38,-130.89 110.41,-125.26 113.91,-135.26 119.38,-130.89"/>
</g>
-<!-- C -->
+<!-- a3 -->
<g id="node3" class="node">
-<title>C</title>
-<ellipse fill="none" stroke="black" cx="137.12" cy="-106.5" rx="27" ry="18"/>
-<text xml:space="preserve" text-anchor="middle" x="137.12" y="-101.45" font-family="Times,serif" font-size="14.00">on</text>
+<title>a3</title>
+<polygon fill="lightblue" stroke="black" points="224,-124 152,-124 152,-88 224,-88 224,-124"/>
+<text xml:space="preserve" text-anchor="middle" x="188" y="-100.2" font-family="sans-serif" font-size="14.00">on</text>
</g>
-<!-- B->C -->
-<g id="edge2" class="edge">
-<title>B->C</title>
-<path fill="none" stroke="black" d="M137.12,-161.31C137.12,-153.73 137.12,-144.6 137.12,-136.04"/>
-<polygon fill="black" stroke="black" points="140.62,-136.04 137.12,-126.04 133.62,-136.04 140.62,-136.04"/>
+<!-- m1->a3 -->
+<g id="edge8" class="edge">
+<title>m1->a3</title>
+<path fill="none" stroke="black" d="M232.35,-168.18C225.36,-158.54 215.68,-145.18 207.14,-133.41"/>
+<polygon fill="black" stroke="black" points="210.21,-131.68 201.51,-125.64 204.54,-135.79 210.21,-131.68"/>
+</g>
+<!-- b1 -->
+<g id="node6" class="node">
+<title>b1</title>
+<polygon fill="lightblue" stroke="black" points="327,-124 255,-124 255,-88 327,-88 327,-124"/>
+<text xml:space="preserve" text-anchor="middle" x="291" y="-100.2" font-family="sans-serif" font-size="14.00">off</text>
+</g>
+<!-- m1->b1 -->
+<g id="edge9" class="edge">
+<title>m1->b1</title>
+<path fill="none" stroke="black" d="M250.99,-167.84C257.29,-158.25 265.92,-145.13 273.55,-133.53"/>
+<polygon fill="black" stroke="black" points="276.25,-135.79 278.82,-125.51 270.4,-131.95 276.25,-135.79"/>
+</g>
+<!-- b2 -->
+<g id="node7" class="node">
+<title>b2</title>
+<polygon fill="lightblue" stroke="black" points="417,-124 345,-124 345,-88 417,-88 417,-124"/>
+<text xml:space="preserve" text-anchor="middle" x="381" y="-100.2" font-family="sans-serif" font-size="14.00">on</text>
+</g>
+<!-- m1->b2 -->
+<g id="edge10" class="edge">
+<title>m1->b2</title>
+<path fill="none" stroke="black" d="M275.15,-177.42C294.03,-173.98 317.53,-167.73 336,-156.5 346.02,-150.41 355.13,-141.42 362.49,-132.78"/>
+<polygon fill="black" stroke="black" points="365.04,-135.2 368.56,-125.2 359.57,-130.82 365.04,-135.2"/>
</g>
-<!-- D -->
+<!-- a2 -->
<g id="node4" class="node">
-<title>D</title>
-<ellipse fill="none" stroke="black" cx="63.12" cy="-18" rx="63.12" ry="18"/>
-<text xml:space="preserve" text-anchor="middle" x="63.12" y="-12.95" font-family="Times,serif" font-size="14.00">inprogress-off</text>
+<title>a2</title>
+<polygon fill="lightgreen" stroke="black" points="114.25,-52 15.75,-52 15.75,-16 114.25,-16 114.25,-52"/>
+<text xml:space="preserve" text-anchor="middle" x="65" y="-28.2" font-family="sans-serif" font-size="14.00">inprogress-on</text>
</g>
-<!-- C->D -->
-<g id="edge4" class="edge">
-<title>C->D</title>
-<path fill="none" stroke="black" d="M124.23,-90.43C113.36,-77.73 97.58,-59.28 84.77,-44.31"/>
-<polygon fill="black" stroke="black" points="87.78,-42.44 78.62,-37.12 82.46,-46.99 87.78,-42.44"/>
-<text xml:space="preserve" text-anchor="middle" x="214.75" y="-57.2" font-family="Times,serif" font-size="14.00">SELECT pg_disable_data_checksums()</text>
+<!-- a1->a2 -->
+<g id="edge1" class="edge">
+<title>a1->a2</title>
+<path fill="none" stroke="black" d="M89.84,-87.7C86.25,-80.07 81.93,-70.92 77.92,-62.4"/>
+<polygon fill="black" stroke="black" points="81.14,-61.03 73.71,-53.47 74.81,-64.01 81.14,-61.03"/>
+</g>
+<!-- a4 -->
+<g id="node5" class="node">
+<title>a4</title>
+<polygon fill="lightgreen" stroke="black" points="231.25,-52 132.75,-52 132.75,-16 231.25,-16 231.25,-52"/>
+<text xml:space="preserve" text-anchor="middle" x="182" y="-28.2" font-family="sans-serif" font-size="14.00">inprogress-off</text>
</g>
-<!-- D->A -->
+<!-- a3->a4 -->
<g id="edge3" class="edge">
-<title>D->A</title>
-<path fill="none" stroke="black" d="M62.52,-36.28C61.62,-68.21 60.54,-138.57 66.12,-197.5 67.43,-211.24 70.27,-226.28 73.06,-238.85"/>
-<polygon fill="black" stroke="black" points="69.64,-239.59 75.32,-248.54 76.46,-238 69.64,-239.59"/>
+<title>a3->a4</title>
+<path fill="none" stroke="black" d="M186.52,-87.7C185.89,-80.41 185.15,-71.73 184.45,-63.54"/>
+<polygon fill="black" stroke="black" points="187.94,-63.28 183.6,-53.61 180.96,-63.87 187.94,-63.28"/>
</g>
-<!-- E -->
-<g id="node5" class="node">
-<title>E</title>
-<ellipse fill="none" stroke="black" cx="198.12" cy="-356.5" rx="32.41" ry="18"/>
-<text xml:space="preserve" text-anchor="middle" x="198.12" y="-351.45" font-family="Times,serif" font-size="14.00">initdb</text>
+<!-- a2->a3 -->
+<g id="edge2" class="edge">
+<title>a2->a3</title>
+<path fill="none" stroke="black" d="M95.33,-52.26C111.09,-61.23 130.56,-72.31 147.59,-82"/>
+<polygon fill="black" stroke="black" points="145.54,-84.86 155.96,-86.77 149,-78.78 145.54,-84.86"/>
+</g>
+<!-- a4->a1 -->
+<g id="edge4" class="edge">
+<title>a4->a1</title>
+<path fill="none" stroke="black" d="M161.18,-52.35C150.98,-60.85 138.51,-71.24 127.36,-80.53"/>
+<polygon fill="black" stroke="black" points="125.37,-77.64 119.93,-86.73 129.85,-83.01 125.37,-77.64"/>
</g>
-<!-- E->A -->
+<!-- b1->b2 -->
<g id="edge5" class="edge">
-<title>E->A</title>
-<path fill="none" stroke="black" d="M179.16,-341.6C159.64,-327.29 129.05,-304.86 107.03,-288.72"/>
-<polygon fill="black" stroke="black" points="109.23,-286 99.1,-282.91 105.09,-291.64 109.23,-286"/>
-<text xml:space="preserve" text-anchor="middle" x="208.57" y="-307.2" font-family="Times,serif" font-size="14.00"> --no-data-checksums</text>
+<title>b1->b2</title>
+<path fill="none" stroke="black" d="M327.21,-93.01C329.33,-92.77 331.44,-92.61 333.55,-92.54"/>
+<polygon fill="black" stroke="black" points="333.21,-96.03 343.35,-92.96 333.51,-89.03 333.21,-96.03"/>
</g>
-<!-- E->C -->
+<!-- b2->b1 -->
<g id="edge6" class="edge">
-<title>E->C</title>
-<path fill="none" stroke="black" d="M227.13,-348.04C242.29,-342.72 259.95,-334.06 271.12,-320.5 301.5,-283.62 316.36,-257.78 294.12,-215.5 268.41,-166.6 209.42,-135.53 171.52,-119.85"/>
-<polygon fill="black" stroke="black" points="172.96,-116.65 162.37,-116.21 170.37,-123.16 172.96,-116.65"/>
-<text xml:space="preserve" text-anchor="middle" x="350.87" y="-218.7" font-family="Times,serif" font-size="14.00"> --data-checksums</text>
+<title>b2->b1</title>
+<path fill="none" stroke="black" d="M344.86,-118.98C342.75,-119.23 340.63,-119.39 338.52,-119.46"/>
+<polygon fill="black" stroke="black" points="338.86,-115.97 328.72,-119.05 338.57,-122.96 338.86,-115.97"/>
</g>
</g>
</svg>
diff --git a/doc/src/sgml/images/meson.build b/doc/src/sgml/images/meson.build
index 220e3eaafb8..54e80fb402c 100644
--- a/doc/src/sgml/images/meson.build
+++ b/doc/src/sgml/images/meson.build
@@ -11,6 +11,7 @@ image_targets = []
fixup_svg_xsl = files('fixup-svg.xsl')
all_files = [
+ 'datachecksums.gv',
'genetic-algorithm.gv',
'gin.gv',
'pagelayout.txt',
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index 6a4cb9bb144..8614338fcab 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -8339,7 +8339,7 @@ FROM pg_stat_get_backend_idset() AS backendid;
<structfield>blocks_total</structfield> <type>bigint</type>
</para>
<para>
- The number of blocks in the current relation which will be processed,
+ The number of blocks in the current relation fork which will be processed,
or <literal>NULL</literal> if the worker process hasn't
calculated the number of blocks yet. The launcher process has
this set to <literal>NULL</literal>.
@@ -8353,7 +8353,7 @@ FROM pg_stat_get_backend_idset() AS backendid;
<structfield>blocks_done</structfield> <type>bigint</type>
</para>
<para>
- The number of blocks in the current relation which have been processed.
+ The number of blocks in the current relation fork which have been processed.
The launcher process has this set to <literal>NULL</literal>.
</para>
</entry>
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index abc035e5409..5a0bda2eca2 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -288,7 +288,7 @@ PostgreSQL documentation
</procedure>
<para>
- The replay requirement in <xref linkend="shut-down-nodes"/>exists
+ The replay requirement in <xref linkend="shut-down-nodes"/> exists
because an offline change is recorded only in the control file and
has no defined ordering against WAL the node has not replayed yet; see
<xref linkend="checksums-offline-enable-disable"/>. A node stopped
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index dbbfdae5e4f..edd39ee437f 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -316,30 +316,6 @@
application can be used to enable or disable data checksums, as well as
verify checksums, on an offline cluster.
</para>
-
- <para>
- An offline change provides durability differently from an
- <link linkend="checksums-online-enable-disable">online change</link>.
- An online transition is WAL-logged: it is ordered against all other
- WAL records, it is replayed after a crash, and it propagates to
- standbys. An offline change is recorded only in the cluster's
- control file: it writes no WAL, it is invisible to replication, and
- it has no defined ordering against WAL the node has not replayed
- yet. When a node later replays WAL that contains an online checksum
- state change, that change takes effect on the node even if it was
- written before the offline change was made.
- </para>
-
- <para>
- An offline change only affects the data directory it is run on; the
- new state does not propagate over replication. In a replication setup
- the same change must be applied to all nodes while all of them are
- stopped, as described in <xref linkend="app-pgchecksums"/>. A standby
- whose state diverges logs a warning but keeps its local setting. The
- mismatch persists until the states are brought together again, with
- the offline procedure or with an online transition; do this promptly.
- </para>
-
</sect2>
<sect2 id="checksums-online-enable-disable" xreflabel="Online Enabling of Checksums">
@@ -456,7 +432,43 @@
still required.
</para>
</sect3>
+ </sect2>
+
+ <sect2 id="checksums-mismatched-states">
+ <title>Mismatched Data Checksums States in a Replicated Cluster</title>
+ <para>
+ The primary and secondaries can end up with different data checksums
+ states due to the different durability models between offline and online
+ checksum operations.
+ </para>
+ <para>
+ An online transition is WAL-logged, which means that it is replayed after a
+ crash, and the transition along with all checksum updates are propagated to
+ standbys. An offline change is recorded only in the cluster's control
+ file, it writes no WAL and is invisible to replication. The changes made
+ to the datafiles to write checksums are not WAL logged. This means that it
+ has no defined ordering against WAL the node has not replayed yet. When a
+ node later replays WAL that contains an online checksum state change, that
+ change takes effect on the node even if it was written before the offline
+ change was made.
+ </para>
+ <para>
+ An offline change only affects the data directory it is run on; the new
+ state does not propagate over replication. This means that all nodes can
+ be operated on in parallel, and no additional network traffic or WAL
+ archive traffic will occur. In a replication setup the same change must be
+ applied to all nodes while all of them are stopped, as described in
+ <xref linkend="app-pgchecksums"/>. A standby whose state diverges logs a
+ warning but keeps its local setting. The mismatch persists until the
+ states are brought together again, with the offline procedure or with an
+ online transition.
+ </para>
+
+ <para>
+ Running a replicated cluster with different data checksum states on the
+ different nodes is not supported or recommended.
+ </para>
</sect2>
</sect1>
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v16-0004-pg_combinebackup-Refuse-mixed-data-checksum-stat.patch (10.0K, ../../0E4734FA-4764-405C-9670-C45DF440A718@yesql.se/7-v16-0004-pg_combinebackup-Refuse-mixed-data-checksum-stat.patch)
download | inline diff:
From cb0fa8e654947cf711d6c33c0e0507146fa1e8d6 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 3 Sep 2026 23:32:04 +0200
Subject: [PATCH v16 4/9] pg_combinebackup: Refuse mixed data checksum states
in a backup chain
check_control_files() detected a chain whose backups were taken under
different data checksum states, warned, and proceeded. The output
directory keeps the last backup's control file, so a full backup taken
with checksums off combined with an incremental taken after an offline
enable produces a cluster whose control file says "on" while most of
its blocks carry no checksums; it starts, and then every connection
dies on the first unchecksummed catalog page. An offline enable
between two backups of a chain is all it takes, since it rewrites
every page without logging anything, so the incremental backup does
not re-ship the pages.
Turn the warning into an error, matching what pg_rewind does for the
equivalent combinations. The check stays asymmetric on purpose: when
the last backup was taken with checksums off, stale checksums from an
earlier backup are never verified, and that chain remains usable.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Reviewed-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Discussion: https://postgr.es/m/anwm6UPxoVS41QA2@bdtpg
---
doc/src/sgml/ref/pg_combinebackup.sgml | 19 +--
src/bin/pg_combinebackup/pg_combinebackup.c | 14 +-
src/test/modules/test_checksums/meson.build | 1 +
.../t/024_combinebackup_mixed.pl | 146 ++++++++++++++++++
4 files changed, 167 insertions(+), 13 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
diff --git a/doc/src/sgml/ref/pg_combinebackup.sgml b/doc/src/sgml/ref/pg_combinebackup.sgml
index 9a6d201e0b8..f5d7e177b36 100644
--- a/doc/src/sgml/ref/pg_combinebackup.sgml
+++ b/doc/src/sgml/ref/pg_combinebackup.sgml
@@ -306,18 +306,19 @@ PostgreSQL documentation
<para>
<literal>pg_combinebackup</literal> does not recompute page checksums when
- writing the output directory. Therefore, if any of the backups used for
- reconstruction were taken with checksums disabled, but the final backup was
- taken with checksums enabled, the resulting directory may contain pages
- with invalid checksums.
+ writing the output directory. It therefore refuses a chain in which some
+ of the backups used for reconstruction were taken with checksums disabled
+ but the final backup was taken with checksums enabled: the resulting
+ directory would contain pages with invalid checksums, and the cluster
+ would fail checksum verification as soon as it reads them. Take a new
+ full backup after enabling data checksums with
+ <xref linkend="app-pgchecksums"/>.
</para>
<para>
- To avoid this problem, taking a new full backup after changing the checksum
- state of the cluster using <xref linkend="app-pgchecksums"/> is
- recommended. Otherwise, you can disable and then optionally reenable
- checksums on the directory produced by <literal>pg_combinebackup</literal>
- in order to correct the problem.
+ The reverse case is accepted: when the final backup was taken with
+ checksums disabled, stale checksums copied from an earlier backup are
+ never verified.
</para>
</refsect1>
diff --git a/src/bin/pg_combinebackup/pg_combinebackup.c b/src/bin/pg_combinebackup/pg_combinebackup.c
index 86e2ee37c40..254a27b125b 100644
--- a/src/bin/pg_combinebackup/pg_combinebackup.c
+++ b/src/bin/pg_combinebackup/pg_combinebackup.c
@@ -670,13 +670,19 @@ check_control_files(int n_backups, char **backup_dirs)
pg_log_debug("system identifier is %" PRIu64, system_identifier);
/*
- * Warn the user if not all backups are in the same state with regards to
- * checksums.
+ * Reject a chain whose backups were not all in the same state with
+ * regards to checksums. The loop above only flags that when the last
+ * backup has checksums enabled: the blocks taken from the older backups
+ * have no checksums, and pg_combinebackup does not recompute them, so the
+ * combined cluster would fail verification as soon as it reads them. The
+ * other direction is harmless, since stale checksums are never verified
+ * while checksums are disabled.
*/
if (data_checksum_mismatch)
{
- pg_log_warning("only some backups have checksums enabled");
- pg_log_warning_hint("Disable, and optionally reenable, checksums on the output directory to avoid failures.");
+ pg_log_error("only some backups have checksums enabled");
+ pg_log_error_hint("Take a new full backup after changing the data checksum state with pg_checksums.");
+ exit(1);
}
return system_identifier;
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index 7d07c757052..7eccd5156b9 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -47,6 +47,7 @@ tests += {
't/021_rewind_divergent_transitions.pl',
't/022_rewind_state.pl',
't/023_rewind_standby_target.pl',
+ 't/024_combinebackup_mixed.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
new file mode 100644
index 00000000000..2f51c7fd9dc
--- /dev/null
+++ b/src/test/modules/test_checksums/t/024_combinebackup_mixed.pl
@@ -0,0 +1,146 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_combinebackup refuses a chain whose backups were taken under
+# different data checksum states when the final state has them enabled.
+#
+# The output directory keeps the last backup's control file, so a full
+# backup taken with checksums off plus an incremental taken after an
+# offline enable would produce a directory whose control file says "on"
+# while most of its blocks have no checksums. An offline enable between
+# the two backups is enough to get there: it rewrites every page but logs
+# nothing, so the incremental backup does not re-ship the pages.
+#
+# The reverse order stays allowed: with the final backup taken with
+# checksums off, stale checksums from an earlier backup are never verified.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('node');
+$node->init(
+ no_data_checksums => 1,
+ has_archiving => 1,
+ allows_streaming => 1);
+$node->append_conf('postgresql.conf', 'summarize_wal = on');
+$node->append_conf('postgresql.conf', 'autovacuum = off');
+$node->start;
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+test_checksum_state($node, 'off');
+
+$node->backup('full');
+
+# Offline enable: rewrites every page, logs nothing.
+$node->stop;
+system_or_bail('pg_checksums', '--enable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'on');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full/backup_manifest',
+ ],
+ 'incremental backup with checksums on');
+
+# Combining the chain must be refused: the result would say "on" while the
+# blocks inherited from the full backup have no checksums.
+command_fails_like(
+ [
+ 'pg_combinebackup',
+ $node->backup_dir . '/full',
+ $node->backup_dir . '/incr',
+ '--output' => $node->backup_dir . '/combined',
+ ],
+ qr/only some backups have checksums enabled/,
+ 'pg_combinebackup refuses a chain crossing an offline enable');
+
+# The other direction: full backup with checksums on, offline disable, then
+# an incremental. The last backup wins, so the combined cluster comes up off.
+$node->backup('full2');
+
+$node->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $node->data_dir);
+$node->start;
+test_checksum_state($node, 'off');
+
+$node->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$node->safe_psql('postgres', 'CHECKPOINT;');
+
+$node->command_ok(
+ [
+ 'pg_basebackup',
+ '--pgdata' => $node->backup_dir . '/incr2',
+ '--dbname' => $node->connstr('postgres'),
+ '--no-sync',
+ '--checkpoint' => 'fast',
+ '--incremental' => $node->backup_dir . '/full2/backup_manifest',
+ ],
+ 'incremental backup with checksums off');
+
+my $restored = PostgreSQL::Test::Cluster->new('restored');
+$restored->init_from_backup(
+ $node, 'incr2',
+ combine_with_prior => ['full2'],
+ has_restoring => 1,
+ standby => 0);
+
+$restored->start;
+
+my ($rc, $stdout, $stderr) = $restored->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'combined backup accepts connections') or diag("stderr: $stderr");
+is($stdout, 'off', 'combined backup follows the last backup state');
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t3;');
+is($rc, 0, 'blocks from the incremental backup are readable')
+ or diag("stderr: $stderr");
+
+($rc, $stdout, $stderr) =
+ $restored->psql('postgres', 'SELECT count(*) FROM t1;');
+is($rc, 0, 'blocks inherited from the full backup are readable')
+ or diag("stderr: $stderr");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($restored->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures in the combined backup');
+
+$restored->stop('immediate');
+$node->stop;
+
+done_testing();
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v16-0003-pg_rewind-Check-the-data-checksum-states-of-sour.patch (22.5K, ../../0E4734FA-4764-405C-9670-C45DF440A718@yesql.se/8-v16-0003-pg_rewind-Check-the-data-checksum-states-of-sour.patch)
download | inline diff:
From 16d9b4ab0bd94df601cf2cbe2644536f3c2394f5 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 3 Sep 2026 23:31:52 +0200
Subject: [PATCH v16 3/9] pg_rewind: Check the data checksum states of source
and target
pg_rewind installs the source's control file on the target, but every
block it does not copy keeps the target's content. With checksums
enabled on the target and disabled on the source, the rewound server
claims enabled checksums while the blocks copied from the source have
none, and fails checksum verification as soon as it reads them; in the
worst case every connection attempt dies on an unverifiable catalog
page.
Comparing the control files is not enough. Replay on the rewound
server resumes from the last common checkpoint and adopts the data
checksum state recorded there, so an offline disable on the source
after the divergence leaves both control files saying "off" while the
rewound server still resumes with verification enabled. Check the
state carried by the divergence checkpoint as well, which pg_rewind
already reads.
Refuse both cases, and an interrupted online transition on either
side, same as pg_checksums does. The opposite mismatch, checksums on
the source only, is allowed with a warning: recovery under the backup
label adopts the divergence state, so the rewound server keeps
checksums disabled (or converges through the replayed WAL if the
source enabled them online), and no page can fail verification.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Reviewed-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Discussion: https://postgr.es/m/anwm6UPxoVS41QA2@bdtpg
---
doc/src/sgml/ref/pg_rewind.sgml | 14 ++
src/bin/pg_rewind/parsexlog.c | 4 +-
src/bin/pg_rewind/pg_rewind.c | 93 ++++++++++-
src/bin/pg_rewind/pg_rewind.h | 1 +
src/test/modules/test_checksums/meson.build | 2 +
.../test_checksums/t/022_rewind_state.pl | 145 +++++++++++++++++
.../t/023_rewind_standby_target.pl | 153 ++++++++++++++++++
7 files changed, 410 insertions(+), 2 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/022_rewind_state.pl
create mode 100644 src/test/modules/test_checksums/t/023_rewind_standby_target.pl
diff --git a/doc/src/sgml/ref/pg_rewind.sgml b/doc/src/sgml/ref/pg_rewind.sgml
index b95e3868e9d..ac3d0c9328f 100644
--- a/doc/src/sgml/ref/pg_rewind.sgml
+++ b/doc/src/sgml/ref/pg_rewind.sgml
@@ -124,6 +124,20 @@ PostgreSQL documentation
<literal>on</literal>, but is enabled by default.
</para>
+ <para>
+ The data checksum states of the source and target server must be
+ compatible. <application>pg_rewind</application> refuses to run if an
+ online checksum state transition was interrupted on either server, or
+ if data checksums are enabled on the target server, or were enabled at
+ the point of divergence, while the source server runs without them: the
+ rewound server would fail checksum verification on the blocks copied
+ from the source. The opposite combination is allowed with a warning;
+ the rewound server keeps data checksums disabled unless the source
+ server enabled them online after the point of divergence. See
+ <xref linkend="app-pgchecksums"/> for how to change the state
+ consistently across a replication setup.
+ </para>
+
<warning>
<title>Warning: Failures While Rewinding</title>
<para>
diff --git a/src/bin/pg_rewind/parsexlog.c b/src/bin/pg_rewind/parsexlog.c
index 023e23b063c..6e87b00f8c2 100644
--- a/src/bin/pg_rewind/parsexlog.c
+++ b/src/bin/pg_rewind/parsexlog.c
@@ -167,7 +167,8 @@ readOneRecord(const char *datadir, XLogRecPtr ptr, int tliIndex,
void
findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
- XLogRecPtr *lastchkptredo, const char *restoreCommand)
+ XLogRecPtr *lastchkptredo, uint32 *lastchkptdatachecksums,
+ const char *restoreCommand)
{
/* Walk backwards, starting from the given record */
XLogRecord *record;
@@ -255,6 +256,7 @@ findLastCheckpoint(const char *datadir, XLogRecPtr forkptr, int tliIndex,
*lastchkptrec = searchptr;
*lastchkpttli = checkPoint.ThisTimeLineID;
*lastchkptredo = checkPoint.redo;
+ *lastchkptdatachecksums = checkPoint.dataChecksumState;
break;
}
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 0e4c2df3f4b..d2521dab333 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -145,6 +145,7 @@ main(int argc, char **argv)
XLogRecPtr chkptrec;
TimeLineID chkpttli;
XLogRecPtr chkptredo;
+ uint32 chkptdatachecksums;
TimeLineID source_tli;
TimeLineID target_tli;
XLogRecPtr target_wal_endrec;
@@ -472,10 +473,51 @@ main(int argc, char **argv)
keepwal_init();
findLastCheckpoint(datadir_target, divergerec, lastcommontliIndex,
- &chkptrec, &chkpttli, &chkptredo, restore_command);
+ &chkptrec, &chkpttli, &chkptredo, &chkptdatachecksums,
+ restore_command);
pg_log_info("rewinding from last common checkpoint at %X/%08X on timeline %u",
LSN_FORMAT_ARGS(chkptrec), chkpttli);
+ /*
+ * Replay on the rewound server resumes from the last common checkpoint
+ * and adopts the data checksum state recorded there, not the state in the
+ * control file installed from the source. If checksums were enabled at
+ * the divergence point but the source runs without them, the rewound
+ * server would verify checksums while the blocks copied from the source
+ * have none. The control files cannot reveal this: an offline disable on
+ * the source after the divergence leaves both of them saying "off".
+ *
+ * Only the fully enabled state needs checking. An in-progress state at
+ * the divergence point means a transition was still running there, and
+ * all of its page rewrites are logged after that point, so replay brings
+ * the rewound server to whatever state the copied WAL ends in.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the source server or recreate the target server from a base backup.");
+ exit(1);
+ }
+
+ /*
+ * The same comparison is needed against the target: every block the
+ * rewind does not copy keeps the target's content. The control files
+ * cannot reveal this case either, in the other direction: a standby's
+ * checkpoints are written by its upstream primary, so a standby whose
+ * checksums were disabled offline still has "on" checkpoints in its WAL,
+ * while both control files may agree.
+ */
+ if (chkptdatachecksums == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_target.data_checksum_version != PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_error("data checksums were enabled at the point of divergence but are disabled on the target server");
+ pg_log_error_detail("Blocks kept from the target would have no checksums, but the rewound server would resume with checksum verification enabled.");
+ pg_log_error_hint("Either enable data checksums on the target server with pg_checksums or recreate the target server from a base backup.");
+ exit(1);
+ }
+
/* Initialize the hash table to track the status of each file */
filehash_init();
@@ -799,6 +841,55 @@ sanityChecks(void)
pg_fatal("target server needs to use either data checksums or \"wal_log_hints = on\"");
}
+ /*
+ * The rewound target keeps its control file fields from the source, but
+ * every block the rewind does not copy keeps the target's content. If
+ * checksums are enabled on the target and disabled on the source, the
+ * result would claim enabled checksums while the blocks copied from the
+ * source have none, and the target would fail checksum verification as
+ * soon as it reads them. Refuse that combination, and refuse an
+ * interrupted online transition on either side, same as pg_checksums.
+ *
+ * The opposite mismatch is allowed: recovery under the backup label
+ * written by pg_rewind adopts the checksum state as of the divergence
+ * point, so a target whose own checkpoints carry no checksums stays
+ * without them, or converges through the replayed WAL if the source
+ * enabled them online. Only warn about it, so that an offline change on
+ * the source is not overlooked.
+ *
+ * That reasoning does not hold when the target is a standby, whose
+ * checkpoints were written by its upstream primary; see the checks after
+ * findLastCheckpoint().
+ */
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted on the target server");
+ pg_log_error_hint("Start the server and shut it down cleanly to reset the state, then retry; on a standby, let replication complete the transition first.");
+ exit(1);
+ }
+ if (ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition is incomplete on the source server");
+ pg_log_error_hint("Let the transition complete, or reset the state with a clean restart, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_OFF)
+ {
+ pg_log_error("data checksums are enabled on the target server but disabled on the source server");
+ pg_log_error_detail("Blocks copied from the source would have no checksums, and the target would fail checksum verification after the rewind.");
+ pg_log_error_hint("Either disable data checksums on the target server or enable them on the source server, then retry.");
+ exit(1);
+ }
+ if (ControlFile_target.data_checksum_version == PG_DATA_CHECKSUM_OFF &&
+ ControlFile_source.data_checksum_version == PG_DATA_CHECKSUM_VERSION)
+ {
+ pg_log_warning("data checksums are disabled on the target server but enabled on the source server");
+ pg_log_warning_detail("The rewound server will keep data checksums disabled unless the source server enabled them online after the point of divergence.");
+ }
+
/*
* Target cluster better not be running. This doesn't guard against
* someone starting the cluster concurrently. Also, this is probably more
diff --git a/src/bin/pg_rewind/pg_rewind.h b/src/bin/pg_rewind/pg_rewind.h
index 9a981f7f246..9187c3bf72f 100644
--- a/src/bin/pg_rewind/pg_rewind.h
+++ b/src/bin/pg_rewind/pg_rewind.h
@@ -39,6 +39,7 @@ extern void findLastCheckpoint(const char *datadir, XLogRecPtr forkptr,
int tliIndex,
XLogRecPtr *lastchkptrec, TimeLineID *lastchkpttli,
XLogRecPtr *lastchkptredo,
+ uint32 *lastchkptdatachecksums,
const char *restoreCommand);
extern XLogRecPtr readOneRecord(const char *datadir, XLogRecPtr ptr,
int tliIndex, const char *restoreCommand);
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index e5d38fafb7d..7d07c757052 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -45,6 +45,8 @@ tests += {
't/019_standby_shutdown_catchup.pl',
't/020_cascade_divergence.pl',
't/021_rewind_divergent_transitions.pl',
+ 't/022_rewind_state.pl',
+ 't/023_rewind_standby_target.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/022_rewind_state.pl b/src/test/modules/test_checksums/t/022_rewind_state.pl
new file mode 100644
index 00000000000..e6d9c645a0b
--- /dev/null
+++ b/src/test/modules/test_checksums/t/022_rewind_state.pl
@@ -0,0 +1,145 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind checks the data checksum states of source and target.
+# Checksums enabled on the target, or at the point of divergence, with
+# a source running without them are refused: the rewound server would
+# verify checksums on blocks copied from a source that has none. The
+# opposite mismatch only warns; replay keeps the target's own state.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+# Scenarios 1 and 2: checksums on from initdb. wal_log_hints keeps the
+# target eligible for pg_rewind after checksums are disabled on it.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b);
+
+# Failover to B, and divergence on A.
+$node_b->promote;
+$node_b->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_a->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_a->stop('fast');
+
+# Offline disable on the source only: enabled target, disabled source.
+$node_b->stop;
+$node_b->checksum_disable_offline;
+$node_b->start;
+test_checksum_state($node_b, 'off');
+
+my @rewind_cmd = (
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'));
+
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums are enabled on the target server but disabled on the source server/,
+ 'refuses an enabled target with a disabled source');
+
+# The refusal happens before any modification: the target still starts
+# on its own timeline.
+$node_a->start;
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'target untouched by the refused rewind');
+$node_a->stop('fast');
+
+# Scenario 2: disabling the target too makes the control files match,
+# but checksums were still enabled at the point of divergence, which is
+# the state replay on the rewound server would resume with.
+$node_a->checksum_disable_offline;
+command_fails_like(
+ \@rewind_cmd,
+ qr/data checksums were enabled at the point of divergence but are disabled on the source server/,
+ 'refuses when checksums were enabled at the point of divergence');
+
+# Scenario 3: fresh pair without checksums; offline enable on the
+# source only. Allowed with a warning, and the rewound server keeps
+# checksums disabled.
+my $node_c = PostgreSQL::Test::Cluster->new('node_c');
+$node_c->init(allows_streaming => 1, no_data_checksums => 1);
+$node_c->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$node_c->start;
+$node_c->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$node_c->backup('backup');
+my $node_d = PostgreSQL::Test::Cluster->new('node_d');
+$node_d->init_from_backup($node_c, 'backup', has_streaming => 1);
+$node_d->start;
+$node_c->wait_for_catchup($node_d);
+
+$node_d->promote;
+$node_d->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$node_c->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$node_c->stop('fast');
+
+$node_d->stop;
+$node_d->checksum_enable_offline;
+$node_d->start;
+test_checksum_state($node_d, 'on');
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_c->data_dir,
+ '--source-server' => $node_d->connstr('postgres'),
+ ]);
+like(
+ $stderr,
+ qr/data checksums are disabled on the target server but enabled on the source server/,
+ 'warns for a disabled target with an enabled source');
+like($stderr, qr/Done!/, 'rewind completed despite the warning');
+
+# The rewound server follows D and keeps its own state.
+$node_c->append_conf('postgresql.conf', 'port = ' . $node_c->port);
+$node_c->enable_streaming($node_d);
+$node_c->set_standby_mode;
+$node_c->start;
+$node_d->wait_for_catchup($node_c);
+is($node_c->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+test_checksum_state($node_c, 'off');
+
+$node_c->stop;
+$node_d->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/023_rewind_standby_target.pl b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
new file mode 100644
index 00000000000..9071114a9ae
--- /dev/null
+++ b/src/test/modules/test_checksums/t/023_rewind_standby_target.pl
@@ -0,0 +1,153 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that pg_rewind compares the data checksum state at the point of
+# divergence against the target as well as the source.
+#
+# A standby's checkpoints are written by its upstream primary, so a standby
+# whose checksums were disabled offline still has "on" checkpoints in its
+# WAL. Rewinding it onto a promoted sibling passes every control file
+# check (target off + source on only warns), yet recovery after the rewind
+# adopts the divergence checkpoint's state and the rewound server would
+# verify checksums over the pages it wrote while it was locally "off".
+# pg_rewind must refuse; after an offline enable on the target, which
+# rewrites all of its pages, the rewind goes through.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum_extended' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum_extended\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t0 AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+
+my $lagging = PostgreSQL::Test::Cluster->new('lagging');
+$lagging->init_from_backup($primary, 'backup', has_streaming => 1);
+$lagging->start;
+
+my $failover = PostgreSQL::Test::Cluster->new('failover');
+$failover->init_from_backup($primary, 'backup', has_streaming => 1);
+$failover->start;
+
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+
+test_checksum_state($primary, 'on');
+test_checksum_state($lagging, 'on');
+test_checksum_state($failover, 'on');
+
+# Offline-disable checksums on the lagging standby only. This is a
+# divergence 012_offline_standby.pl declares survivable: the standby keeps
+# its own state, warns, and stays readable.
+$lagging->stop;
+system_or_bail('pg_checksums', '--disable', '--pgdata', $lagging->data_dir);
+$lagging->start;
+test_checksum_state($lagging, 'off');
+
+# Give the lagging standby plenty of pages to write while it is locally
+# "off", and force a restartpoint so they reach disk without checksums.
+# The other standby replays the same WAL with checksums on.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t1 AS SELECT generate_series(1,200000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($lagging);
+$primary->wait_for_catchup($failover);
+$lagging->safe_psql('postgres', "CHECKPOINT;");
+
+# Take the "failover" standby out of the picture. It is promoted from
+# this position later, so everything replayed from here on is the part of
+# the WAL that diverges and that pg_rewind will roll back. t1 is already
+# replayed everywhere, so its blocks are *not* rolled back.
+$failover->stop;
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,1000) AS a;");
+$primary->wait_for_catchup($lagging);
+
+# Failover: promote the other standby, which forks the timeline behind the
+# replay position of the lagging standby.
+$primary->stop('fast');
+$failover->start;
+$failover->promote;
+$failover->safe_psql('postgres',
+ "CREATE TABLE t3 AS SELECT generate_series(1,1000) AS a;");
+$failover->safe_psql('postgres', "CHECKPOINT;");
+
+# Re-attaching the lagging standby with pg_rewind must be refused: the
+# divergence checkpoint says "on" while the target is "off", so the rewound
+# server would verify checksums over the blocks it keeps.
+$lagging->stop;
+
+command_fails_like(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ],
+ qr/data checksums were enabled at the point of divergence but are disabled on the target server/,
+ 'pg_rewind refuses a target whose divergence checkpoint has checksums');
+
+# Enabling checksums offline on the target rewrites all of its pages, after
+# which the rewind is safe.
+system_or_bail('pg_checksums', '--enable', '--pgdata', $lagging->data_dir);
+
+my ($stdout, $stderr) = run_command(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $lagging->data_dir,
+ '--source-server' => $failover->connstr('postgres'),
+ ]);
+like($stderr, qr/Done!/, 'pg_rewind completes after the offline enable');
+
+$lagging->append_conf('postgresql.conf', 'port = ' . $lagging->port);
+$lagging->enable_streaming($failover);
+$lagging->set_standby_mode;
+$lagging->start;
+
+my ($rc, $out, $err) =
+ $lagging->psql('postgres',
+ "SELECT setting FROM pg_settings " . "WHERE name = 'data_checksums';");
+is($rc, 0, 'rewound standby accepts connections') or diag("stderr: $err");
+is($out, 'on', 'rewound standby resumes with checksums on');
+
+($rc, $out, $err) = $lagging->psql('postgres', "SELECT count(*) FROM t1;");
+is($rc, 0, 'blocks kept from the target stay readable')
+ or diag("stderr: $err");
+
+my $log = PostgreSQL::Test::Utils::slurp_file($lagging->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the rewound standby');
+
+$lagging->stop('immediate');
+$failover->stop;
+done_testing();
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v16-0002-pg_checksums-Refuse-interrupted-transitions-note.patch (7.9K, ../../0E4734FA-4764-405C-9670-C45DF440A718@yesql.se/9-v16-0002-pg_checksums-Refuse-interrupted-transitions-note.patch)
download | inline diff:
From 782c04b1c65e992fe1769d0f371837aad7555698 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 3 Sep 2026 23:31:33 +0200
Subject: [PATCH v16 2/9] pg_checksums: Refuse interrupted transitions, note
change is local
An inprogress-on or inprogress-off control file means an online
transition was cut short; running the offline tool on top of it mixes
two procedures. Refuse it in every mode, with a hint pointing at the
start-stop cycle (or, on a standby, at letting replication finish)
that resets the state. Also print that the change applies to one
data directory only, with standby-aware wording, since in a
replication setup the same change must be made on every node.
On a primary, inprogress-on cannot survive a graceful stop: the
checksums launcher resolves it from its exit cleanup, and a crashed
primary is already rejected by the existing "cluster must be shut
down" check. inprogress-off can, however: it is set by the backend
running pg_disable_data_checksums(), and a fast shutdown arriving
between its two barriers leaves it behind in a cleanly shut down
control file, where the next start-stop cycle resolves it at end of
recovery. A standby can be stopped with either state, since it has
no launcher and only carries forward whatever state the last replayed
WAL record left it in, with a restartpoint persisting that as-is.
The standby is also the deterministic way to reach the new guard,
which is why the test coverage uses one; the primary window would
need an injection point between the two barriers.
The pg_checksums page said the tool still processes all relation files
regardless of an interrupted online transition; describe the refusal
instead.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Reviewed-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Discussion: https://postgr.es/m/anwm6UPxoVS41QA2@bdtpg
---
doc/src/sgml/ref/pg_checksums.sgml | 11 ++++---
src/bin/pg_checksums/pg_checksums.c | 21 +++++++++++++
.../test_checksums/t/012_offline_standby.pl | 31 ++++++++++++++-----
3 files changed, 51 insertions(+), 12 deletions(-)
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index 2d2057a8c19..abc035e5409 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -49,11 +49,12 @@ PostgreSQL documentation
</para>
<para>
- When enabling checksums with <application>pg_checksums</application>, if
- checksums were in the process of being enabled using
- <xref linkend="checksums-online-enable-disable"/> when the cluster was shut
- down, <application>pg_checksums</application> will still process all
- relation files regardless of the progress of online checksum processing.
+ If checksums were in the process of being enabled or disabled using
+ <xref linkend="checksums-online-enable-disable"/> when the cluster was
+ shut down, the control file still records that interrupted state, and
+ <application>pg_checksums</application> refuses to run in any mode.
+ Start the cluster and shut it down cleanly to reset the state, then
+ retry; on a standby, let replication complete the transition first.
</para>
<para>
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 5d6ea318784..3b58c6ca608 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -586,6 +586,21 @@ main(int argc, char *argv[])
ControlFile->state != DB_SHUTDOWNED_IN_RECOVERY)
pg_fatal("cluster must be shut down");
+ /*
+ * An inprogress state means an online transition was cut short. A
+ * standby stopped mid-transition carries either state; a cleanly shut
+ * down primary can still carry inprogress-off, which a fast shutdown
+ * during pg_disable_data_checksums() leaves behind, while inprogress-on
+ * is always resolved by the launcher's exit cleanup.
+ */
+ if (ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ ControlFile->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ {
+ pg_log_error("an online data checksum state transition was interrupted");
+ pg_log_error_hint("Start and cleanly shut down the cluster once to reset the data checksum state, then retry. On a standby, let replication complete the transition first.");
+ exit(1);
+ }
+
if (ControlFile->data_checksum_version != PG_DATA_CHECKSUM_VERSION &&
mode == PG_MODE_CHECK)
pg_fatal("data checksums are not enabled in cluster");
@@ -673,6 +688,12 @@ main(int argc, char *argv[])
printf(_("Checksums enabled in cluster\n"));
else
printf(_("Checksums disabled in cluster\n"));
+
+ printf(_("This change applies to this data directory only.\n"));
+ if (ControlFile->state == DB_SHUTDOWNED_IN_RECOVERY)
+ printf(_("This node appears to be a standby; apply the same change to the primary and all other standbys.\n"));
+ else
+ printf(_("In a replication setup, apply the same change to every node while all are stopped, before restarting any of them.\n"));
}
return 0;
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
index c0ec374c8be..ba1f14e0399 100644
--- a/src/test/modules/test_checksums/t/012_offline_standby.pl
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -164,7 +164,10 @@ $newnode->stop('immediate');
# Converge the cluster: enable offline on the standby too.
$standby->stop;
-$standby->checksum_enable_offline;
+command_checks_all(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ 0, [qr/appears to be a standby/],
+ [], 'standby-role notice on offline enable');
$standby->start;
test_checksum_state($standby, 'on');
$primary->wait_for_catchup($standby);
@@ -232,11 +235,11 @@ unlike(
);
# Scenario 4: a standby stopped while replaying an interrupted online
-# transition keeps the interrupted state in its own control file. A
-# primary is never caught this way, as its checksums launcher resolves
-# inprogress-on back to off from its exit cleanup. A standby has no
-# launcher; it carries forward whatever the last replayed record left
-# it in.
+# transition keeps the interrupted state in its own control file, and
+# pg_checksums must refuse to touch it. A primary is never caught this
+# way, as its checksums launcher resolves inprogress-on back to off from
+# its exit cleanup. A standby has no launcher; it carries forward
+# whatever the last replayed record left it in.
# Block an online enable on the primary at inprogress-on with a
# blocking temp table, same trick as in 004_offline.pl.
@@ -250,6 +253,20 @@ wait_for_checksum_state($standby, 'inprogress-on');
# Stop the standby cleanly; its restartpoint persists inprogress-on to
# its own control file, since nothing on a standby resolves it away.
$standby->stop;
+
+command_fails_like(
+ [ 'pg_checksums', '--enable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --enable refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --check refuses a standby stopped mid-transition');
+command_fails_like(
+ [ 'pg_checksums', '--disable', '-D', $standby->data_dir ],
+ qr/online data checksum state transition was interrupted/,
+ 'pg_checksums --disable refuses a standby stopped mid-transition');
+
$standby->start;
wait_for_checksum_state($standby, 'inprogress-on');
@@ -272,7 +289,7 @@ wait_for_checksum_state($primary, 'on');
$primary->wait_for_catchup($standby);
wait_for_checksum_state($standby, 'on');
-is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
'10001', 'standby readable once the transition completes');
# Scenario 4 continued: the backup taken mid-transition must also
--
2.39.3 (Apple Git-146)
=
[application/octet-stream] v16-0001-Do-not-adopt-data-checksum-state-from-another-no.patch (137.0K, ../../0E4734FA-4764-405C-9670-C45DF440A718@yesql.se/10-v16-0001-Do-not-adopt-data-checksum-state-from-another-no.patch)
download | inline diff:
From 7cbd6adadbc3360a3f191755ffc35a9ef7fb3cd8 Mon Sep 17 00:00:00 2001
From: Daniel Gustafsson <dgustafsson@postgresql.org>
Date: Thu, 3 Sep 2026 23:29:17 +0200
Subject: [PATCH v16 1/9] Do not adopt data checksum state from another node
during replay
Offline pg_checksums changes are local to one node, but replay adopted
the data checksum state carried by checkpoint records unconditionally.
After an offline enable on the primary, a standby whose pages were
never rewritten started verifying checksums it does not have; after an
offline change on a standby, the next replayed checkpoint silently
reverted it.
To fix, make the control file track this node's state alone, and have
replay cross-check the replayed state against it instead of adopting
it, warning once per divergent value and reporting when the states
agree again. pg_control gains a watermark, normally the end LSN of
the newest XLOG2_CHECKSUMS record the node has written or applied, so
that replay can skip transition records whose effect the control file
already contains, and a flag marking a state last written by
pg_checksums, which recovery must never overwrite with a replayed one.
Since the control file may only claim "on" once every page on disk
carries a checksum, persisting the state is tied to flushes:
XLOG2_CHECKSUMS replay persists every state but "on" immediately and
leaves "on" to the next restartpoint, checkpoints and restartpoints
only persist a state their flush ran under from beginning to end, and
transitions publish their state under the new
DataChecksumTransitionLock so that the states carried by WAL records
match their WAL order. The comments in xlog.c spell out the
individual rules.
Recovery from a base backup may be an exception to not adopting: its
control file was copied at an arbitrary moment, so the state carried
by the starting checkpoint is the one the WAL from there on was
written under. That holds for a backup taken from a primary, and only
when the copied state is not node local and its watermark is below the
starting checkpoint; a backup taken from a standby keeps the copied
state. pg_rewind keeps the target's own state and clamps the
watermark to the divergence point, since a watermark set on the
target's abandoned fork could numerically cover transition records the
source wrote after the divergence.
Bump PG_CONTROL_VERSION.
Document the offline procedure for replication setups, the lockstep
one: stop all nodes, run pg_checksums on each of them, and only then
restart. An offline change writes no WAL and has no ordering against
WAL a node has not replayed yet, so a node must not stop before
replaying all WAL of its upstream, or the change is overridden on
restart.
Author: Zsolt Parragi <zsolt.parragi@percona.com>
Author: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reported-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Discussion: https://postgr.es/m/anwm6UPxoVS41QA2@bdtpg
---
doc/src/sgml/ref/pg_checksums.sgml | 87 ++-
doc/src/sgml/wal.sgml | 23 +
src/backend/access/transam/xlog.c | 571 ++++++++++++++++--
src/backend/postmaster/datachecksum_state.c | 26 +-
.../utils/activity/wait_event_names.txt | 1 +
src/bin/pg_checksums/pg_checksums.c | 10 +
src/bin/pg_controldata/pg_controldata.c | 4 +
src/bin/pg_resetwal/pg_resetwal.c | 7 +
src/bin/pg_rewind/pg_rewind.c | 35 +-
src/bin/pg_upgrade/controldata.c | 2 +-
src/include/catalog/pg_control.h | 29 +-
src/include/storage/lwlocklist.h | 1 +
src/test/modules/test_checksums/Makefile | 2 +-
src/test/modules/test_checksums/meson.build | 10 +
.../test_checksums/t/012_offline_standby.pl | 304 ++++++++++
.../modules/test_checksums/t/013_rewind.pl | 201 ++++++
.../modules/test_checksums/t/014_lockstep.pl | 182 ++++++
.../t/015_standby_crash_after_disable.pl | 136 +++++
.../t/016_promote_enable_crash.pl | 146 +++++
.../test_checksums/t/017_restartpoint_race.pl | 151 +++++
.../t/018_enable_crash_windows.pl | 508 ++++++++++++++++
.../t/019_standby_shutdown_catchup.pl | 138 +++++
.../t/020_cascade_divergence.pl | 127 ++++
.../t/021_rewind_divergent_transitions.pl | 205 +++++++
24 files changed, 2822 insertions(+), 84 deletions(-)
create mode 100644 src/test/modules/test_checksums/t/012_offline_standby.pl
create mode 100644 src/test/modules/test_checksums/t/013_rewind.pl
create mode 100644 src/test/modules/test_checksums/t/014_lockstep.pl
create mode 100644 src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
create mode 100644 src/test/modules/test_checksums/t/016_promote_enable_crash.pl
create mode 100644 src/test/modules/test_checksums/t/017_restartpoint_race.pl
create mode 100644 src/test/modules/test_checksums/t/018_enable_crash_windows.pl
create mode 100644 src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
create mode 100644 src/test/modules/test_checksums/t/020_cascade_divergence.pl
create mode 100644 src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
diff --git a/doc/src/sgml/ref/pg_checksums.sgml b/doc/src/sgml/ref/pg_checksums.sgml
index ae66fad3f0f..2d2057a8c19 100644
--- a/doc/src/sgml/ref/pg_checksums.sgml
+++ b/doc/src/sgml/ref/pg_checksums.sgml
@@ -242,15 +242,79 @@ PostgreSQL documentation
data directory must not be started or else data loss may occur.
</para>
<para>
- When using a replication setup with tools which perform direct copies
- of relation file blocks (for example <xref linkend="app-pgrewind"/>),
- enabling or disabling checksums can lead to page corruptions in the
- shape of incorrect checksums if the operation is not done consistently
- across all nodes. When enabling or disabling checksums in a replication
- setup, it is thus recommended to stop all the clusters before switching
- them all consistently. Destroying all standbys, performing the operation
- on the primary and finally recreating the standbys from scratch is also
- safe.
+ Enabling or disabling checksums with
+ <application>pg_checksums</application> changes only the local data
+ directory; the new state is not replicated to any other node. In a
+ replication setup the same change must be applied to every node:
+ </para>
+
+ <procedure>
+ <step id="shut-down-nodes">
+ <title>Shut down all nodes</title>
+ <para>
+ All nodes participating in the replication must be stopped with a
+ clean shutdown; <application>pg_checksums</application> refuses to run
+ on a data directory left behind by an immediate shutdown. Before
+ stopping a standby, make sure it has replayed all WAL of the primary:
+ stop the primary first, read its <quote>Latest checkpoint
+ location</quote> with <xref linkend="app-pgcontroldata"/>, and check
+ that <function>pg_last_wal_replay_lsn()</function> on the standby has
+ advanced past it. Comparing the replay position with
+ <function>pg_last_wal_receive_lsn()</function> is not enough, as it
+ only shows that the WAL the standby received has been replayed.
+ </para>
+ </step>
+
+ <step>
+ <title>Enable or disable data checksums on each node</title>
+ <para>
+ Run <application>pg_checksums</application> on the data directory of
+ each node in the replication setup. Nodes can be processed in
+ parallel while they are shut down. Processing must complete
+ successfully on all nodes before continuing.
+ </para>
+ </step>
+
+ <step>
+ <title>Restart all nodes</title>
+ <para>
+ Start the nodes normally, verify that
+ <xref linkend="guc-data-checksums"/> matches on all of them, and
+ monitor the logs of the standbys for data checksum state mismatch
+ warnings.
+ </para>
+ </step>
+ </procedure>
+
+ <para>
+ The replay requirement in <xref linkend="shut-down-nodes"/>exists
+ because an offline change is recorded only in the control file and
+ has no defined ordering against WAL the node has not replayed yet; see
+ <xref linkend="checksums-offline-enable-disable"/>. A node stopped
+ before replaying an online checksum state change applies that change
+ when it is restarted, overriding the offline change, and the states of
+ the nodes silently diverge until a later checkpoint record triggers
+ the warning described below. Because of this it is best not to mix
+ the two mechanisms: change the state of a replication setup either
+ with the offline procedure above or with an online transition, and
+ make sure the previous change has reached every node before starting
+ the next one.
+ </para>
+ <para>
+ If the change is applied inconsistently, each node keeps its own state,
+ and a standby logs a warning when the state recorded in the replayed WAL
+ differs from its own. The same warning can appear transiently while a
+ standby catches up over WAL written before a consistent change; it stops
+ once a checkpoint record carrying the new state has been replayed.
+ </para>
+ <para>
+ A standby whose data directory was never checksummed must not have
+ checksums enabled by catching up this way. Converge the cluster by
+ running <application>pg_checksums</application> on it while stopped as
+ outlined above, by enabling checksums online, or by recreating it from
+ a base backup. Note
+ that an online enable only starts from a primary whose checksums are off,
+ so if they are already enabled there, disable them online first.
</para>
<para>
If <application>pg_checksums</application> is aborted or killed while
@@ -258,6 +322,11 @@ PostgreSQL documentation
remains unchanged, and <application>pg_checksums</application> can be
re-run to perform the same operation.
</para>
+ <para>
+ Tools that copy relation file blocks directly between nodes, such as
+ <xref linkend="app-pgrewind"/>, require all nodes to be in
+ the same data checksum state, else there is risk for data corruption.
+ </para>
<para>
The target cluster must have the same major version as
<application>pg_checksums</application>.
diff --git a/doc/src/sgml/wal.sgml b/doc/src/sgml/wal.sgml
index ec62d17fbcc..dbbfdae5e4f 100644
--- a/doc/src/sgml/wal.sgml
+++ b/doc/src/sgml/wal.sgml
@@ -317,6 +317,29 @@
verify checksums, on an offline cluster.
</para>
+ <para>
+ An offline change provides durability differently from an
+ <link linkend="checksums-online-enable-disable">online change</link>.
+ An online transition is WAL-logged: it is ordered against all other
+ WAL records, it is replayed after a crash, and it propagates to
+ standbys. An offline change is recorded only in the cluster's
+ control file: it writes no WAL, it is invisible to replication, and
+ it has no defined ordering against WAL the node has not replayed
+ yet. When a node later replays WAL that contains an online checksum
+ state change, that change takes effect on the node even if it was
+ written before the offline change was made.
+ </para>
+
+ <para>
+ An offline change only affects the data directory it is run on; the
+ new state does not propagate over replication. In a replication setup
+ the same change must be applied to all nodes while all of them are
+ stopped, as described in <xref linkend="app-pgchecksums"/>. A standby
+ whose state diverges logs a warning but keeps its local setting. The
+ mismatch persists until the states are brought together again, with
+ the offline procedure or with an online transition; do this promptly.
+ </para>
+
</sect2>
<sect2 id="checksums-online-enable-disable" xreflabel="Online Enabling of Checksums">
diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c
index c7a8d64c1f7..d1a8a16bd2d 100644
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -556,14 +556,22 @@ typedef struct XLogCtlData
*/
XLogRecPtr lastFpwDisableRecPtr;
- /* last data_checksum_version we've seen */
+ /* current data checksum state of this node */
uint32 data_checksum_version;
+ /*
+ * Copies of control file fields with the same names, see pg_control.h for
+ * an in-depth description of these fields. Must be updated together with
+ * data_checksum_version under info_lck.
+ */
+ XLogRecPtr data_checksum_lsn;
+ bool data_checksum_is_local;
+
slock_t info_lck; /* locks shared variables shown above */
/*
* lastChecksumChangeRecPtr points to the end of the last XLOG2_CHECKSUMS
- * record inserted or replayed, i.e. the last change of
+ * record inserted or replayed which corresponds to the last change of
* data_checksum_version. InvalidXLogRecPtr if the state hasn't changed
* since the server started.
*/
@@ -690,6 +698,20 @@ static ChecksumStateType LocalDataChecksumState = 0;
*/
int data_checksums = 0;
+/*
+ * Whether replay of the next checkpoint-family record must adopt the data
+ * checksum state it carries. Set when recovery starts from a base backup,
+ * where the state at the redo point takes precedence over the control file
+ * copied with the backup later.
+ */
+static bool adoptChecksumStateFromNextCheckpoint = false;
+
+/*
+ * Sentinel for the data checksum mismatch warning tracking in
+ * CheckReplayedDataChecksumState(): no warning is outstanding.
+ */
+#define NO_WARNING_ISSUED PG_UINT32_MAX
+
/* For WALInsertLockAcquire/Release functions */
static int MyLockNo = 0;
static bool holdingAllLocks = false;
@@ -730,6 +752,8 @@ static void ValidateXLOGDirectoryStructure(void);
static void CleanupBackupHistory(void);
static void UpdateMinRecoveryPoint(XLogRecPtr lsn, bool force);
static bool PerformRecoveryXLogAction(void);
+static void CheckReplayedDataChecksumState(uint32 replayed_version);
+static void AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn);
static void InitControlFile(uint64 sysidentifier, uint32 data_checksum_version);
static void WriteControlFile(void);
static void ReadControlFile(void);
@@ -757,7 +781,7 @@ static void WALInsertLockAcquireExclusive(void);
static void WALInsertLockRelease(void);
static void WALInsertLockUpdateInsertingAt(XLogRecPtr insertingAt);
-static void XLogChecksums(uint32 new_type);
+static XLogRecPtr XLogChecksums(uint32 new_type);
/*
* Insert an XLOG record represented by an already-constructed chain of data
@@ -4292,6 +4316,8 @@ InitControlFile(uint64 sysidentifier, uint32 data_checksum_version)
ControlFile->track_commit_timestamp = track_commit_timestamp;
ControlFile->data_checksum_version = data_checksum_version;
ControlFile->data_checksum_version_init = data_checksum_version;
+ ControlFile->data_checksum_is_local = false;
+ ControlFile->data_checksum_lsn = InvalidXLogRecPtr;
/*
* Set the data_checksum_version value into XLogCtl, which is where all
@@ -4774,6 +4800,7 @@ void
SetDataChecksumsOnInProgress(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
/*
* The state transition is performed in a critical section with
@@ -4782,14 +4809,12 @@ SetDataChecksumsOnInProgress(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_ON);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_ON;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4827,6 +4852,8 @@ void
SetDataChecksumsOn(void)
{
uint64 barrier;
+ bool persist;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4847,23 +4874,11 @@ SetDataChecksumsOn(void)
SpinLockRelease(&XLogCtl->info_lck);
INJECTION_POINT("datachecksums-enable-checksums-delay", NULL);
+ INJECTION_POINT_LOAD("datachecksums-on-before-publish");
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_VERSION);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- SpinLockRelease(&XLogCtl->info_lck);
-
- /*
- * Update the controlfile before waiting since if we have an immediate
- * shutdown while waiting we want to come back up with checksums enabled.
- */
- LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
- UpdateControlFile();
- LWLockRelease(ControlFileLock);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_VERSION);
barrier = EmitProcSignalBarrier(PROCSIGNAL_BARRIER_CHECKSUM_ON);
@@ -4873,6 +4888,41 @@ SetDataChecksumsOn(void)
INJECTION_POINT("datachecksums-on-before-checkpoint", NULL);
RequestCheckpoint(CHECKPOINT_FORCE | CHECKPOINT_WAIT | CHECKPOINT_FAST);
+
+ INJECTION_POINT("datachecksums-on-after-checkpoint", NULL);
+
+ /*
+ * Persist "on" only now that the checkpoint has flushed the pages the
+ * transition rewrote. Crash recovery initializes verification from the
+ * control file but resumes from a checkpoint that can predate the
+ * transition, so an "on" persisted earlier would verify pages during
+ * replay whose rewrite never reached the disk: the pages found by the
+ * data checksums worker in shared buffers are not written back by its
+ * ring buffer.
+ *
+ * The checkpoint above normally persists the state itself; this covers
+ * the case where it started before the record written above and left the
+ * field alone. Crashing before this point is safe, as replay then
+ * re-establishes "on" from the full page images of the rewrite. Skip the
+ * write if the state moved on meanwhile, since whatever moved it persists
+ * its own. Compare the watermark rather than the state: a state
+ * comparison could not tell our transition from a later round trip back
+ * to "on".
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ persist = (XLogCtl->data_checksum_lsn == recptr);
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (persist)
+ {
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = PG_DATA_CHECKSUM_VERSION;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+ }
+
WaitForProcSignalBarrier(barrier);
}
@@ -4893,6 +4943,7 @@ void
SetDataChecksumsOff(void)
{
uint64 barrier;
+ XLogRecPtr recptr;
SpinLockAcquire(&XLogCtl->info_lck);
@@ -4918,14 +4969,12 @@ SetDataChecksumsOff(void)
START_CRIT_SECTION();
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_INPROGRESS_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_INPROGRESS_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -4956,14 +5005,12 @@ SetDataChecksumsOff(void)
/* Ensure that we don't incur a checkpoint during disabling checksums */
MyProc->delayChkptFlags |= DELAY_CHKPT_START;
- XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SpinLockRelease(&XLogCtl->info_lck);
+ recptr = XLogChecksums(PG_DATA_CHECKSUM_OFF);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->data_checksum_version = PG_DATA_CHECKSUM_OFF;
+ ControlFile->data_checksum_lsn = recptr;
+ ControlFile->data_checksum_is_local = false;
UpdateControlFile();
LWLockRelease(ControlFileLock);
@@ -5001,6 +5048,138 @@ SetLocalDataChecksumState(uint32 data_checksum_version)
data_checksums = data_checksum_version;
}
+/*
+ * CheckReplayedDataChecksumState
+ * Cross-check the data checksum state carried by a replayed checkpoint
+ * record against the state of this node.
+ *
+ * The state in the record belongs to whichever node wrote the WAL, and must
+ * not be adopted: an offline state change made with pg_checksums on one node
+ * of a replication set generates no WAL, and adopting would leak it into the
+ * other nodes through replay. Backup label recovery is the one exception,
+ * see AdoptReplayedDataChecksumState(). XLOG_CHECKPOINT_ONLINE needs no call
+ * here, as an online checkpoint's state already traveled in the preceding
+ * XLOG_CHECKPOINT_REDO record.
+ *
+ * Only archive recovery can see a lasting mismatch, as only there can the WAL
+ * and the control file come from different nodes or different times. In
+ * crash recovery a mismatch means replay resumed from a restartpoint
+ * predating an already-applied XLOG2_CHECKSUMS record, and replaying forward
+ * re-establishes the same state.
+ */
+static void
+CheckReplayedDataChecksumState(uint32 replayed_version)
+{
+ /*
+ * Warn once per remote value, so a lasting mismatch does not flood the
+ * log. Matching states re-arm the warning. Backend-local state is
+ * enough: replay only runs in the startup process, and restarting it
+ * re-arms as well.
+ */
+ static uint32 last_warned_version = NO_WARNING_ISSUED;
+ uint32 local_version;
+
+ if (!ArchiveRecoveryRequested)
+ return;
+
+ /*
+ * Re-replayed WAL below the consistency point was already cross-checked
+ * before minRecoveryPoint was last persisted, and the persisted state can
+ * legitimately be newer than what checkpoint records there carry:
+ * XLOG2_CHECKSUMS replay persists most states ahead of the restartpoint
+ * horizon. In particular the checkpoint record recovery restarts from is
+ * such a re-replay.
+ */
+ if (!reachedConsistency)
+ return;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ local_version = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (replayed_version == local_version)
+ {
+ /*
+ * Report convergence if this process warned before. Nothing else
+ * tells the operator that running pg_checksums on the other nodes, or
+ * a rebuild, took effect.
+ */
+ if (last_warned_version != NO_WARNING_ISSUED)
+ ereport(LOG,
+ errmsg("data checksum state \"%s\" of this node now agrees with the replayed WAL",
+ get_checksum_state_string(local_version)));
+
+ last_warned_version = NO_WARNING_ISSUED;
+ return;
+ }
+
+ /* the nodes legitimately differ while an online transition runs */
+ if (replayed_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ replayed_version == PG_DATA_CHECKSUM_INPROGRESS_OFF ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_ON ||
+ local_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
+ return;
+
+ if (replayed_version == last_warned_version)
+ return;
+ last_warned_version = replayed_version;
+
+ ereport(WARNING,
+ errmsg("data checksum state \"%s\" of this node does not match the state \"%s\" in the replayed WAL",
+ get_checksum_state_string(local_version),
+ get_checksum_state_string(replayed_version)),
+ errdetail("The data checksum state may have been changed with pg_checksums on another node."),
+ errhint("Apply the same change with pg_checksums on the primary and all standby servers, or rebuild this server from a base backup."));
+}
+
+/*
+ * AdoptReplayedDataChecksumState
+ * Adopt the data checksum state at the redo point of backup label
+ * recovery.
+ *
+ * The state is persisted immediately so that a crash before the first
+ * restartpoint does not resurrect the state copied with the backup; a crash
+ * at this point restarts from the same redo point, so the control file does
+ * not run ahead of the replay position. If the value is unchanged the
+ * control file already carries it, so both the barrier and the persist are
+ * skipped.
+ *
+ * lsn is the location of the checkpoint-family record the state was taken
+ * from and becomes the new watermark: the adopted state covers everything
+ * below the redo point, which replay never revisits. For the same reason
+ * the old watermark can stay when the value is unchanged: the records
+ * between the two positions are never replayed, so nothing depends on
+ * which one is recorded.
+ */
+static void
+AdoptReplayedDataChecksumState(uint32 new_version, XLogRecPtr lsn)
+{
+ bool changed = false;
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (XLogCtl->data_checksum_version != new_version)
+ {
+ XLogCtl->data_checksum_version = new_version;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(new_version);
+ changed = true;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ if (!changed)
+ return;
+
+ EmitAndWaitDataChecksumsBarrier(new_version);
+
+ LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
+ ControlFile->data_checksum_version = new_version;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ UpdateControlFile();
+ LWLockRelease(ControlFileLock);
+}
+
/* guc hook */
const char *
show_data_checksums(void)
@@ -5454,6 +5633,8 @@ XLOGShmemInit(void *arg)
/* Use the checksum info from control file */
XLogCtl->data_checksum_version = ControlFile->data_checksum_version;
+ XLogCtl->data_checksum_lsn = ControlFile->data_checksum_lsn;
+ XLogCtl->data_checksum_is_local = ControlFile->data_checksum_is_local;
SetLocalDataChecksumState(XLogCtl->data_checksum_version);
SpinLockInit(&XLogCtl->Insert.insertpos_lck);
@@ -6031,6 +6212,51 @@ StartupXLOG(void)
SetCommitTsLimit(checkPoint.oldestCommitTsXid,
checkPoint.newestCommitTsXid);
+ /*
+ * When recovery starts from a base backup, the control file was copied at
+ * an arbitrary moment and its data checksum state may differ from the
+ * state at the redo point, which is what the WAL from there on was
+ * written under. Adopt the state of the starting checkpoint: a shutdown
+ * checkpoint is not replayed, so take it from the record read above; the
+ * redo point of an online checkpoint is its CHECKPOINT_REDO record, so
+ * let the replay of that record adopt it. Check backupStartPoint in
+ * addition to the label: on a crash restart during backup recovery the
+ * label file is already renamed away, but the start point persists until
+ * the backup end record.
+ *
+ * Not for a base backup taken from a standby, though. Its starting
+ * checkpoint is the standby's last restartpoint, a record written by the
+ * upstream primary, whose state is not the one the copied files were
+ * written under. The copied control file is already correct: a standby
+ * persists its state only at restartpoint horizons and never claims more
+ * than what reached disk. Such backups are recognized by backupEndPoint
+ * together with backupEndRequired; backupEndPoint is only set for "BACKUP
+ * FROM: standby" labels and persists across a crash restart. pg_rewind
+ * writes a standby label as well, but no backupEndPoint, and its recovery
+ * keeps adopting: the control file it installs carries the target's own
+ * checksum state, which can lag the redo point of the last common
+ * checkpoint the same way a restartpoint horizon can.
+ *
+ * Never adopt over a state the control file's watermark or local flag
+ * marks as newer than the starting checkpoint. A pg_checksums change is
+ * local to the node and generates no WAL, so nothing in the replayed WAL
+ * could ever restore it once overwritten; and a watermark above the redo
+ * point means the control file already contains the effect of every
+ * transition record up to there.
+ */
+ if ((haveBackupLabel || XLogRecPtrIsValid(ControlFile->backupStartPoint)) &&
+ !(XLogRecPtrIsValid(ControlFile->backupEndPoint) &&
+ ControlFile->backupEndRequired) &&
+ !ControlFile->data_checksum_is_local &&
+ checkPoint.redo > ControlFile->data_checksum_lsn)
+ {
+ if (wasShutdown)
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ checkPoint.redo);
+ else
+ adoptChecksumStateFromNextCheckpoint = true;
+ }
+
/*
* Clear out any old relcache cache files. This is *necessary* if we do
* any WAL replay, since that would probably result in the cache files
@@ -6634,11 +6860,7 @@ StartupXLOG(void)
if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_ON)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
ereport(WARNING,
@@ -6655,11 +6877,7 @@ StartupXLOG(void)
else if (XLogCtl->data_checksum_version == PG_DATA_CHECKSUM_INPROGRESS_OFF)
{
XLogChecksums(PG_DATA_CHECKSUM_OFF);
-
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = PG_DATA_CHECKSUM_OFF;
- SetLocalDataChecksumState(XLogCtl->data_checksum_version);
- SpinLockRelease(&XLogCtl->info_lck);
+ SetLocalDataChecksumState(PG_DATA_CHECKSUM_OFF);
EmitAndWaitDataChecksumsBarrier(PG_DATA_CHECKSUM_OFF);
}
@@ -6684,6 +6902,8 @@ StartupXLOG(void)
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
XLogCtl->SharedRecoveryState = RECOVERY_STATE_DONE;
SpinLockRelease(&XLogCtl->info_lck);
@@ -6815,6 +7035,23 @@ static bool
PerformRecoveryXLogAction(void)
{
bool promoted = false;
+ bool flushForChecksums;
+ uint32 checksum_state;
+
+ /*
+ * The end-of-recovery record persists the data checksum state without
+ * flushing the buffer pool, but the control file may only claim "on" once
+ * every page on disk carries a checksum. If replay entered that state
+ * without a restartpoint following it, the pages rewritten by the
+ * transition are still only in the buffer pool, so take the full
+ * checkpoint below instead of the lightweight record.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ flushForChecksums = (checksum_state == PG_DATA_CHECKSUM_VERSION &&
+ ControlFile->data_checksum_version != checksum_state);
/*
* Perform a checkpoint to update all our recovery activity to disk.
@@ -6830,7 +7067,7 @@ PerformRecoveryXLogAction(void)
* fully out of recovery mode and already accepting queries.
*/
if (ArchiveRecoveryRequested && IsUnderPostmaster &&
- PromoteIsTriggered())
+ PromoteIsTriggered() && !flushForChecksums)
{
promoted = true;
@@ -7437,6 +7674,7 @@ CreateCheckPoint(int flags)
uint32 freespace;
XLogRecPtr PriorRedoPtr;
XLogRecPtr last_important_lsn;
+ XLogRecPtr checksumLsn;
VirtualTransactionId *vxids;
int nvxids;
int oldXLogAllowed = 0;
@@ -7549,11 +7787,14 @@ CreateCheckPoint(int flags)
checkPoint.wal_level = wal_level;
/*
- * Get the current data_checksum_version value from xlogctl, valid at the
- * time of the checkpoint.
+ * Get the current data_checksum_version value from xlogctl. This is
+ * final only for a shutdown checkpoint, where no concurrent transition is
+ * possible; an online checkpoint resamples it together with the redo
+ * record below.
*/
SpinLockAcquire(&XLogCtl->info_lck);
checkPoint.dataChecksumState = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
if (shutdown)
@@ -7610,10 +7851,21 @@ CreateCheckPoint(int flags)
{
xl_checkpoint_redo redo_rec;
+ /*
+ * Sample the data checksum state and insert the redo record under
+ * DataChecksumTransitionLock, so that a concurrent transition cannot
+ * insert its XLOG2_CHECKSUMS record between the sampling and the
+ * insertion below. Without this, the redo record could follow the
+ * transition record in WAL while carrying the pre-transition state,
+ * and recovery resuming here would never learn about the transition.
+ * See XLogChecksums().
+ */
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
WALInsertLockAcquire();
redo_rec.wal_level = wal_level;
SpinLockAcquire(&XLogCtl->info_lck);
redo_rec.data_checksum_version = XLogCtl->data_checksum_version;
+ checksumLsn = XLogCtl->data_checksum_lsn;
SpinLockRelease(&XLogCtl->info_lck);
WALInsertLockRelease();
@@ -7621,6 +7873,14 @@ CreateCheckPoint(int flags)
XLogBeginInsert();
XLogRegisterData(&redo_rec, sizeof(xl_checkpoint_redo));
(void) XLogInsert(RM_XLOG_ID, XLOG_CHECKPOINT_REDO);
+ LWLockRelease(DataChecksumTransitionLock);
+
+ /*
+ * The checkpoint record must carry the same state as the redo record
+ * just inserted: the sample taken before redo determination can be
+ * stale by now, and the pair would otherwise disagree.
+ */
+ checkPoint.dataChecksumState = redo_rec.data_checksum_version;
/*
* XLogInsertRecord will have updated XLogCtl->Insert.RedoRecPtr in
@@ -7824,6 +8084,41 @@ CreateCheckPoint(int flags)
ControlFile->minRecoveryPoint = InvalidXLogRecPtr;
ControlFile->minRecoveryPointTLI = 0;
+ /*
+ * Persist the data checksum state this node runs under into the control
+ * file. Only the top-level field tracks this node, checkPointCopy is a
+ * historical record used to resume replay.
+ *
+ * checkPoint.dataChecksumState was sampled while holding the
+ * DataChecksumTransitionLock together with the redo record, so it is the
+ * state in effect at the redo point. If it was "on", the XLOG2_CHECKSUMS
+ * record announcing that precedes the redo point and every page the
+ * transition rewrote was dirtied before it, so CheckPointGuts() has just
+ * written all of them out. Recording the state here is what keeps a
+ * finished transition from being resolved as interrupted when this
+ * checkpoint is the one crash recovery resumes from: replay never sees
+ * the record announcing it.
+ *
+ * Persist it only if the state did not change while the flush was in
+ * progress. If it changed in between, the pages written out straddle two
+ * states, and the newer one could claim checksums that pages already on
+ * disk do not carry; leave the field to the next checkpoint then, the
+ * transition itself has already persisted every state that is safe
+ * without a flush.
+ *
+ * Compare the watermark rather than the state: record positions are
+ * unique, so a full round trip back to the sampled state cannot alias,
+ * while its flushed pages straddle the intermediate states all the same.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksumLsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checkPoint.dataChecksumState;
+ ControlFile->data_checksum_lsn = checksumLsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
+
/*
* Persist unloggedLSN value. It's reset on crash recovery, so this goes
* unused on non-shutdown checkpoints, but seems useful to store it always
@@ -7968,9 +8263,11 @@ CreateEndOfRecoveryRecord(void)
ControlFile->minRecoveryPoint = recptr;
ControlFile->minRecoveryPointTLI = xlrec.ThisTimeLineID;
- /* start with the latest checksum version (as of the end of recovery) */
+ /* persist the data checksum state this node ended recovery with */
SpinLockAcquire(&XLogCtl->info_lck);
ControlFile->data_checksum_version = XLogCtl->data_checksum_version;
+ ControlFile->data_checksum_lsn = XLogCtl->data_checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
@@ -8175,6 +8472,9 @@ CreateRestartPoint(int flags)
XLogRecPtr endptr;
XLogSegNo _logSegNo;
TimestampTz xtime;
+ uint32 checksum_state;
+ XLogRecPtr checksum_lsn;
+ bool checksum_is_local;
/* Concurrent checkpoint/restartpoint cannot happen */
Assert(!IsUnderPostmaster || MyBackendType == B_CHECKPOINTER);
@@ -8221,8 +8521,47 @@ CreateRestartPoint(int flags)
UpdateMinRecoveryPoint(InvalidXLogRecPtr, true);
if (flags & CHECKPOINT_IS_SHUTDOWN)
{
+ bool catchUpChecksums;
+
+ /*
+ * There is no new restartpoint to persist the data checksum state
+ * with, but a cleanly stopped node should not leave the control
+ * file behind the state replay reached: pg_checksums and
+ * pg_rewind read it, and an in-progress state there makes them
+ * refuse to run. Catching it up needs the same guarantee a
+ * restartpoint gives: that every page on disk carries a checksum,
+ * so flush the buffer pool first. Only a transition to "on" that
+ * no restartpoint followed can get here; the other states are
+ * already persisted by XLOG2_CHECKSUMS replay. Replay has ended
+ * by now, so the state cannot change under us.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ checksum_is_local = XLogCtl->data_checksum_is_local;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockAcquire(ControlFileLock, LW_SHARED);
+ catchUpChecksums =
+ (checksum_lsn != ControlFile->data_checksum_lsn &&
+ XLogRecPtrIsValid(lastCheckPointRecPtr));
+ LWLockRelease(ControlFileLock);
+
+ if (catchUpChecksums)
+ {
+ MemSet(&CheckpointStats, 0, sizeof(CheckpointStats));
+ CheckpointStats.ckpt_start_t = GetCurrentTimestamp();
+ CheckPointGuts(lastCheckPoint.redo, flags);
+ }
+
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
+ if (catchUpChecksums)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = checksum_is_local;
+ }
UpdateControlFile();
LWLockRelease(ControlFileLock);
}
@@ -8263,6 +8602,17 @@ CreateRestartPoint(int flags)
/* Update the process title */
update_checkpoint_display(flags, true, false);
+ /*
+ * Note the data checksum state the flush below starts under. Replay runs
+ * concurrently and can change the state while the flush is in progress,
+ * in which case the flush covers pages written under both states; see
+ * where the state is persisted further down.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ checksum_state = XLogCtl->data_checksum_version;
+ checksum_lsn = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
CheckPointGuts(lastCheckPoint.redo, flags);
/*
@@ -8322,8 +8672,26 @@ CreateRestartPoint(int flags)
ControlFile->state = DB_SHUTDOWNED_IN_RECOVERY;
}
- /* we shall start with the latest checksum version */
- ControlFile->data_checksum_version = lastCheckPoint.dataChecksumState;
+ /*
+ * Persist the data checksum state of this node. Not the state of the
+ * replayed checkpoint: that one belongs to the node that wrote it and
+ * may differ after an offline change on either side.
+ * ControlFile->checkPointCopy above keeps the replayed value on
+ * purpose, being a historical record used to resume replay rather
+ * than a tracker of node state.
+ *
+ * Persist only if the flush above ran under one state throughout; see
+ * CreateCheckPoint() for why, including why this compares the
+ * watermark and not the state.
+ */
+ SpinLockAcquire(&XLogCtl->info_lck);
+ if (checksum_lsn == XLogCtl->data_checksum_lsn)
+ {
+ ControlFile->data_checksum_version = checksum_state;
+ ControlFile->data_checksum_lsn = checksum_lsn;
+ ControlFile->data_checksum_is_local = XLogCtl->data_checksum_is_local;
+ }
+ SpinLockRelease(&XLogCtl->info_lck);
UpdateControlFile();
}
@@ -8764,9 +9132,22 @@ XLogReportParameters(void)
}
/*
- * Log the new state of checksums
+ * XLogChecksums
+ * Log and publish the new state of checksums
+ *
+ * Inserting the record and publishing the new state must be atomic with
+ * respect to a checkpoint sampling the state for its XLOG_CHECKPOINT_REDO
+ * record: without that, a checkpoint could read the old state after the
+ * record is already in WAL and insert a redo record that both follows the
+ * transition in WAL order and carries the pre-transition state. Recovery
+ * resuming from such a redo point would never replay the transition record
+ * and resolve the finished transition as interrupted.
+ * DataChecksumTransitionLock serializes the two; see CreateCheckPoint().
+ *
+ * Returns the end LSN of the inserted record, which the caller persists
+ * together with the new state as the data checksum watermark.
*/
-static void
+static XLogRecPtr
XLogChecksums(uint32 new_type)
{
xl_checksum_state xlrec;
@@ -8774,12 +9155,28 @@ XLogChecksums(uint32 new_type)
xlrec.new_checksum_state = new_type;
+ LWLockAcquire(DataChecksumTransitionLock, LW_EXCLUSIVE);
+
XLogBeginInsert();
XLogRegisterData((char *) &xlrec, sizeof(xl_checksum_state));
recptr = XLogInsert(RM_XLOG2_ID, XLOG2_CHECKSUMS);
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, recptr);
+
+ /* only loaded by SetDataChecksumsOn(), a no-op for the other callers */
+ INJECTION_POINT_CACHED("datachecksums-on-before-publish", NULL);
+
+ SpinLockAcquire(&XLogCtl->info_lck);
+ XLogCtl->data_checksum_version = new_type;
+ XLogCtl->data_checksum_lsn = recptr;
+ XLogCtl->data_checksum_is_local = false;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ LWLockRelease(DataChecksumTransitionLock);
+
XLogFlush(recptr);
+
+ return recptr;
}
/*
@@ -8968,11 +9365,19 @@ xlog_redo(XLogReaderState *record)
/* ControlFile->checkPointCopy always tracks the latest ckpt XID */
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
ControlFile->checkPointCopy.nextXid = checkPoint.nextXid;
- ControlFile->data_checksum_version = checkPoint.dataChecksumState;
UpdateControlFile();
LWLockRelease(ControlFileLock);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(checkPoint.dataChecksumState,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(checkPoint.dataChecksumState);
+
/*
* We should've already switched to the new TLI before replaying this
* record.
@@ -9207,19 +9612,17 @@ xlog_redo(XLogReaderState *record)
else if (info == XLOG_CHECKPOINT_REDO)
{
xl_checkpoint_redo redo_rec;
- bool new_state = false;
memcpy(&redo_rec, XLogRecGetData(record), sizeof(xl_checkpoint_redo));
- SpinLockAcquire(&XLogCtl->info_lck);
- XLogCtl->data_checksum_version = redo_rec.data_checksum_version;
- SetLocalDataChecksumState(redo_rec.data_checksum_version);
- if (redo_rec.data_checksum_version != ControlFile->data_checksum_version)
- new_state = true;
- SpinLockRelease(&XLogCtl->info_lck);
-
- if (new_state)
- EmitAndWaitDataChecksumsBarrier(redo_rec.data_checksum_version);
+ if (adoptChecksumStateFromNextCheckpoint)
+ {
+ adoptChecksumStateFromNextCheckpoint = false;
+ AdoptReplayedDataChecksumState(redo_rec.data_checksum_version,
+ record->ReadRecPtr);
+ }
+ else
+ CheckReplayedDataChecksumState(redo_rec.data_checksum_version);
}
else if (info == XLOG_LOGICAL_DECODING_STATUS_CHANGE)
{
@@ -9281,25 +9684,63 @@ xlog2_redo(XLogReaderState *record)
{
xl_checksum_state state;
XLogRecPtr lsn = record->EndRecPtr;
+ XLogRecPtr watermark;
memcpy(&state, XLogRecGetData(record), sizeof(xl_checksum_state));
+ SpinLockAcquire(&XLogCtl->info_lck);
+ watermark = XLogCtl->data_checksum_lsn;
+ SpinLockRelease(&XLogCtl->info_lck);
+
+ /*
+ * Skip records this node has already applied. The control file
+ * carries the watermark, so this holds across restarts: recovery
+ * resuming below a record whose effect the control file already
+ * contains must not re-apply it, or it would revert a state change
+ * made with pg_checksums in between, which moves the state without
+ * writing any record of its own.
+ */
+ if (lsn <= watermark)
+ return;
+
/* advertise the location before the new state becomes visible */
pg_atomic_write_u64(&XLogCtl->lastChecksumChangeRecPtr, lsn);
SpinLockAcquire(&XLogCtl->info_lck);
XLogCtl->data_checksum_version = state.new_checksum_state;
+ XLogCtl->data_checksum_lsn = lsn;
+ XLogCtl->data_checksum_is_local = false;
+ SetLocalDataChecksumState(state.new_checksum_state);
SpinLockRelease(&XLogCtl->info_lck);
LWLockAcquire(ControlFileLock, LW_EXCLUSIVE);
- ControlFile->data_checksum_version = state.new_checksum_state;
+
+ /*
+ * Persist the new state, except when it is "on". Only "on" verifies
+ * checksums during reads, and between the last restartpoint and this
+ * record there may be pages on disk flushed under the old state; a
+ * crash-restart initializes verification from the control file and
+ * replay reads those pages back, so the control file may only say
+ * "on" once everything written under the transition has been flushed,
+ * as restartpoints and the end of recovery do. The opposite direction
+ * cannot wait for the restartpoint: once this record is replayed,
+ * evicted pages are written without checksums, and a control file
+ * still saying "on" would fail verification on exactly those pages
+ * after a crash.
+ */
+ if (state.new_checksum_state != PG_DATA_CHECKSUM_VERSION)
+ {
+ ControlFile->data_checksum_version = state.new_checksum_state;
+ ControlFile->data_checksum_lsn = lsn;
+ ControlFile->data_checksum_is_local = false;
+ }
/*
* Update minRecoveryPoint to ensure that if recovery is aborted, we
* recover back up to this point before allowing hot standby again.
- * The new state is durable in pg_control while its location is only
- * tracked in shared memory; a standby becoming consistent below this
- * record would let base backups resume checksum verification with the
+ * The change location is only tracked in shared memory and is lost
+ * over a restart; a standby becoming consistent below this record
+ * would let base backups resume checksum verification with the
* location unknown. The local copies cannot be updated as long as
* crash recovery is happening and we expect all the WAL to be
* replayed.
diff --git a/src/backend/postmaster/datachecksum_state.c b/src/backend/postmaster/datachecksum_state.c
index f69258bc33d..099a6b4fe2e 100644
--- a/src/backend/postmaster/datachecksum_state.c
+++ b/src/backend/postmaster/datachecksum_state.c
@@ -94,9 +94,9 @@
*
* If processing is started in an online cluster then all backends are in Bd.
* If processing was halted by the cluster shutting down (due to a crash or
- * intentional restart), the controlfile state "inprogress-on" will be observed
- * on system startup and all backends will be placed in Bd. The controlfile
- * state will also be set to "off".
+ * intentional restart), the control file state "inprogress-on" will be
+ * observed on system startup and all backends will be placed in Bd. The
+ * control file state will also be set to "off".
*
* Backends transition Bd -> Bi via a procsignalbarrier which is emitted by the
* DataChecksumsWorkerLauncherMain. When all backends have acknowledged the
@@ -146,7 +146,25 @@
* stop writing data checksums as no backend is enforcing data checksum
* validation any longer.
*
- * 4. Future opportunities for optimizations
+ * 4. Interaction with offline data checksum changes
+ * -------------------------------------------------
+ * Enabling or disabling checksums offline with pg_checksums uses none of the
+ * machinery in this file, but the two mechanisms share the state kept in the
+ * control file, so their interaction is documented here.
+ *
+ * pg_checksums writes the new state to the control file and sets
+ * data_checksum_is_local, marking a state that no WAL record accounts for.
+ * Recovery then does not adopt the state carried by a replayed checkpoint
+ * record over it. The control file also carries a watermark, the WAL
+ * position through which data checksum transitions are covered. Replay skips
+ * transition records ending at or below the watermark, as their effect is
+ * already contained in the control file, and applies records above it as
+ * usual, whether they were written before or after an offline change. This
+ * is why an offline change in a replicated setup must be made on every node
+ * while all of them are stopped and caught up; see the pg_checksums
+ * documentation for the procedure.
+ *
+ * 5. Future opportunities for optimizations
* -----------------------------------------
* Below are some potential optimizations and improvements which were brought
* up during reviews of this feature, but which weren't implemented in the
diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt
index 0a70cffa081..3d366fd1114 100644
--- a/src/backend/utils/activity/wait_event_names.txt
+++ b/src/backend/utils/activity/wait_event_names.txt
@@ -371,6 +371,7 @@ WaitLSN "Waiting to read or update shared Wait-for-LSN state."
LogicalDecodingControl "Waiting to read or update logical decoding status information."
DataChecksumsWorker "Waiting for data checksums worker."
AioWorkerControl "Waiting to update AIO worker information."
+DataChecksumTransition "Waiting for a data checksum state transition to be written to WAL."
#
# END OF PREDEFINED LWLOCKS (DO NOT CHANGE THIS LINE)
diff --git a/src/bin/pg_checksums/pg_checksums.c b/src/bin/pg_checksums/pg_checksums.c
index 3b3ae23f1a6..5d6ea318784 100644
--- a/src/bin/pg_checksums/pg_checksums.c
+++ b/src/bin/pg_checksums/pg_checksums.c
@@ -648,6 +648,16 @@ main(int argc, char *argv[])
ControlFile->data_checksum_version =
(mode == PG_MODE_ENABLE) ? PG_DATA_CHECKSUM_VERSION : PG_DATA_CHECKSUM_OFF;
+ /*
+ * Mark the state as changed locally, without a WAL record. Recovery
+ * then does not let a replayed checkpoint overwrite it, as no record
+ * could restore the change afterwards. The watermark is left alone:
+ * XLOG2_CHECKSUMS records at or below it stay covered, while records
+ * above it, which this node has not applied yet, still take effect on
+ * replay no matter when they were written.
+ */
+ ControlFile->data_checksum_is_local = true;
+
if (do_sync)
{
pg_log_info("syncing data directory");
diff --git a/src/bin/pg_controldata/pg_controldata.c b/src/bin/pg_controldata/pg_controldata.c
index b785f7f4070..33f0e9e4ea9 100644
--- a/src/bin/pg_controldata/pg_controldata.c
+++ b/src/bin/pg_controldata/pg_controldata.c
@@ -349,6 +349,10 @@ main(int argc, char *argv[])
(ControlFile->float8ByVal ? _("by value") : _("by reference")));
printf(_("Data page checksum version: %u\n"),
ControlFile->data_checksum_version);
+ printf(_("Data checksum watermark: %X/%08X\n"),
+ LSN_FORMAT_ARGS(ControlFile->data_checksum_lsn));
+ printf(_("Data checksum state is node-local: %s\n"),
+ (ControlFile->data_checksum_is_local ? _("yes") : _("no")));
printf(_("Default char data signedness: %s\n"),
(ControlFile->default_char_signedness ? _("signed") : _("unsigned")));
printf(_("Mock authentication nonce: %s\n"),
diff --git a/src/bin/pg_resetwal/pg_resetwal.c b/src/bin/pg_resetwal/pg_resetwal.c
index 63e4381e03f..b11f76a401e 100644
--- a/src/bin/pg_resetwal/pg_resetwal.c
+++ b/src/bin/pg_resetwal/pg_resetwal.c
@@ -923,6 +923,13 @@ RewriteControlFile(void)
ControlFile.backupEndPoint = InvalidXLogRecPtr;
ControlFile.backupEndRequired = false;
+ /*
+ * The old WAL is gone and the new position may lie below the old
+ * watermark, which would make replay ignore future checksum transition
+ * records. The state itself is kept.
+ */
+ ControlFile.data_checksum_lsn = InvalidXLogRecPtr;
+
/*
* Force the defaults for max_* settings. The values don't really matter
* as long as wal_level='minimal'; the postmaster will reset these fields
diff --git a/src/bin/pg_rewind/pg_rewind.c b/src/bin/pg_rewind/pg_rewind.c
index 2e86fd158d0..0e4c2df3f4b 100644
--- a/src/bin/pg_rewind/pg_rewind.c
+++ b/src/bin/pg_rewind/pg_rewind.c
@@ -37,7 +37,8 @@ static void usage(const char *progname);
static void perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo);
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec);
static void createBackupLabel(XLogRecPtr startpoint, TimeLineID starttli,
XLogRecPtr checkpointloc);
@@ -531,7 +532,8 @@ main(int argc, char **argv)
* This is the point of no return. Once we start copying things, there is
* no turning back!
*/
- perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo);
+ perform_rewind(filemap, source, chkptrec, chkpttli, chkptredo,
+ divergerec);
if (showprogress)
pg_log_info("syncing target data directory");
@@ -566,7 +568,8 @@ static void
perform_rewind(filemap_t *filemap, rewind_source *source,
XLogRecPtr chkptrec,
TimeLineID chkpttli,
- XLogRecPtr chkptredo)
+ XLogRecPtr chkptredo,
+ XLogRecPtr divergerec)
{
XLogRecPtr endrec;
TimeLineID endtli;
@@ -738,6 +741,32 @@ perform_rewind(filemap_t *filemap, rewind_source *source,
ControlFile_new.minRecoveryPoint = endrec;
ControlFile_new.minRecoveryPointTLI = endtli;
ControlFile_new.state = DB_IN_ARCHIVE_RECOVERY;
+
+ /*
+ * Keep the target's own data checksum state. Most of the data directory
+ * is still the target's: only blocks it changed since the divergence were
+ * copied from the source, so the source's state says nothing about the
+ * pages that stay. Replay from the last common checkpoint applies any
+ * WAL-logged transition the target has not seen (the watermark tells them
+ * apart), which converges the rewound server to the source's state
+ * whenever the WAL carries it.
+ */
+ ControlFile_new.data_checksum_version = ControlFile_target.data_checksum_version;
+ ControlFile_new.data_checksum_lsn = ControlFile_target.data_checksum_lsn;
+ ControlFile_new.data_checksum_is_local = ControlFile_target.data_checksum_is_local;
+
+ /*
+ * The watermark is only meaningful within the history the node replays.
+ * Records at or below the divergence point are common to both histories
+ * and stay covered, but a watermark above it was set by a transition
+ * record on the target's own abandoned fork: numerically it can cover
+ * transition records the source wrote after the divergence, and replay
+ * would skip them as already applied. Clamp it to the divergence point,
+ * so that every transition record on the source's history takes effect.
+ */
+ if (ControlFile_new.data_checksum_lsn > divergerec)
+ ControlFile_new.data_checksum_lsn = divergerec;
+
if (!dry_run)
update_controlfile(datadir_target, &ControlFile_new, do_sync);
}
diff --git a/src/bin/pg_upgrade/controldata.c b/src/bin/pg_upgrade/controldata.c
index f0fa2b1689f..759bd4ef77c 100644
--- a/src/bin/pg_upgrade/controldata.c
+++ b/src/bin/pg_upgrade/controldata.c
@@ -431,7 +431,7 @@ get_control_data(ClusterInfo *cluster)
cluster->controldata.date_is_int = strstr(p, "64-bit integers") != NULL;
got_date_is_int = true;
}
- else if ((p = strstr(bufin, "checksum")) != NULL)
+ else if ((p = strstr(bufin, "Data page checksum version:")) != NULL)
{
p = strchr(p, ':');
diff --git a/src/include/catalog/pg_control.h b/src/include/catalog/pg_control.h
index 89ab43dd4fc..c3c934d0012 100644
--- a/src/include/catalog/pg_control.h
+++ b/src/include/catalog/pg_control.h
@@ -22,7 +22,7 @@
/* Version identifier for this pg_control format */
-#define PG_CONTROL_VERSION 2000
+#define PG_CONTROL_VERSION 2001
/* Nonce key length, see below */
#define MOCK_AUTH_NONCE_LEN 32
@@ -237,6 +237,33 @@ typedef struct ControlFileData
/* Current data checksums state */
uint32 data_checksum_version;
+ /*
+ * WAL position through which data checksum transitions are covered.
+ * Replay ignores XLOG2_CHECKSUMS records ending at or below this point:
+ * their effect is already contained in data_checksum_version, or an
+ * offline pg_checksums change made after they were first applied
+ * supersedes them. Ordinarily this is the end of the newest such record
+ * this node has written or applied, but a tool may store any position
+ * that covers the same set of records. If the node has never written or
+ * applied such a record this field shall be set to InvalidXLogRecPtr.
+ *
+ * The comparison has no timeline context, so the value is only valid
+ * within the WAL history this node replays. A tool that moves the node
+ * to another history must clamp the watermark to the point where the
+ * histories fork, as pg_rewind does, or reset it, as pg_resetwal does.
+ */
+ XLogRecPtr data_checksum_lsn;
+
+ /*
+ * True when data_checksum_version was last set by pg_checksums in an
+ * offline operation rather than by an online, WAL-logged transition. Such
+ * a state is local to this node and not derived from WAL, so recovery
+ * must not replace it with a state taken from a checkpoint record;
+ * nothing in the WAL could restore the change once it is overwritten.
+ * Cleared by the next WAL-logged transition.
+ */
+ bool data_checksum_is_local;
+
/*
* True if the default signedness of char is "signed" on a platform where
* the cluster is initialized.
diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h
index d7eb648bd27..8d858be9927 100644
--- a/src/include/storage/lwlocklist.h
+++ b/src/include/storage/lwlocklist.h
@@ -89,6 +89,7 @@ PG_LWLOCK(54, WaitLSN)
PG_LWLOCK(55, LogicalDecodingControl)
PG_LWLOCK(56, DataChecksumsWorker)
PG_LWLOCK(57, AioWorkerControl)
+PG_LWLOCK(58, DataChecksumTransition)
/*
* There also exist several built-in LWLock tranches. As with the predefined
diff --git a/src/test/modules/test_checksums/Makefile b/src/test/modules/test_checksums/Makefile
index 71455cd5577..80f54bf6d8c 100644
--- a/src/test/modules/test_checksums/Makefile
+++ b/src/test/modules/test_checksums/Makefile
@@ -9,7 +9,7 @@
#
#-------------------------------------------------------------------------
-EXTRA_INSTALL = src/test/modules/injection_points
+EXTRA_INSTALL = contrib/pg_buffercache src/test/modules/injection_points
export enable_injection_points
diff --git a/src/test/modules/test_checksums/meson.build b/src/test/modules/test_checksums/meson.build
index fb7129d796f..e5d38fafb7d 100644
--- a/src/test/modules/test_checksums/meson.build
+++ b/src/test/modules/test_checksums/meson.build
@@ -35,6 +35,16 @@ tests += {
't/009_fpi.pl',
't/010_backup_straddle.pl',
't/011_standby_straddle.pl',
+ 't/012_offline_standby.pl',
+ 't/013_rewind.pl',
+ 't/014_lockstep.pl',
+ 't/015_standby_crash_after_disable.pl',
+ 't/016_promote_enable_crash.pl',
+ 't/017_restartpoint_race.pl',
+ 't/018_enable_crash_windows.pl',
+ 't/019_standby_shutdown_catchup.pl',
+ 't/020_cascade_divergence.pl',
+ 't/021_rewind_divergent_transitions.pl',
],
},
}
diff --git a/src/test/modules/test_checksums/t/012_offline_standby.pl b/src/test/modules/test_checksums/t/012_offline_standby.pl
new file mode 100644
index 00000000000..c0ec374c8be
--- /dev/null
+++ b/src/test/modules/test_checksums/t/012_offline_standby.pl
@@ -0,0 +1,304 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Offline checksum changes with pg_checksums are local to one node. A
+# standby must neither adopt the state of the primary from replayed
+# checkpoint records, nor lose its own offline change to them.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Scenario 1: enable offline on the primary only. The standby must
+# stay off, warn about the mismatch, and remain readable.
+$standby->stop;
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+test_checksum_state($primary, 'on');
+test_checksum_state($standby, 'off');
+
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after offline enable on the primary');
+
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+my @warnings = $log =~ /(does not match the state)/g;
+is(scalar(@warnings), 1, 'mismatch warned once per remote value');
+
+# Matching states re-arm the warning: undo the divergence on the primary,
+# then diverge again to the same value, all without restarting the standby.
+$primary->stop;
+$primary->checksum_disable_offline;
+$logstart = -s $standby->logfile;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no warning while the states match again');
+like($log, qr/now agrees with the replayed WAL/,
+ 'convergence is reported once the states match again');
+
+$primary->stop;
+$primary->checksum_enable_offline;
+$primary->start;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $logstart);
+test_checksum_state($standby, 'off');
+
+# The local state survives both clean and immediate restarts.
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+# Still in scenario 1's divergence: take a base backup *from the
+# standby*. A backup taken on a standby uses the last restartpoint as
+# its starting checkpoint (do_pg_backup_start()), so the record at the
+# redo point was written by the upstream primary and carries the
+# primary's state. Recovery from such a backup must not adopt that
+# state: the files were copied from the standby, and the new node
+# would come up verifying checksums its files do not have.
+
+# Make sure the standby has written pages under its own "off" state, so
+# its files really do lack checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE t2 AS SELECT generate_series(1,50000) AS a;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', "SELECT count(*) FROM t2;");
+
+$standby->backup('from_standby');
+my $newnode = PostgreSQL::Test::Cluster->new('newnode');
+$newnode->init_from_backup($standby, 'from_standby');
+
+# Stream from the primary, not from the standby the files were copied
+# from, so the new node replays the "on" primary's records.
+$newnode->enable_streaming($primary);
+$newnode->start;
+
+# The new node's files all came from a cluster running with checksums
+# off. Anything but "off" here means it adopted the primary's state
+# through the checkpoint record at the redo point.
+my ($rc, $stdout, $stderr) = $newnode->psql('postgres',
+ "SELECT setting FROM pg_settings WHERE name = 'data_checksums';");
+is($rc, 0, 'the node copied from the standby accepts connections')
+ or diag("stderr: $stderr");
+is($stdout, 'off', 'backup of an "off" standby comes up with checksums off');
+
+# A checkpoint record streamed from the "on" primary must not flip the
+# state either.
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($newnode);
+test_checksum_state($newnode, 'off');
+
+# And it must be able to read the pages the standby wrote without
+# checksums.
+(undef, $stdout, $stderr) =
+ $newnode->psql('postgres', "SELECT count(*) FROM t2;");
+is($stdout, '50000', 'pages copied from the standby are readable')
+ or diag("stderr: $stderr");
+
+my $newnode_log = PostgreSQL::Test::Utils::slurp_file($newnode->logfile);
+unlike(
+ $newnode_log,
+ qr/page verification failed/,
+ 'no checksum verification failures on the node copied from the standby');
+
+$newnode->stop('immediate');
+
+# Converge the cluster: enable offline on the standby too.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+$primary->wait_for_catchup($standby);
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable after converging');
+
+# Scenario 2: disable offline on the standby only. The replayed
+# checkpoint records of the still-enabled primary must not override it.
+$standby->stop;
+$standby->checksum_disable_offline;
+$standby->start;
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'off');
+
+# Restartpoints must persist the local state, not the replayed copy.
+$standby->safe_psql('postgres', "CHECKPOINT;");
+$standby->restart;
+test_checksum_state($standby, 'off');
+$standby->stop('immediate');
+$standby->start;
+test_checksum_state($standby, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable with checksums disabled locally');
+
+# Scenario 3: crash-restart right after an online transition, before the
+# next restartpoint. Replay then resumes from an older restartpoint whose
+# checkpoint records still carry the pre-transition state. Those must
+# still match the state seeded from the control file, and the transition
+# itself must be re-established by re-replaying the XLOG2_CHECKSUMS record,
+# without a spurious mismatch warning along the way.
+
+# Converge first: bring the standby back to "on" offline.
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+disable_data_checksums($primary, wait => 'off');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+# Crash-restart the standby immediately, before any restartpoint has had a
+# chance to persist the new state to its control file.
+$logstart = -s $standby->logfile;
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001',
+ 'standby readable after crash-restart across an online transition');
+
+$log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no spurious mismatch warning after crash-restart across an online transition'
+);
+
+# Scenario 4: a standby stopped while replaying an interrupted online
+# transition keeps the interrupted state in its own control file. A
+# primary is never caught this way, as its checksums launcher resolves
+# inprogress-on back to off from its exit cleanup. A standby has no
+# launcher; it carries forward whatever the last replayed record left
+# it in.
+
+# Block an online enable on the primary at inprogress-on with a
+# blocking temp table, same trick as in 004_offline.pl.
+my $bsession = $primary->background_psql('postgres');
+$bsession->query_safe('CREATE TEMPORARY TABLE tt (a integer);');
+enable_data_checksums($primary, wait => 'inprogress-on');
+
+# The standby picks up the in-progress state from the XLOG2_CHECKSUMS record.
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# Stop the standby cleanly; its restartpoint persists inprogress-on to
+# its own control file, since nothing on a standby resolves it away.
+$standby->stop;
+$standby->start;
+wait_for_checksum_state($standby, 'inprogress-on');
+
+# The primary still sits at inprogress-on, so a base backup taken now
+# copies a control file and a redo point that both carry the
+# in-progress state; replay of the XLOG2_CHECKSUMS records completes
+# the transition on the new standby.
+$primary->backup('inprogress_backup');
+my $standby2 = PostgreSQL::Test::Cluster->new('standby2');
+$standby2->init_from_backup($primary, 'inprogress_backup',
+ has_streaming => 1);
+$standby2->start;
+
+# Backup label recovery adopts the redo point's state immediately,
+# before any further WAL is replayed.
+wait_for_checksum_state($standby2, 'inprogress-on');
+
+$bsession->quit;
+wait_for_checksum_state($primary, 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+is( $standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'standby readable once the transition completes');
+
+# Scenario 4 continued: the backup taken mid-transition must also
+# complete the transition and stay readable.
+$primary->wait_for_catchup($standby2);
+wait_for_checksum_state($standby2, 'on');
+
+is($standby2->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'mid-transition backup standby readable after the transition');
+
+$standby2->stop;
+
+# Nothing along this standby's replay can legitimately disagree: it
+# started out at the same inprogress-on state as the primary.
+my $standby2_log = PostgreSQL::Test::Utils::slurp_file($standby2->logfile);
+unlike(
+ $standby2_log,
+ qr/does not match the state/,
+ 'no mismatch warning on a backup taken mid-transition');
+
+# No extra CHECKPOINT needed: the transition's completion checkpoint
+# wrote every page with a checksum, and the shutdown restartpoint
+# flushes the rest.
+command_ok([ 'pg_checksums', '--check', '-D', $standby2->data_dir ],
+ 'checksums valid on the mid-transition backup standby');
+
+$standby->stop;
+$primary->stop;
+done_testing();
diff --git a/src/test/modules/test_checksums/t/013_rewind.pl b/src/test/modules/test_checksums/t/013_rewind.pl
new file mode 100644
index 00000000000..a791e24317d
--- /dev/null
+++ b/src/test/modules/test_checksums/t/013_rewind.pl
@@ -0,0 +1,201 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test pg_rewind across an online data checksum enable.
+#
+# A clean switchover leaves the shutdown checkpoint of the old primary
+# as the last common checkpoint between the two nodes. When data
+# checksums are enabled online on the new primary before the old one is
+# rewound, pg_rewind installs the control file of the new primary, which
+# already claims checksums are fully enabled, while replay begins at the
+# shutdown checkpoint whose record still carries the old state. Replay
+# of the WAL stretch from before the enable must run with checksums off,
+# as recorded in the checkpoint, else it would verify pages which never
+# had checksums written and fail recovery.
+#
+# The new primary requests a checkpoint right after promotion, and
+# replaying its CHECKPOINT_REDO record would repair the state before
+# any interesting WAL is reached. Hold the checkpointer on the standby
+# in a restartpoint over the promotion, like in the recovery test
+# 041_checkpoint_at_promote, so that the post-promotion writes end up
+# in WAL before the first checkpoint of the new timeline.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Old primary. full_page_writes is off so that the updates done on the
+# promoted node do not carry full page images, forcing replay on the
+# rewound node to read the pages from disk. wal_log_hints is required
+# by pg_rewind on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+full_page_writes = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+log_checkpoints = on
+]);
+$node_a->start;
+
+if (!$node_a->check_extension('injection_points'))
+{
+ plan skip_all => 'Extension injection_points not installed';
+}
+
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+$node_a->safe_psql('postgres', "CREATE EXTENSION injection_points;");
+
+# Set the hint bits on t before taking the backup, so that reads on the
+# promoted node do not emit full page images for its pages later.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->safe_psql('postgres', "SELECT count(*) FROM t;");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_a, 'off');
+test_checksum_state($node_b, 'off');
+
+# Hold the next restartpoint on the standby.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point', 'wait');");
+
+# Give the restartpoint a checkpoint record to work on, then start it
+# in a background session; it will block on the injection point with
+# the checkpointer busy until released.
+$node_a->safe_psql('postgres', "CHECKPOINT;");
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+my $bg_psql = $node_b->background_psql('postgres', on_error_stop => 0);
+$bg_psql->query_until(
+ qr/starting_restartpoint/, q(
+ \echo starting_restartpoint
+ CHECKPOINT;
+));
+$node_b->wait_for_event('checkpointer', 'create-restart-point');
+
+# Clean switchover: the shutdown checkpoint of A streams to B and
+# becomes the last common checkpoint.
+$node_a->stop('fast');
+
+my ($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+$stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+my $shutdown_ckpt = $1;
+$node_b->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_ckpt'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+
+my $logstart = -s $node_b->logfile;
+$node_b->promote;
+
+# Accidental restart of the old primary, diverging its timeline.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->stop('fast');
+
+# Updates on the new primary before its first checkpoint; replay of
+# these on the rewound node has to read the pages from disk.
+$node_b->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 25 = 0;");
+ok( !$node_b->log_contains("checkpoint complete", $logstart),
+ "no checkpoint on the new timeline before the updates");
+
+# Release the checkpointer; the queued post-promotion checkpoint runs
+# after the updates.
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$node_b->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+$bg_psql->quit;
+
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+# pg_rewind refuses to run with full_page_writes disabled on the
+# source; the updates it was disabled for are already in WAL.
+$node_b->safe_psql('postgres', "ALTER SYSTEM SET full_page_writes = on;");
+$node_b->safe_psql('postgres', "SELECT pg_reload_conf();");
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the new primary');
+
+# Replay on the rewound node must start at the shutdown checkpoint of
+# the switchover, with the control file of the new primary.
+my $backup_label = slurp_file($node_a->data_dir . '/backup_label');
+$backup_label =~ /^CHECKPOINT LOCATION: ([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from backup_label";
+is($1, $shutdown_ckpt, 'replay starts at the switchover checkpoint');
+
+($stdout, $stderr) = run_command(
+ [
+ 'pg_waldump',
+ '-p' => $node_a->data_dir . '/pg_wal',
+ '-t' => 1,
+ '-s' => $shutdown_ckpt,
+ '-n' => 1,
+ ]);
+like($stdout, qr/CHECKPOINT_SHUTDOWN/,
+ 'last common checkpoint is a shutdown checkpoint');
+
+# pg_rewind keeps the target's own checksum state in the control file it
+# writes; the online enable reaches the rewound node through WAL replay
+# below, not through the copied control file.
+($stdout, $stderr) = run_command([ 'pg_controldata', $node_a->data_dir ]);
+like(
+ $stdout,
+ qr/^Data page checksum version:\s*0$/m,
+ 'rewound node keeps its own checksum state in the control file');
+
+# Start the rewound node as a standby of the new primary. Replay runs
+# through the pre-enable WAL stretch and the online enable.
+#
+# The rewind replaced the configuration files with those of the new
+# primary, so put the port back.
+my $connstr = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_b->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/014_lockstep.pl b/src/test/modules/test_checksums/t/014_lockstep.pl
new file mode 100644
index 00000000000..713f37fb22c
--- /dev/null
+++ b/src/test/modules/test_checksums/t/014_lockstep.pl
@@ -0,0 +1,182 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# The lockstep procedure for offline checksum changes in a replication
+# setup: stop all nodes, run pg_checksums on all of them, restart.
+# Replay of checkpoint records written before the change must not
+# revert the state of the standby.
+#
+# The same pair then runs the procedure in the disable direction, and
+# finally across a divergence: after a failover both nodes get their
+# checksums enabled offline, while the last common checkpoint still
+# carries "off". pg_rewind must keep the offline "on" on the rewound
+# node, since no record in the replayed WAL could ever restore it.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# wal_log_hints keeps the pair eligible for pg_rewind without checksums.
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_keep_size = '1GB'
+wal_log_hints = on
+]);
+$primary->start;
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Part 1: the lockstep procedure in the enable direction. Stop the
+# standby first: the WAL written after this point is replayed only
+# after the offline switch, and every checkpoint record in it still
+# carries the old state.
+$standby->stop;
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 0;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 1;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+# The lockstep procedure.
+$primary->checksum_enable_offline;
+$standby->checksum_enable_offline;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to off.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "off" in the replayed WAL/,
+ 0);
+test_checksum_state($standby, 'on');
+test_checksum_state($primary, 'on');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep enable');
+
+# Crash the standby and replay the same stretch again.
+$standby->stop('immediate');
+$standby->start;
+$primary->wait_for_catchup($standby);
+test_checksum_state($standby, 'on');
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after crash restart');
+
+# Once a post-switch checkpoint has been replayed the states match and
+# no warning may be logged.
+my $logstart = -s $standby->logfile;
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->wait_for_catchup($standby);
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile, $logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no mismatch warning once the states match');
+
+# Every page the standby wrote in this window must carry a checksum.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$standby->stop;
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'checksums valid on the standby');
+
+# Part 2: the lockstep procedure in the other direction. The standby
+# is already stopped; give it a pre-switch WAL stretch to replay, with
+# every checkpoint record in it still carrying "on" -- an explicit
+# checkpoint plus the shutdown checkpoint written by stop().
+$primary->safe_psql('postgres', "UPDATE t SET a = a WHERE a % 10 = 2;");
+$primary->safe_psql('postgres', "CHECKPOINT;");
+$primary->stop;
+
+$primary->checksum_disable_offline;
+$standby->checksum_disable_offline;
+
+my $disable_logstart = -s $standby->logfile;
+$primary->start;
+$standby->start;
+
+# The standby replays the pre-switch checkpoints; its state must not
+# revert to on.
+$primary->wait_for_catchup($standby);
+$standby->wait_for_log(qr/does not match the state "on" in the replayed WAL/,
+ $disable_logstart);
+test_checksum_state($standby, 'off');
+test_checksum_state($primary, 'off');
+
+is($standby->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'standby readable after lockstep disable');
+
+# Part 3: failover, divergence and pg_rewind across offline enables.
+# The roles swap from here on: the standby becomes the new primary and
+# the old primary is rewound to follow it, but the variables keep the
+# names they had above.
+#
+# Failover: promote the standby, then diverge the old primary.
+$standby->promote;
+$standby->safe_psql('postgres', "INSERT INTO t VALUES (0);");
+$primary->safe_psql('postgres', "INSERT INTO t VALUES (-1);");
+$primary->stop;
+
+# The lockstep procedure across the divergence: both nodes get their
+# checksums enabled offline.
+$primary->checksum_enable_offline;
+$standby->stop;
+$standby->checksum_enable_offline;
+$standby->start;
+test_checksum_state($standby, 'on');
+
+# The states match, so the rewind proceeds without complaint.
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $primary->data_dir,
+ '--source-server' => $standby->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled offline on both nodes');
+
+# The rewind replaced the configuration files with those of the source,
+# so put the port back. The copied file also carries the source's own
+# stale primary_conninfo; enable_streaming() below appends ours, which
+# wins as the later entry.
+$primary->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$primary->port]}
+]);
+$primary->enable_streaming($standby);
+
+my $rewind_logstart = -s $primary->logfile;
+$primary->start;
+$standby->wait_for_catchup($primary);
+
+is($primary->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10001', 'rewound server readable as a standby');
+
+# The offline enable survives: replay from the common checkpoint must
+# not resurrect the pre-divergence "off".
+test_checksum_state($primary, 'on');
+
+$log =
+ PostgreSQL::Test::Utils::slurp_file($primary->logfile, $rewind_logstart);
+unlike(
+ $log,
+ qr/does not match the state/,
+ 'no divergence reported between the rewound server and its source');
+
+$primary->stop;
+$standby->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
new file mode 100644
index 00000000000..4ab51c34d8f
--- /dev/null
+++ b/src/test/modules/test_checksums/t/015_standby_crash_after_disable.pl
@@ -0,0 +1,136 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that a standby crashing after a replayed online disable, but before
+# its next restartpoint, restarts cleanly.
+#
+# Pages dirtied before the XLOG2_CHECKSUMS record and evicted after it are
+# written out under the new "off" state, without checksums. The control
+# file must follow the record immediately in this direction: a crash-restart
+# initializes checksum verification from the control file, and one still
+# saying "on" would fail verification on exactly those pages while replaying
+# records older than the state change.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$primary->start;
+
+test_checksum_state($primary, 'on');
+
+$primary->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,1000) AS a;");
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+
+# Small buffer pool so that replaying a bulk load evicts the pages we care
+# about, and no restartpoints so the control file stays where it is.
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 1MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_delay = 10000
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'on');
+
+# No full page images, so that replay has to read the pages of "t" back from
+# disk instead of overwriting them from the WAL.
+$primary->append_conf('postgresql.conf', 'full_page_writes = off');
+$primary->reload;
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" on the standby, *before* the state change. They are
+# not flushed: the standby has no restartpoint from here on.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+# Online disable. The standby replays it and switches to "off".
+$primary->safe_psql('postgres', 'SELECT pg_disable_data_checksums();');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'off');
+
+my ($ctl_before) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($ctl_state) = $ctl_before =~ /Data page checksum version:\s+(\d+)/;
+note(
+ "standby control file data_checksum_version after the disable: "
+ . "$ctl_state (0 = off, 1 = on)");
+
+# Force the standby to evict the dirty pages of "t" now that it is "off":
+# they are written back without checksums.
+$primary->safe_psql('postgres',
+ "CREATE TABLE filler AS SELECT generate_series(1,300000) AS a;");
+$primary->wait_for_catchup($standby);
+
+# Crash the standby before it gets a chance to run a restartpoint.
+$standby->stop('immediate');
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'standby restarts after crashing between a checksum state '
+ . 'change and the next restartpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'the pages written while "off" are readable')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("standby log tail:\n" . join("\n", @lines[ -12 .. -1 ]))
+ if @lines >= 12;
+ fail('the pages written while "off" are readable');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/016_promote_enable_crash.pl b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
new file mode 100644
index 00000000000..13c983136ef
--- /dev/null
+++ b/src/test/modules/test_checksums/t/016_promote_enable_crash.pl
@@ -0,0 +1,146 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a standby promoted right after replaying an online enable, before any
+# restartpoint has flushed the rewritten pages.
+#
+# The end-of-recovery record persists the checksum state without flushing the
+# buffer pool, while the control file still points at a restartpoint older than
+# the transition. Persisting "on" there would make a crash before the
+# post-promotion checkpoint verify checksums over the pages that were flushed
+# while the state was still "off"; promotion must take the full checkpoint
+# instead.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+
+# Establish the restartpoint that recovery will resume from after the crash.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Dirty the pages of "t" without full page images, then push them out to disk
+# on the standby while checksums are still off.
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Online enable on the primary; the standby replays it into its own buffers,
+# where the rewritten pages stay dirty (no restartpoint, no bgwriter).
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+note("standby control file before the promotion: $before");
+
+# Promote. The end-of-recovery record persists the live state without any
+# buffer flush; the checkpoint requested afterwards is not immediate.
+$standby->promote;
+$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery();')
+ or die 'timed out waiting for the promotion';
+$standby->stop('immediate');
+
+($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($ckpt) = $ctl =~ /Latest checkpoint location:\s+(\S+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note(
+ "standby control file after the promotion: $after, checkpoint $ckpt, redo $redo"
+);
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started,
+ 'promoted node restarts after crashing before its first checkpoint');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/017_restartpoint_race.pl b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
new file mode 100644
index 00000000000..10919c9741c
--- /dev/null
+++ b/src/test/modules/test_checksums/t/017_restartpoint_race.pl
@@ -0,0 +1,151 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test a restartpoint whose flush races the replay of an online enable.
+#
+# Replay keeps running while CheckPointGuts() writes out the buffer pool, so
+# the state at the end of the flush can be newer than the one the flush ran
+# under. Persisting it would claim checksums for pages the flush wrote out
+# before the transition, against a redo pointer that predates it.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+full_page_writes = off
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+shared_buffers = 512MB
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+bgwriter_lru_maxpages = 0
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# C1: the checkpoint the pending restartpoint will target.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+
+# Dirty the pages of "t" without full page images and push them to disk on
+# the standby while it is still "off".
+$primary->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+$primary->wait_for_catchup($standby);
+
+my $relpath = $standby->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+my $evicted = $standby->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+note("evict_relation on the standby: $evicted, relpath $relpath");
+
+# Start a restartpoint for C1 and hold it right after CheckPointGuts().
+$standby->safe_psql('postgres',
+ "SELECT injection_points_attach('create-restart-point','wait');");
+my $bg = $standby->background_psql('postgres');
+$bg->query_until(qr//, "\\echo restartpoint\nCHECKPOINT;\n");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint injection point';
+
+# The startup process keeps replaying while the checkpointer is held: the
+# whole online enable lands, and the rewritten pages stay dirty.
+enable_data_checksums($primary, wait => 'on');
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+# Let the restartpoint finish; it persists the state it samples now.
+$standby->safe_psql('postgres',
+ "SELECT injection_points_wakeup('create-restart-point');");
+$standby->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_stat_activity WHERE wait_event = 'create-restart-point';"
+) or die 'timed out waiting for the restartpoint to finish';
+$standby->safe_psql('postgres',
+ "SELECT injection_points_detach('create-restart-point');");
+
+$standby->stop('immediate');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+my ($redo) = $ctl =~ /Latest checkpoint's REDO location:\s+(\S+)/;
+note("standby control file after the restartpoint: $after, redo $redo");
+
+my $page;
+open(my $fh, '<', $standby->data_dir . '/' . $relpath) or die $!;
+binmode $fh;
+read($fh, $page, 8192);
+close($fh);
+my ($pd_checksum) = unpack('x8 v', $page);
+note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+my $started = $standby->start(fail_ok => 1);
+ok($started, 'standby restarts after a restartpoint that raced the enable');
+
+my $log = PostgreSQL::Test::Utils::slurp_file($standby->logfile);
+unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+unlike($log, qr/invalid page in block/, 'no invalid pages while replaying');
+
+if ($started)
+{
+ my ($rc, $stdout, $stderr) =
+ $standby->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ $standby->stop('immediate');
+}
+else
+{
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+}
+
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/018_enable_crash_windows.pl b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
new file mode 100644
index 00000000000..4dac590e819
--- /dev/null
+++ b/src/test/modules/test_checksums/t/018_enable_crash_windows.pl
@@ -0,0 +1,508 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Crash and checkpoint windows inside an online data checksum enable on
+# a single node. Each scenario starts from "off" on the same cluster,
+# runs the enable into a held injection point, and checks what a crash
+# or a concurrent checkpoint in that window may and may not do to the
+# control file and the pages on disk.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+use IPC::Run;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum_extended' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum_extended\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+# Bring the node back to "off" with no leftover table between
+# scenarios. The crash restarts already resolve an interrupted enable
+# to "off", so only disable when something is left to disable. Those
+# restarts also clear any attached injection points. The checkpoint
+# bounds the WAL the next scenario's crash recovery has to replay.
+sub reset_scenario
+{
+ my ($node) = @_;
+
+ BAIL_OUT('node is not running after the previous scenario')
+ unless $node->is_alive;
+
+ my $state = $node->safe_psql('postgres', 'SHOW data_checksums;');
+ disable_data_checksums($node, wait => 'off') if $state ne 'off';
+ $node->safe_psql('postgres', 'DROP TABLE IF EXISTS t;');
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+ test_checksum_state($node, 'off');
+}
+
+my $node = PostgreSQL::Test::Cluster->new('crash_windows');
+$node->init(no_data_checksums => 1);
+
+# Scenario 1 needs the pages of "t" to reach disk without a full page
+# image, hence full_page_writes and wal_log_hints off. Scenario 2
+# needs them to stay dirty in shared buffers until a checkpoint, hence
+# no bgwriter and a large shared_buffers. wal_level is pinned because
+# init() writes "minimal" for a non-streaming node. The rest merely
+# tolerate the settings.
+$node->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+wal_level = replica
+max_wal_size = 10GB
+full_page_writes = off
+wal_log_hints = off
+shared_buffers = 512MB
+bgwriter_lru_maxpages = 0
+));
+$node->start;
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$node->safe_psql('postgres', 'CREATE EXTENSION pg_buffercache;');
+
+test_checksum_state($node, 'off');
+
+# Scenario 1: a primary crashes inside SetDataChecksumsOn(), just before the
+# forced checkpoint that flushes the rewritten pages, with crash recovery
+# resuming from a checkpoint older than the transition. The control file
+# must still say "inprogress-on" there: replay does not adopt the state of
+# the checkpoint record it resumes from, so an "on" written before the flush
+# would stay in effect while replay reads pages whose rewrite never reached
+# disk. Here the pages went out through the rewriting worker's ring buffer,
+# so only the control file state discriminates; see the next scenario for
+# the case where they did not.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Establish the checkpoint that crash recovery will resume from.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty the pages of "t" without full page images, then push them out
+ # to disk while checksums are still off.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+ my $evicted = $node->safe_psql('postgres',
+ "SELECT pg_buffercache_evict_relation('t'::regclass);");
+ note("evict_relation on t: $evicted, relpath $relpath");
+
+ # Stop the enable right after the control file has been updated to
+ # "on" but before the checkpoint that flushes the rewritten pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+ $node->safe_psql('postgres', 'SELECT pg_enable_data_checksums();');
+
+ $node->poll_query_until('postgres',
+ "SELECT count(*) > 0 FROM pg_stat_activity WHERE wait_event = 'datachecksums-on-before-checkpoint';"
+ ) or die 'timed out waiting for the injection point';
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ $node->stop('immediate');
+
+ # Show the on-disk checksum field of the first page of "t".
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 2: the same crash as above, but with the rewritten pages resident
+# in shared buffers instead of gone out through the ring. The rewriting
+# worker reads through a BAS_VACUUM ring, which writes pages back as the
+# ring recycles, but a page already resident in shared buffers is not read
+# through the ring: ReadBufferExtended() hands back the existing buffer, and
+# it stays dirty until a checkpoint. The control file may therefore not
+# say "on" before that checkpoint has run, or crash recovery would resume
+# from a checkpoint older than the transition with verification already
+# enabled, and replay records that read those still-unchecksummed pages.
+# full_page_writes is off so that the records do not simply overwrite the
+# pages with a full page image.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,100000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Establish the checkpoint that crash recovery will resume from, with the
+ # pages of "t" written out while checksums are still off.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ # Dirty those pages again without emitting full page images, and leave
+ # them in shared buffers. Replay of these records has to read the
+ # pages from disk.
+ $node->safe_psql('postgres', 'UPDATE t SET a = a + 1;');
+
+ my $dirty_before = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_before, '>', 0,
+ 'pages of t are resident and dirty before enabling checksums');
+
+ # Hold the enabling right before the checkpoint that flushes the rewritten
+ # pages.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($ctl_state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ note("control file data_checksum_version before the crash: $ctl_state");
+ is($ctl_state, '3',
+ 'control file still says "inprogress-on" before the checkpoint');
+
+ # The rewritten pages must still be sitting dirty in shared buffers, or
+ # the window this test is about does not exist.
+ my $dirty_after = $node->safe_psql('postgres',
+ "SELECT count(*) FROM pg_buffercache "
+ . "WHERE relfilenode = pg_relation_filenode('t'::regclass) "
+ . "AND relforknumber = 0 AND isdirty;");
+ cmp_ok($dirty_after, '>', 0,
+ 'rewritten pages of t are still dirty before the checkpoint');
+
+ $node->stop('immediate');
+
+ # The on-disk copy is the one written before the transition, without a
+ # checksum.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ is($pd_checksum, 0, 'block 0 of t on disk carries no checksum');
+
+ my $started = $node->start(fail_ok => 1);
+ ok($started, 'primary restarts after crashing inside the online enable');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/page verification failed/,
+ 'no checksum verification failures while replaying');
+ unlike($log, qr/invalid page in block/,
+ 'no invalid pages while replaying');
+
+ if ($started)
+ {
+ my ($rc, $stdout, $stderr) =
+ $node->psql('postgres', 'SELECT count(*) FROM t;');
+ is($rc, 0, 'table readable after the crash restart')
+ or diag("stderr: $stderr");
+ }
+ else
+ {
+ my @lines = grep { /FATAL|PANIC|invalid page|verification failed/ }
+ split(/\n/, $log);
+ diag("log tail:\n" . join("\n", @lines));
+ fail('table readable after the crash restart');
+ }
+}
+reset_scenario($node);
+
+# Scenario 3: a checkpoint that runs between the XLOG2_CHECKSUMS("on")
+# record and the control file write at the end of SetDataChecksumsOn() must
+# not make a crash throw the completed transition away. SetDataChecksumsOn()
+# writes the record, flips shared memory to "on", emits the barrier and only
+# then requests the checkpoint that flushes the rewritten pages; the control
+# file is written after that checkpoint returns. A crash in that window is
+# harmless only as long as recovery still starts before the record. Any
+# checkpoint completing in the window moves the redo point past the record,
+# so recovery would never see it, would come up with the control file's
+# "inprogress-on" and StartupXLOG() would demote that to "off", even though
+# every page on disk carries a checksum by then. CreateCheckPoint()
+# therefore persists the state the checkpoint ran under, the same way
+# CreateRestartPoint() does. Here the window is made deterministic by
+# holding the launcher at the datachecksums-on-before-checkpoint injection
+# point and checkpointing from another session.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+ my $relpath = $node->safe_psql('postgres',
+ "SELECT pg_relation_filepath('t'::regclass);");
+
+ # Hold the launcher after the record, the shared memory flip and the
+ # barrier, but before the checkpoint SetDataChecksumsOn() requests
+ # itself.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+
+ # Every backend already sees "on" and writes checksums.
+ test_checksum_state($node, 'on');
+
+ # ... while the control file still says "inprogress-on".
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($before) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($before, '3', 'control file says "inprogress-on" inside the window');
+
+ # A concurrent checkpoint. It flushes the rewritten pages and moves
+ # the redo point past the XLOG2_CHECKSUMS("on") record, so it has to
+ # record the "on" state in the control file as well.
+ $node->safe_psql('postgres', 'CHECKPOINT;');
+
+ ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($after) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($after, '1',
+ 'the concurrent checkpoint records "on" in the control file');
+
+ $node->stop('immediate');
+
+ # The checkpoint flushed the rewritten pages, so they carry a checksum on
+ # disk: the transition really did complete.
+ my $page;
+ open(my $fh, '<', $node->data_dir . '/' . $relpath) or die $!;
+ binmode $fh;
+ read($fh, $page, 8192);
+ close($fh);
+ my ($pd_checksum) = unpack('x8 v', $page);
+ note("on-disk pd_checksum of t block 0: $pd_checksum");
+ isnt($pd_checksum, 0, 'block 0 of t on disk carries a checksum');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition is complete on disk, so the cluster has to come back
+ # "on".
+ test_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+}
+reset_scenario($node);
+
+# Scenario 4: a crash between the checkpoint an online enable requests and
+# the control file write that follows it must not lose the transition. The
+# last steps of SetDataChecksumsOn() are
+#
+# WAL record -> shmem -> barrier -> checkpoint -> persist
+#
+# The checkpoint is what makes "on" safe to persist, but it also moves the
+# redo point above the XLOG2_CHECKSUMS record that carries the new state.
+# Crash recovery started from that checkpoint therefore never replays the
+# record, so without the state the checkpoint itself persists, a crash in
+# the remaining window would bring the cluster back at "inprogress-on",
+# which StartupXLOG() resolves to "off", discarding a transition whose
+# pages are all on disk with a checksum. The previous scenario exercises
+# the same window through a checkpoint requested by another session; this
+# one closes it from the other side: the transition's own checkpoint is the
+# one that moves the redo point, so persisting the state from
+# CreateCheckPoint() is what has to save it, not the write below the
+# injection point.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # Hold the launcher after the checkpoint that licenses "on" has
+ # completed and before the state reaches the control file.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-after-checkpoint','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-after-checkpoint');
+
+ # The transition is complete as far as the running cluster is concerned.
+ test_checksum_state($node, 'on');
+
+ # The checkpoint has already recorded it, so the pending write below the
+ # injection point has nothing left to do.
+ my ($ctl) = run_command([ 'pg_controldata', $node->data_dir ]);
+ my ($version) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+ is($version, '1',
+ 'the requested checkpoint recorded "on" in the control file');
+
+ # Crash before the control file write that follows the checkpoint.
+ $node->stop('immediate');
+
+ $node->start;
+
+ # Every page on disk carries a checksum and the checkpoint that
+ # flushed them completed, so the cluster has to come back verifying
+ # them.
+ test_checksum_state($node, 'on');
+
+ is($node->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '10000', 'relation readable after the crash');
+
+ my $log = PostgreSQL::Test::Utils::slurp_file($node->logfile);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ # The data directory must be one the offline tools accept.
+ $node->stop;
+ $node->command_ok([ 'pg_checksums', '--check', '-D', $node->data_dir ],
+ 'pg_checksums accepts the data directory');
+
+ # Leave the node running for the reset.
+ $node->start;
+}
+reset_scenario($node);
+
+# Scenario 5: a checkpoint racing SetDataChecksumsOn() between the insertion
+# of the XLOG2_CHECKSUMS("on") record and the shared memory update must not
+# insert an XLOG_CHECKPOINT_REDO record that follows the transition in WAL
+# order while still carrying "inprogress-on". Recovery resuming from such a
+# redo point never replays the preceding transition record, comes up in
+# "inprogress-on" and resolves the finished transition as interrupted.
+# XLogChecksums() closes the window by inserting the record and publishing
+# the new state under DataChecksumTransitionLock, which CreateCheckPoint()
+# takes around sampling the state and inserting the redo record. Here the
+# launcher is held between the two steps at the
+# datachecksums-on-before-publish injection point, and a concurrent
+# CHECKPOINT has to block on the lock instead of completing inside the
+# window.
+{
+ $node->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+ # The datachecksums-on-before-publish point fires inside a critical
+ # section, where the wait machinery must not allocate. Waiting once
+ # at the datachecksums-enable-checksums-delay point, which the
+ # launcher runs outside the critical section, initializes it; see
+ # 050_redo_segment_missing.pl for the same recipe around
+ # create-checkpoint-run.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-enable-checksums-delay','wait');"
+ );
+
+ # Hold the launcher after the XLOG2_CHECKSUMS("on") record is in WAL but
+ # before the new state is published in shared memory.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-publish','wait');"
+ );
+
+ enable_data_checksums($node);
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-enable-checksums-delay');
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-enable-checksums-delay');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-enable-checksums-delay');");
+
+ $node->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-publish');
+
+ # The record is in WAL, the published state is still the old one.
+ test_checksum_state($node, 'inprogress-on');
+
+ # A concurrent checkpoint. It must block on DataChecksumTransitionLock
+ # before inserting its redo record rather than complete inside the window.
+ my $checkpointer = IPC::Run::start(
+ [ 'psql', '-XAtq', '-d', $node->connstr('postgres'), '-c', 'CHECKPOINT;' ],
+ '<' => '/dev/null',
+ '>' => '/dev/null',
+ '2>' => '/dev/null',
+ IPC::Run::timer($PostgreSQL::Test::Utils::timeout_default));
+
+ ok( $node->poll_query_until(
+ 'postgres',
+ "SELECT wait_event = 'DataChecksumTransition' "
+ . "FROM pg_stat_activity WHERE backend_type = 'checkpointer';"),
+ 'concurrent checkpoint blocks on DataChecksumTransitionLock');
+
+ # Release the launcher; the checkpoint then samples the published "on" and
+ # its redo record follows the transition record.
+ $node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-publish');");
+ $node->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-publish');");
+
+ $checkpointer->finish;
+
+ # Crash while the launcher's own checkpoint may still be in flight.
+ # Recovery resumes from the concurrent checkpoint's redo point, which
+ # now lies above the transition record and carries "on".
+ $node->stop('immediate');
+
+ my $log_offset = -s $node->logfile;
+ $node->start;
+
+ # The transition completed, so the cluster has to come back "on".
+ wait_for_checksum_state($node, 'on');
+
+ my $log =
+ PostgreSQL::Test::Utils::slurp_file($node->logfile, $log_offset);
+ unlike(
+ $log,
+ qr/enabling data checksums was interrupted/,
+ 'the completed transition is not reported as interrupted');
+
+ $node->stop;
+}
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
new file mode 100644
index 00000000000..abcc234ab22
--- /dev/null
+++ b/src/test/modules/test_checksums/t/019_standby_shutdown_catchup.pl
@@ -0,0 +1,138 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A standby stopped after replaying an online enable, but before the
+# checkpoint record that follows it on the primary.
+#
+# The shutdown restartpoint has no new checkpoint record to work from
+# and is skipped, so the control file would keep the in-progress state
+# the transition passed through, even though replay left the node at
+# "on" and rewrote every page. pg_checksums reads that field and
+# refuses to run on an in-progress state, so the shutdown has to flush
+# and catch it up instead.
+#
+# The caught-up control file then carries the watermark of the "on"
+# record, which the second half of this test relies on: an offline
+# disable made while the standby is down must survive the re-replay of
+# that record on the next startup, or the change would be silently
+# reverted.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+wal_keep_size = 1GB
+));
+$primary->start;
+
+$primary->safe_psql('postgres', 'CREATE EXTENSION injection_points;');
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,10000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', qq(
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+));
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+# Anchor the standby's last restartpoint here, so that nothing replayed
+# from now on gives the shutdown restartpoint a newer checkpoint record
+# to use.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+
+# Hold the enable right after the state change record, before the
+# checkpoint it requests once the transition is complete.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_attach('datachecksums-on-before-checkpoint','wait');"
+);
+enable_data_checksums($primary);
+$primary->wait_for_event('datachecksums launcher',
+ 'datachecksums-on-before-checkpoint');
+wait_for_checksum_state($primary, 'on');
+
+# The state change record is already flushed and streams on its own;
+# no checkpoint record follows while the launcher is held.
+$primary->wait_for_catchup($standby);
+wait_for_checksum_state($standby, 'on');
+
+$standby->stop('fast');
+
+my ($ctl) = run_command([ 'pg_controldata', $standby->data_dir ]);
+my ($state) = $ctl =~ /Data page checksum version:\s+(\d+)/;
+is($state, '1',
+ 'shutdown catches the control file up with the replayed state');
+
+command_ok([ 'pg_checksums', '--check', '-D', $standby->data_dir ],
+ 'pg_checksums verifies the stopped standby');
+
+# Disable checksums offline while the standby is down. The restart
+# below resumes replay under the last restartpoint, re-reading the "on"
+# record; the watermark persisted with the catch-up above marks it as
+# already applied, so the offline change must survive.
+$standby->checksum_disable_offline;
+
+# Let the enable finish on the primary.
+$primary->safe_psql('postgres',
+ "SELECT injection_points_wakeup('datachecksums-on-before-checkpoint');");
+$primary->safe_psql('postgres',
+ "SELECT injection_points_detach('datachecksums-on-before-checkpoint');");
+
+# The launcher exits only after the checkpoint it requests completes,
+# so a checkpoint record now follows the transition in WAL.
+$primary->poll_query_until('postgres',
+ "SELECT count(*) = 0 FROM pg_catalog.pg_stat_activity "
+ . "WHERE backend_type = 'datachecksums launcher';");
+
+my $log_offset = -s $standby->logfile;
+$standby->start;
+$primary->wait_for_catchup($standby);
+
+test_checksum_state($standby, 'off');
+
+# The nodes legitimately diverged, which replayed checkpoints report.
+$standby->wait_for_log(
+ qr/data checksum state "off" of this node does not match the state "on" in the replayed WAL/,
+ $log_offset);
+
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/020_cascade_divergence.pl b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
new file mode 100644
index 00000000000..8f2816ac405
--- /dev/null
+++ b/src/test/modules/test_checksums/t/020_cascade_divergence.pl
@@ -0,0 +1,127 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# An offline data checksum state change made on one node of a cascading setup
+# is reported all the way down the chain.
+#
+# CheckReplayedDataChecksumState() is reached from xlog_redo() when a
+# checkpoint record (XLOG_CHECKPOINT_SHUTDOWN or XLOG_CHECKPOINT_REDO) is
+# replayed after consistency has been reached. Those records are written by
+# the root primary only and are relayed verbatim by every intermediate
+# standby, so a cascaded standby that was changed offline learns about the
+# divergence too. An intermediate standby's own restartpoints are not
+# WAL-logged and cannot surface it.
+#
+# Note that this makes detection checkpoint-driven: an idle or read-only
+# primary surfaces the divergence no sooner than its next checkpoint, so up to
+# checkpoint_timeout may pass before the operator sees anything. Closing that
+# window would require comparing the state when a walreceiver connects.
+
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum(_extended)?\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+# checkpoint_timeout is deliberately long so that the only checkpoint record in
+# this test is the one requested explicitly below.
+my $conf = qq(
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+);
+
+my $primary = PostgreSQL::Test::Cluster->new('cascade_primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf('postgresql.conf', $conf);
+$primary->start;
+$primary->safe_psql('postgres',
+ 'CREATE TABLE t AS SELECT generate_series(1,1000) AS a;');
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('cascade_standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf('postgresql.conf', $conf);
+$standby->start;
+
+$standby->backup('backup');
+my $cascade = PostgreSQL::Test::Cluster->new('cascade_cascade');
+$cascade->init_from_backup($standby, 'backup', has_streaming => 1);
+$cascade->append_conf('postgresql.conf', $conf);
+$cascade->start;
+
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+test_checksum_state($primary, 'off');
+test_checksum_state($standby, 'off');
+test_checksum_state($cascade, 'off');
+
+# Change the cascaded standby offline, and only it.
+$cascade->stop;
+$cascade->command_ok([ 'pg_checksums', '--enable', '-D', $cascade->data_dir ],
+ 'pg_checksums enables checksums on the cascaded standby only');
+
+my $logstart = -s $cascade->logfile;
+$cascade->start;
+
+test_checksum_state($cascade, 'on');
+
+# Ordinary WAL traffic carries no state, so the cascaded standby keeps
+# streaming from a chain whose state is "off" without noticing.
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (1);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+is($cascade->safe_psql('postgres', 'SELECT count(*) FROM t;'),
+ '1001', 'cascaded standby keeps streaming from a divergent chain');
+
+# Restartpoints on the intermediate standby are not WAL-logged, so this does
+# not surface anything on the cascaded standby either.
+$standby->safe_psql('postgres', 'CHECKPOINT;');
+$primary->safe_psql('postgres', 'INSERT INTO t VALUES (2);');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+my $log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+unlike(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'an intermediate restartpoint does not surface the divergence');
+
+# A checkpoint on the root primary does: the record is relayed down the whole
+# chain, so both the direct and the cascaded standby compare it against their
+# own state. wait_for_catchup() below waits for replay, not just receipt, so
+# the record has been through xlog_redo() on both nodes once it returns.
+$primary->safe_psql('postgres', 'CHECKPOINT;');
+$primary->wait_for_catchup($standby);
+$standby->wait_for_catchup($cascade);
+
+$log = PostgreSQL::Test::Utils::slurp_file($cascade->logfile, $logstart);
+like(
+ $log,
+ qr/data checksum state .* does not match/,
+ 'the cascaded standby reports the divergence on the primary checkpoint');
+
+$cascade->stop;
+$standby->stop;
+$primary->stop;
+
+done_testing();
diff --git a/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
new file mode 100644
index 00000000000..3f660623bd3
--- /dev/null
+++ b/src/test/modules/test_checksums/t/021_rewind_divergent_transitions.pl
@@ -0,0 +1,205 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# pg_rewind across online data checksum transitions on both sides of a
+# divergence, with the two nodes trading roles between the scenarios.
+#
+# Scenario 1: after a switchover the new primary enables checksums
+# online, while the old primary restarts on its old timeline, advances
+# its WAL beyond the enable records and runs an online enable/disable
+# cycle of its own. The old primary ends "off" with a checksum
+# watermark numerically above every checksum record the new primary has
+# written. pg_rewind clamps the watermark it keeps to the divergence
+# point; without the clamp, replay on the rewound node would skip the
+# source's enable as already applied and stay "off" under an "on"
+# primary.
+#
+# Scenario 2: checksums are disabled again, and after another
+# switchover both nodes enable them online independently, so the
+# divergence checkpoint carries "off" while both control files say
+# "on". The rewind is allowed, the target keeps its own "on" state,
+# and replay re-walks the source's enable onto the already enabled
+# node.
+use strict;
+use warnings FATAL => 'all';
+
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+use FindBin;
+use lib $FindBin::RealBin;
+
+use DataChecksums::Utils;
+
+# This test suite is expensive to execute, require PG_TEST_EXTRA to contain
+# 'checksum_extended' to run it.
+if ($ENV{PG_TEST_EXTRA})
+{
+ plan skip_all => 'Expensive data checksums test disabled'
+ unless ($ENV{PG_TEST_EXTRA} =~ /\bchecksum_extended\b/);
+}
+else
+{
+ plan skip_all => 'Expensive data checksums test disabled';
+}
+
+sub controldata_watermark
+{
+ my ($node) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $node->data_dir ]);
+ $stdout =~ /^Data checksum watermark:\s*([0-9A-F]+)\/([0-9A-F]+)$/m
+ or die "watermark missing from pg_controldata output";
+ return (hex($1) << 32) + hex($2);
+}
+
+# Wait until the standby has replayed the shutdown checkpoint of the
+# stopped primary, so that a subsequent promotion diverges after it and
+# the shutdown checkpoint becomes the last common checkpoint.
+sub wait_for_shutdown_checkpoint_replay
+{
+ my ($primary, $standby) = @_;
+ my ($stdout) = run_command([ 'pg_controldata', $primary->data_dir ]);
+ $stdout =~ /^Latest checkpoint location:\s*([0-9A-F\/]+)$/m
+ or die "checkpoint location missing from pg_controldata output";
+ my $shutdown_checkpoint = $1;
+
+ $standby->poll_query_until('postgres',
+ "SELECT pg_last_wal_replay_lsn() > '$shutdown_checkpoint'::pg_lsn;")
+ or die "standby never replayed the shutdown checkpoint";
+}
+
+# Old primary, checksums off. wal_log_hints is required by pg_rewind
+# on a cluster without data checksums.
+my $node_a = PostgreSQL::Test::Cluster->new('node_a');
+$node_a->init(allows_streaming => 1, no_data_checksums => 1);
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+autovacuum = off
+wal_log_hints = on
+wal_keep_size = '1GB'
+]);
+$node_a->start;
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t AS SELECT generate_series(1,10000) AS a;");
+$node_a->safe_psql('postgres', "CREATE TABLE t_div (a int);");
+
+$node_a->backup('backup');
+my $node_b = PostgreSQL::Test::Cluster->new('node_b');
+$node_b->init_from_backup($node_a, 'backup', has_streaming => 1);
+$node_b->start;
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+
+# Clean switchover to B; enable checksums online on it.
+$node_a->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_a, $node_b);
+$node_b->promote;
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+
+$node_b->stop('fast');
+my $watermark_b = controldata_watermark($node_b);
+$node_b->start;
+
+# Accidental restart of the old primary on the old timeline. Advance
+# its WAL beyond the enable watermark of B, then run an online enable
+# and disable cycle: the node ends "off" with a watermark above every
+# checksum record B has written.
+$node_a->start;
+$node_a->safe_psql('postgres', "INSERT INTO t_div VALUES (1);");
+$node_a->safe_psql('postgres',
+ "CREATE TABLE t_pad AS SELECT generate_series(1,200000) AS a;");
+enable_data_checksums($node_a, wait => 'on');
+disable_data_checksums($node_a, wait => 'off');
+test_checksum_state($node_a, 'off');
+$node_a->stop('fast');
+
+my $watermark_a = controldata_watermark($node_a);
+die "test broken: target watermark not above the source's enable"
+ unless $watermark_a > $watermark_b;
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_a->data_dir,
+ '--source-server' => $node_b->connstr('postgres'),
+ ],
+ 'pg_rewind from the promoted node');
+
+# Start the rewound node as a standby of B. Replay from the last
+# common checkpoint runs through B's online enable, which the target
+# never saw, so the rewound node must converge to "on".
+my $connstr_b = $node_b->connstr;
+$node_a->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_a->port]}
+primary_conninfo = '$connstr_b application_name=@{[$node_a->name]}'
+]);
+$node_a->set_standby_mode;
+$node_a->start;
+
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'on');
+
+is($node_a->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+# Scenario 2, reusing the pair with the roles reversed. Disable
+# checksums online so the next divergence point carries "off", and let
+# A replay the change.
+disable_data_checksums($node_b, wait => 'off');
+$node_b->wait_for_catchup($node_a, 'replay', $node_b->lsn('insert'));
+test_checksum_state($node_a, 'off');
+
+# Clean switchover back to A; enable checksums online on it.
+$node_b->stop('fast');
+wait_for_shutdown_checkpoint_replay($node_b, $node_a);
+$node_a->promote;
+enable_data_checksums($node_a, wait => 'on');
+test_checksum_state($node_a, 'on');
+my $source_enable_watermark = controldata_watermark($node_a);
+
+# The old primary restarts on its old timeline and enables checksums
+# online independently: both control files say "on", the divergence
+# checkpoint says "off".
+$node_b->start;
+$node_b->safe_psql('postgres', "INSERT INTO t_div VALUES (2);");
+enable_data_checksums($node_b, wait => 'on');
+test_checksum_state($node_b, 'on');
+$node_b->stop('fast');
+
+command_ok(
+ [
+ 'pg_rewind',
+ '--target-pgdata' => $node_b->data_dir,
+ '--source-server' => $node_a->connstr('postgres'),
+ ],
+ 'pg_rewind with checksums enabled on both nodes');
+
+my $connstr_a = $node_a->connstr;
+$node_b->append_conf(
+ 'postgresql.conf', qq[
+port = @{[$node_b->port]}
+primary_conninfo = '$connstr_a application_name=@{[$node_b->name]}'
+]);
+$node_b->set_standby_mode;
+$node_b->start;
+
+$node_a->wait_for_catchup($node_b, 'replay', $node_a->lsn('insert'));
+test_checksum_state($node_b, 'on');
+
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t;"),
+ '10000', 'data readable on the rewound node');
+is($node_b->safe_psql('postgres', "SELECT count(*) FROM t_div;"),
+ '0', 'divergent insert was rewound');
+
+$node_b->stop('fast');
+is(controldata_watermark($node_b), $source_enable_watermark,
+ 'rewound node replayed the source checksum transition');
+command_ok([ 'pg_checksums', '--check', '-D', $node_b->data_dir ],
+ 'checksums valid on the rewound node');
+
+$node_a->stop('fast');
+command_ok([ 'pg_checksums', '--check', '-D', $node_a->data_dir ],
+ 'checksums valid on the twice-rewound node');
+
+done_testing();
--
2.39.3 (Apple Git-146)
=
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-14 13:48 Daniel Gustafsson <daniel@yesql.se>
parent: Daniel Gustafsson <daniel@yesql.se>
0 siblings, 1 reply; 43+ messages in thread
From: Daniel Gustafsson @ 2026-09-14 13:48 UTC (permalink / raw)
To: Heikki Linnakangas <hlinnaka@iki.fi>; +Cc: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-hackers@lists.postgresql.org
Pushed to master along with the cleanups from the "Trying to break online
checksums with LLMs" thread. Not backpatched yet, as I wanted a few builds in
the BF first and also some input on the new open item in the other thread.
--
Daniel Gustafsson
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-26 05:00 Alexander Lakhin <exclusion@gmail.com>
parent: Daniel Gustafsson <daniel@yesql.se>
0 siblings, 1 reply; 43+ messages in thread
From: Alexander Lakhin @ 2026-09-26 05:00 UTC (permalink / raw)
To: Daniel Gustafsson <daniel@yesql.se>; +Cc: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; Zsolt Parragi <zsolt.parragi@percona.com>; Heikki Linnakangas <hlinnaka@iki.fi>; pgsql-hackers@lists.postgresql.org
Hello Daniel,
14.09.2026 16:48, Daniel Gustafsson wrote:
> Pushed to master along with the cleanups from the "Trying to break online
> checksums with LLMs" thread. Not backpatched yet, as I wanted a few builds in
> the BF first and also some input on the new open item in the other thread.
BF animal turaco (Raspberry PI, kernel 6.12.75+rpt-rpi-v8) managed to fail
a test added in b52a1c2c8:
[22:51:29.073](0.002s) ok 7 - replay starts at the switchover checkpoint
[22:51:29.087](0.014s) not ok 8 - last common checkpoint is a shutdown checkpoint
[22:51:29.088](0.001s)
[22:51:29.088](0.000s) # Failed test 'last common checkpoint is a shutdown checkpoint'
# at t/013_rewind.pl line 161.
[22:51:29.089](0.001s) # ''
# doesn't match '(?^:CHECKPOINT_SHUTDOWN)'
[22:51:29.100](0.011s) ok 9 - rewound node keeps its own checksum state in the control file
...
Test Summary Report
-------------------
t/013_rewind.pl (Wstat: 256 (exited 1) Tests: 13 Failed: 1)
I've reproduced this failure with the following modification:
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -3446,2 +3446,3 @@ XLogFileInitInternal(XLogSegNo logsegno, TimeLineID logtli,
*/
+pg_usleep(100000);
installed_segno = logsegno;
which makes all-zero 000000020000000000000005 appear inside
$node_a->data_dir . '/pg_wal':
tr -d '\000' < src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal/000000020000000000000005
| wc
0 0 0
and then if readdir() happens to return this file first (I'm observing
this on ext4):
perl -e 'opendir(my $dh, $ARGV[0]) or die; my @d = grep { length($_) == 24 } readdir($dh); print("@d\n");' \
src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal
000000020000000000000005 000000020000000000000004 000000010000000000000002 000000020000000000000003 000000010000000000000003
pg_waldump with no explicit segment specification fails:
.../pg_waldump -p src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal -s 0/03000000
pg_waldump: error: invalid WAL segment size in WAL file "000000020000000000000005" (0 bytes)
pg_waldump: detail: The WAL segment size must be a power of two between 1 MB and 1 GB.
[1] https://buildfarm.postgresql.org/cgi-bin/show_log.pl?nm=turaco&dt=2026-09-24%2019%3A58%3A27
Best regards,
Alexander
^ permalink raw reply [nested|flat] 43+ messages in thread
* Re: Offline data checksum changes can cause incorrect checksum state on standbys
@ 2026-09-28 06:58 Nazir Bilal Yavuz <byavuz81@gmail.com>
parent: Alexander Lakhin <exclusion@gmail.com>
0 siblings, 0 replies; 43+ messages in thread
From: Nazir Bilal Yavuz @ 2026-09-28 06:58 UTC (permalink / raw)
To: Alexander Lakhin <exclusion@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; Zsolt Parragi <zsolt.parragi@percona.com>; Heikki Linnakangas <hlinnaka@iki.fi>; pgsql-hackers@lists.postgresql.org
Hi,
On Sat, 26 Sept 2026 at 08:00, Alexander Lakhin <exclusion@gmail.com> wrote:
>
> 14.09.2026 16:48, Daniel Gustafsson wrote:
>
> BF animal turaco (Raspberry PI, kernel 6.12.75+rpt-rpi-v8) managed to fail
> a test added in b52a1c2c8:
> [22:51:29.073](0.002s) ok 7 - replay starts at the switchover checkpoint
> [22:51:29.087](0.014s) not ok 8 - last common checkpoint is a shutdown checkpoint
> [22:51:29.088](0.001s)
> [22:51:29.088](0.000s) # Failed test 'last common checkpoint is a shutdown checkpoint'
> # at t/013_rewind.pl line 161.
> [22:51:29.089](0.001s) # ''
> # doesn't match '(?^:CHECKPOINT_SHUTDOWN)'
> [22:51:29.100](0.011s) ok 9 - rewound node keeps its own checksum state in the control file
> ...
> Test Summary Report
> -------------------
> t/013_rewind.pl (Wstat: 256 (exited 1) Tests: 13 Failed: 1)
>
> I've reproduced this failure with the following modification:
> --- a/src/backend/access/transam/xlog.c
> +++ b/src/backend/access/transam/xlog.c
> @@ -3446,2 +3446,3 @@ XLogFileInitInternal(XLogSegNo logsegno, TimeLineID logtli,
> */
> +pg_usleep(100000);
> installed_segno = logsegno;
>
> which makes all-zero 000000020000000000000005 appear inside
> $node_a->data_dir . '/pg_wal':
> tr -d '\000' < src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal/000000020000000000000005 | wc
> 0 0 0
>
> and then if readdir() happens to return this file first (I'm observing
> this on ext4):
> perl -e 'opendir(my $dh, $ARGV[0]) or die; my @d = grep { length($_) == 24 } readdir($dh); print("@d\n");' \
> src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal
> 000000020000000000000005 000000020000000000000004 000000010000000000000002 000000020000000000000003 000000010000000000000003
>
> pg_waldump with no explicit segment specification fails:
> .../pg_waldump -p src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal -s 0/03000000
> pg_waldump: error: invalid WAL segment size in WAL file "000000020000000000000005" (0 bytes)
> pg_waldump: detail: The WAL segment size must be a power of two between 1 MB and 1 GB.
Thanks for the report! I think your analysis correct. I encountered
the same problem on my local, there is another thread for fixing this
problem [1].
I can reproduce the problem with your reproducer, though not on the
first try; I needed to run the test a couple of times. Then, I confirm
that the patch in [1] fixes the problem, I run 013_rewind test 100
times and there was no failure.
[1] https://postgr.es/m/CAN55FZ1Yak_xBqMaDQsD7atpBkGLEkF-DKXcs3nLHM1Uq4YRew%40mail.gmail.com
--
Regards,
Nazir Bilal Yavuz
Microsoft
^ permalink raw reply [nested|flat] 43+ messages in thread
end of thread, other threads:[~2026-09-28 06:58 UTC | newest]
Thread overview: 43+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2026-08-12 07:55 Offline data checksum changes can cause incorrect checksum state on standbys Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-08-14 13:59 ` Daniel Gustafsson <daniel@yesql.se>
2026-08-14 15:27 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-08-26 20:37 ` Daniel Gustafsson <daniel@yesql.se>
2026-08-28 05:33 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-08-28 06:53 ` Daniel Gustafsson <daniel@yesql.se>
2026-08-28 09:24 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-08-28 11:16 ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-08-28 15:13 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-08-28 16:04 ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-08-29 02:41 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-08-29 21:36 ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-08-31 05:06 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-08-31 07:16 ` Daniel Gustafsson <daniel@yesql.se>
2026-08-31 08:01 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-08-31 10:55 ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-08-31 11:38 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-08-31 13:10 ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-08-31 15:28 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-08-31 23:31 ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-09-01 08:38 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-09-01 13:50 ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-09-02 07:45 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-09-02 12:13 ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-09-02 14:34 ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-09-03 04:08 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-09-03 11:06 ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-09-03 11:54 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-09-03 23:08 ` Daniel Gustafsson <daniel@yesql.se>
2026-09-07 11:08 ` Heikki Linnakangas <hlinnaka@iki.fi>
2026-09-07 12:26 ` Daniel Gustafsson <daniel@yesql.se>
2026-09-07 15:57 ` Heikki Linnakangas <hlinnaka@iki.fi>
2026-09-08 08:17 ` Daniel Gustafsson <daniel@yesql.se>
2026-09-08 08:58 ` Heikki Linnakangas <hlinnaka@iki.fi>
2026-09-08 09:12 ` Daniel Gustafsson <daniel@yesql.se>
2026-09-08 21:21 ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-09-09 21:45 ` Daniel Gustafsson <daniel@yesql.se>
2026-09-10 10:05 ` Daniel Gustafsson <daniel@yesql.se>
2026-09-10 13:18 ` Daniel Gustafsson <daniel@yesql.se>
2026-09-14 13:48 ` Daniel Gustafsson <daniel@yesql.se>
2026-09-26 05:00 ` Alexander Lakhin <exclusion@gmail.com>
2026-09-28 06:58 ` Nazir Bilal Yavuz <byavuz81@gmail.com>
2026-08-31 08:01 ` Zsolt Parragi <zsolt.parragi@percona.com>
This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox