From: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
To: Zsolt Parragi <zsolt.parragi@percona.com>
Cc: Daniel Gustafsson <daniel@yesql.se>
Cc: pgsql-hackers@lists.postgresql.org
Subject: Re: Offline data checksum changes can cause incorrect checksum state on standbys
Date: Wed, 2 Sep 2026 07:45:10 +0000
Message-ID: <apfUBpSn9Rj1p+f1@bdtpg> (raw)
In-Reply-To: <CAN4CZFNdfb-yFRW7Sh3FkZ3Qc91Lz-JDQN0b-070-6G35g5Ssg@mail.gmail.com>
References: <apUL3N4IE934qJ08@bdtpg>
<8DCA12FF-0199-403D-A203-666528A25DB6@yesql.se>
<apU0vKrffQDmmv68@bdtpg>
<CAN4CZFNqKg9Ts76r922cKWcOgtH32ocyghQa8NLB-qeCeC1RKg@mail.gmail.com>
<apVnwOuJBb4rHk+B@bdtpg>
<CAN4CZFMFcgfgJ99RYhax-T+YJH=CWLy6ZGANMxf77SrSirw3VQ@mail.gmail.com>
<apWdicFsT+iFxgAH@bdtpg>
<CAN4CZFN5sOQvWTXr7Dm1J1HXwczikBfMjkbFQ7QmHsr+VHd5Mg@mail.gmail.com>
<apaPDmrlhtXgKR+E@bdtpg>
<CAN4CZFNdfb-yFRW7Sh3FkZ3Qc91Lz-JDQN0b-070-6G35g5Ssg@mail.gmail.com>
Hi,
On Tue, Sep 01, 2026 at 02:50:04PM +0100, Zsolt Parragi wrote:
> Thanks!
>
> I applied these changes to v9 with some additional comment editing. I
> also squashed 0005 into 0001 because it describes what's implemented
> there, and I also tried to significantly reduce the commit message of
> 0001. Otherwise everything else is unchanged.
Thanks!
I initially thought there could be two more issues: one involving a base backup
spanning an online enable and another involving a crash during the first recovery
after pg_rewind. Further testing showed that neither was an issue.
So I'm happy with the current v9-0001 behavior. I now just have a couple of
wording comments:
=== 1
+ only then restart them. Before stopping a standby, make sure it has
+ replayed all WAL of its upstream node, for example by stopping the
+ primary first and comparing
+ <function>pg_last_wal_replay_lsn()</function> with
+ <function>pg_last_wal_receive_lsn()</function> on the standby.
Equality only proves that all received WAL has been replayed, not that all
upstream WAL was received. Maybe we should compare against the stopped
primary's shutdown checkpoint location, as 021 does?
=== 2
+ /*
+ * Mark the state as changed locally, without a WAL record. Recovery
+ * then knows the state is newer than anything the WAL carries and
+ * does not let a replayed checkpoint overwrite it. The watermark is
+ * left alone: any XLOG2_CHECKSUMS record this node had applied stays
+ * covered, and only records above it, written after this change, take
+ * effect again.
+ */
An offline change has no ordering against WAL not yet replayed, so records above
the watermark may have been written before the offline change. Maybe this should
be worded in terms of records covered by the watermark, without implying
chronological ordering?
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Reply to all the recipients using the --to and --cc options:
reply via email
To: pgsql-hackers@postgresql.org
Cc: bertranddrouvot.pg@gmail.com, zsolt.parragi@percona.com, daniel@yesql.se, pgsql-hackers@lists.postgresql.org
Subject: Re: Offline data checksum changes can cause incorrect checksum state on standbys
In-Reply-To: <apfUBpSn9Rj1p+f1@bdtpg>
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox