pg.ddx.io  pgsql-hackers@postgresql.org mailing list archive  
help / color / mirror / Atom feed
From: Alexander Lakhin <exclusion@gmail.com>
To: Daniel Gustafsson <daniel@yesql.se>
Cc: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Cc: Zsolt Parragi <zsolt.parragi@percona.com>
Cc: Heikki Linnakangas <hlinnaka@iki.fi>
Cc: pgsql-hackers@lists.postgresql.org
Subject: Re: Offline data checksum changes can cause incorrect checksum state on standbys
Date: Sat, 26 Sep 2026 08:00:00 +0300
Message-ID: <ffd837c5-dc22-4db1-a2d8-b623cbe9b30c@gmail.com> (raw)
In-Reply-To: <B151956B-1117-4FD1-A3BC-C6C561EE07F8@yesql.se>
References: <CAN4CZFMFcgfgJ99RYhax-T+YJH=CWLy6ZGANMxf77SrSirw3VQ@mail.gmail.com>
	<apWdicFsT+iFxgAH@bdtpg>
	<CAN4CZFN5sOQvWTXr7Dm1J1HXwczikBfMjkbFQ7QmHsr+VHd5Mg@mail.gmail.com>
	<apaPDmrlhtXgKR+E@bdtpg>
	<CAN4CZFNdfb-yFRW7Sh3FkZ3Qc91Lz-JDQN0b-070-6G35g5Ssg@mail.gmail.com>
	<apfUBpSn9Rj1p+f1@bdtpg>
	<CAN4CZFOOHZmVnL-B2D+V7ODmVFgkYucgQD+-qegDTKZgJ3dtQg@mail.gmail.com>
	<CAN4CZFP_-vMOtHEj0u_v9OYoihxmi1QkfP_wiP_ytdg9FE=KfA@mail.gmail.com>
	<apjysxMee6zvSCnk@bdtpg>
	<CAN4CZFNG57F8DY5NF=y=bdbCju6+_JxrBnUVYg5Xz3LOMCrSug@mail.gmail.com>
	<aplf7lfka5ER+PhB@bdtpg>
	<97139167-3B7B-4B00-B623-2DD5C0656DDB@yesql.se>
	<5b08b4d1-0982-4b77-bac5-3bdffc6583f5@iki.fi>
	<166AD46C-25A4-4AF3-B65F-095E3746A96E@yesql.se>
	<4BCE2F21-B669-46F3-B781-010F71B0B065@yesql.se>
	<0E4734FA-4764-405C-9670-C45DF440A718@yesql.se>
	<B151956B-1117-4FD1-A3BC-C6C561EE07F8@yesql.se>

Hello Daniel,

14.09.2026 16:48, Daniel Gustafsson wrote:
> Pushed to master along with the cleanups from the "Trying to break online
> checksums with LLMs" thread.  Not backpatched yet, as I wanted a few builds in
> the BF first and also some input on the new open item in the other thread.

BF animal turaco (Raspberry PI, kernel 6.12.75+rpt-rpi-v8) managed to fail
a test added in b52a1c2c8:
[22:51:29.073](0.002s) ok 7 - replay starts at the switchover checkpoint
[22:51:29.087](0.014s) not ok 8 - last common checkpoint is a shutdown checkpoint
[22:51:29.088](0.001s)
[22:51:29.088](0.000s) #   Failed test 'last common checkpoint is a shutdown checkpoint'
#   at t/013_rewind.pl line 161.
[22:51:29.089](0.001s) #                   ''
#     doesn't match '(?^:CHECKPOINT_SHUTDOWN)'
[22:51:29.100](0.011s) ok 9 - rewound node keeps its own checksum state in the control file
...
Test Summary Report
-------------------
t/013_rewind.pl                      (Wstat: 256 (exited 1) Tests: 13 Failed: 1)

I've reproduced this failure with the following modification:
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -3446,2 +3446,3 @@ XLogFileInitInternal(XLogSegNo logsegno, TimeLineID logtli,
          */
+pg_usleep(100000);
         installed_segno = logsegno;

which makes all-zero 000000020000000000000005 appear inside
$node_a->data_dir . '/pg_wal':
tr -d '\000' < src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal/000000020000000000000005 
| wc
       0       0       0

and then if readdir() happens to return this file first (I'm observing
this on ext4):
perl -e 'opendir(my $dh, $ARGV[0]) or die; my @d = grep { length($_) == 24 } readdir($dh); print("@d\n");' \
src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal
000000020000000000000005 000000020000000000000004 000000010000000000000002 000000020000000000000003 000000010000000000000003

pg_waldump with no explicit segment specification fails:
.../pg_waldump -p src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal -s 0/03000000
pg_waldump: error: invalid WAL segment size in WAL file "000000020000000000000005" (0 bytes)
pg_waldump: detail: The WAL segment size must be a power of two between 1 MB and 1 GB.


[1] https://buildfarm.postgresql.org/cgi-bin/show_log.pl?nm=turaco&dt=2026-09-24%2019%3A58%3A27

Best regards,
Alexander

view thread (43+ messages)  latest in thread

Message-ID: <ffd837c5-dc22-4db1-a2d8-b623cbe9b30c@gmail.com>
Permalink:  ../ffd837c5-dc22-4db1-a2d8-b623cbe9b30c@gmail.com/
Also on:    postgresql.org/message-id/ffd837c5-dc22-4db1-a2d8-b623cbe9b30c@gmail.com

 · 

reply

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Reply to all the recipients using the --to and --cc options:
  reply via email

  To: pgsql-hackers@postgresql.org
  Cc: exclusion@gmail.com, daniel@yesql.se, bertranddrouvot.pg@gmail.com, zsolt.parragi@percona.com, hlinnaka@iki.fi, pgsql-hackers@lists.postgresql.org
  Subject: Re: Offline data checksum changes can cause incorrect checksum state on standbys
  In-Reply-To: <ffd837c5-dc22-4db1-a2d8-b623cbe9b30c@gmail.com>

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox