Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.98.2) (envelope-from ) id 1xAKW9-00000002a5B-0FvA for pgsql-hackers@arkaria.postgresql.org; Sat, 26 Sep 2026 05:00:09 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.98.2) (envelope-from ) id 1xAKW7-00000003eEO-1aha for pgsql-hackers@arkaria.postgresql.org; Sat, 26 Sep 2026 05:00:07 +0000 Received: from makus.postgresql.org ([2001:4800:3e1:1::229]) by malur.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.98.2) (envelope-from ) id 1xAKW6-00000003eED-41xA for pgsql-hackers@lists.postgresql.org; Sat, 26 Sep 2026 05:00:07 +0000 Received: from mail-wr2-x10.google.com ([2a00:1450:4864:30::10]) by makus.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256 (Exim 4.98.2) (envelope-from ) id 1xAKW4-00000001J2z-0mDm for pgsql-hackers@lists.postgresql.org; Sat, 26 Sep 2026 05:00:05 +0000 Received: by mail-wr2-x10.google.com with SMTP id ffacd0b85a97d-4843796e373so870136f8f.1 for ; Fri, 25 Sep 2026 22:00:04 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790398802; x=1791003602; darn=lists.postgresql.org; h=in-reply-to:from:content-language:references:cc:to:subject :user-agent:mime-version:date:message-id:content-type:from:to:cc :subject:date:message-id:reply-to:content-type; bh=2fKbByqGtGwf9vaUGcvtLdkp9tQFobeUAuzV76ZEtIY=; b=OjfBUxwdy9cP+SHDDaLQEHdWAvEZ9NW+ni9xw/3z76dA9Ps8Pnj/8ixQszzieOYONW IJFvSysPvgbKCCprjzfBzO6JSzD5cAOcoQVGLWCNVEQiRc4gA+uki+hZ4+DaorBllY5Z CifK07vjzo/pxjgifV89IZW1POQJmqtbhB/d4OeRPCIyXgKRqmTLJJtN2vn2/vnBp1lK mp53+VH5Xk+WMN45cvQ0e7PKgDWv+TkR45NWf1QtQkY9I/ZlEGGTwJHO8rFTi+LL6bj1 tO6pk6scZnQEVCZXNqGe5Lk60NxJ9mey/xf2aNkojH13NZM1o8jr7753QvWrQ1IN2YOx vr2g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790398802; x=1791003602; h=in-reply-to:from:content-language:references:cc:to:subject :user-agent:mime-version:date:message-id:content-type:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=2fKbByqGtGwf9vaUGcvtLdkp9tQFobeUAuzV76ZEtIY=; b=gS2tImKQKg7tK4O0Cjo6d7F2R74O1HZ9ziwtin/uQsOXMEXaFLexJN7HRTEwGl3ppg v0ebISDbNVKY9VfTP3wMe3708nNfWQxmlBJJwmLzO9RyIRNYnoiqgLaA3JeVfwtg8NMm smKfNDg5ozXFyltYqdT87ZwzQKVzn5KYMdj+LEBX7ehejWoLmeMsGfn+RFuFqkWZW0Qr Wn/P+9b0JWAgkPqqfVQBdedWSp1eBDtimr5I21wk5O2D7LqOwUtaZW2uX0GemwAk0n+L n7PgV+TCDvEhQe5SXf7/M7NvhATcJpy/zhIEa09TBnGOnJ35tv/0Kkky3SiW2vYRAdWt FGfw== X-Forwarded-Encrypted: i=1; AKwUvBzoDbfp5+DTz2BugwIQn6ZGuwQ8lgD00hG4ivjHHK4QSYOXt57uACxBcQ3q6DQx+Gw4/g/cVTO2ADrVBX+S@lists.postgresql.org X-Gm-Message-State: AFuF++lfwMB+Tn3RrBnkqppAh/i2X0/tbK0qIMAqn+mvO90V7w8KXPqU NNTuvtMpXljqRFjeb/A/dkWhH1fiUCIXSiDlCssWSBvtRQnUfsLg0R19 X-Gm-Gg: AYBFou05e1ZXZYS10q6X/75O40btTDV1fNBV52gl72lt5Y0SE/g1Y3GGxBDUG485Rc/ VJDCtqzyVYP6xFezxBBbZyRX/DCpq8+ZP0ZeJoqAAXZQ+mLXuDbf5zWhWAJUf+g1GWcLbeYUPGB aXvuWmJAe7FdSjJRnvUqjqbiRYHlRJsZEUdOMwAUjv5nfov8jiPtTlgLD+2ftNSgQoY+Z3kMdcM kWf6TCloY7yJXjDWLVLqvRkD2mBHfz3xfWv4qzPmxnFYrrKmuR/wXRVW8Op4Uht9TnniV9U5na4 i0t9H7jNk5RbuGBXR4RvrbxNrOkZlsgtWlkB2j+9IT0g9lag4dT/sYeZroRZeGF4ZeSEdKDSY8U nPVNXgBsFDRb1wwVMr8+mxSqT31bi23zyVBdEoPToIMfA3ER3EBpdDP2ReoQy2SW9S7vw36Ej2r LKCkPFG2QMxjMdxqiisHoM+Q7a1LfCXHPYQCcaQWf0cUm3mIPazhMd5bH/+56DM6CUPeWsdF0Lc ldVPfpjHFE= X-Received: by 2002:adf:e18d:0:b0:488:84bf:eed4 with SMTP id ffacd0b85a97d-48884bff1e6mr2492710f8f.47.1790398802230; Fri, 25 Sep 2026 22:00:02 -0700 (PDT) Received: from [192.168.0.50] ([89.149.68.133]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-4887a6456f8sm11054935f8f.26.2026.09.25.22.00.01 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Fri, 25 Sep 2026 22:00:01 -0700 (PDT) Content-Type: multipart/alternative; boundary="------------IpAZYOvddcD4kX0VoI5BYV80" Message-ID: Date: Sat, 26 Sep 2026 08:00:00 +0300 MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: Offline data checksum changes can cause incorrect checksum state on standbys To: Daniel Gustafsson Cc: Bertrand Drouvot , Zsolt Parragi , Heikki Linnakangas , pgsql-hackers@lists.postgresql.org References: <97139167-3B7B-4B00-B623-2DD5C0656DDB@yesql.se> <5b08b4d1-0982-4b77-bac5-3bdffc6583f5@iki.fi> <166AD46C-25A4-4AF3-B65F-095E3746A96E@yesql.se> <4BCE2F21-B669-46F3-B781-010F71B0B065@yesql.se> <0E4734FA-4764-405C-9670-C45DF440A718@yesql.se> Content-Language: en-US From: Alexander Lakhin In-Reply-To: List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk This is a multi-part message in MIME format. --------------IpAZYOvddcD4kX0VoI5BYV80 Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit Hello Daniel, 14.09.2026 16:48, Daniel Gustafsson wrote: > Pushed to master along with the cleanups from the "Trying to break online > checksums with LLMs" thread. Not backpatched yet, as I wanted a few builds in > the BF first and also some input on the new open item in the other thread. BF animal turaco (Raspberry PI, kernel 6.12.75+rpt-rpi-v8) managed to fail a test added in b52a1c2c8: [22:51:29.073](0.002s) ok 7 - replay starts at the switchover checkpoint [22:51:29.087](0.014s) not ok 8 - last common checkpoint is a shutdown checkpoint [22:51:29.088](0.001s) [22:51:29.088](0.000s) #   Failed test 'last common checkpoint is a shutdown checkpoint' #   at t/013_rewind.pl line 161. [22:51:29.089](0.001s) #                   '' #     doesn't match '(?^:CHECKPOINT_SHUTDOWN)' [22:51:29.100](0.011s) ok 9 - rewound node keeps its own checksum state in the control file ... Test Summary Report ------------------- t/013_rewind.pl                      (Wstat: 256 (exited 1) Tests: 13 Failed: 1) I've reproduced this failure with the following modification: --- a/src/backend/access/transam/xlog.c +++ b/src/backend/access/transam/xlog.c @@ -3446,2 +3446,3 @@ XLogFileInitInternal(XLogSegNo logsegno, TimeLineID logtli,          */ +pg_usleep(100000);         installed_segno = logsegno; which makes all-zero 000000020000000000000005 appear inside $node_a->data_dir . '/pg_wal': tr -d '\000' < src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal/000000020000000000000005 | wc       0       0       0 and then if readdir() happens to return this file first (I'm observing this on ext4): perl -e 'opendir(my $dh, $ARGV[0]) or die; my @d = grep { length($_) == 24 } readdir($dh); print("@d\n");' \ src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal 000000020000000000000005 000000020000000000000004 000000010000000000000002 000000020000000000000003 000000010000000000000003 pg_waldump with no explicit segment specification fails: .../pg_waldump -p src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal -s 0/03000000 pg_waldump: error: invalid WAL segment size in WAL file "000000020000000000000005" (0 bytes) pg_waldump: detail: The WAL segment size must be a power of two between 1 MB and 1 GB. [1] https://buildfarm.postgresql.org/cgi-bin/show_log.pl?nm=turaco&dt=2026-09-24%2019%3A58%3A27 Best regards, Alexander --------------IpAZYOvddcD4kX0VoI5BYV80 Content-Type: text/html; charset=UTF-8 Content-Transfer-Encoding: 8bit
Hello Daniel,

14.09.2026 16:48, Daniel Gustafsson wrote:
Pushed to master along with the cleanups from the "Trying to break online
checksums with LLMs" thread.  Not backpatched yet, as I wanted a few builds in
the BF first and also some input on the new open item in the other thread.

BF animal turaco (Raspberry PI, kernel 6.12.75+rpt-rpi-v8) managed to fail
a test added in b52a1c2c8:
[22:51:29.073](0.002s) ok 7 - replay starts at the switchover checkpoint
[22:51:29.087](0.014s) not ok 8 - last common checkpoint is a shutdown checkpoint
[22:51:29.088](0.001s)
[22:51:29.088](0.000s) #   Failed test 'last common checkpoint is a shutdown checkpoint'
#   at t/013_rewind.pl line 161.
[22:51:29.089](0.001s) #                   ''
#     doesn't match '(?^:CHECKPOINT_SHUTDOWN)'
[22:51:29.100](0.011s) ok 9 - rewound node keeps its own checksum state in the control file
...
Test Summary Report
-------------------
t/013_rewind.pl                      (Wstat: 256 (exited 1) Tests: 13 Failed: 1)

I've reproduced this failure with the following modification:
--- a/src/backend/access/transam/xlog.c
+++ b/src/backend/access/transam/xlog.c
@@ -3446,2 +3446,3 @@ XLogFileInitInternal(XLogSegNo logsegno, TimeLineID logtli,
         */
+pg_usleep(100000);
        installed_segno = logsegno;

which makes all-zero 000000020000000000000005 appear inside
$node_a->data_dir . '/pg_wal':
tr -d '\000' < src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal/000000020000000000000005 | wc
      0       0       0

and then if readdir() happens to return this file first (I'm observing
this on ext4):
perl -e 'opendir(my $dh, $ARGV[0]) or die; my @d = grep { length($_) == 24 } readdir($dh); print("@d\n");' \
src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal
000000020000000000000005 000000020000000000000004 000000010000000000000002 000000020000000000000003 000000010000000000000003

pg_waldump with no explicit segment specification fails:
.../pg_waldump -p src/test/modules/test_checksums/tmp_check/t_013_rewind_node_a_data/pgdata/pg_wal -s 0/03000000
pg_waldump: error: invalid WAL segment size in WAL file "000000020000000000000005" (0 bytes)
pg_waldump: detail: The WAL segment size must be a power of two between 1 MB and 1 GB.


[1] https://buildfarm.postgresql.org/cgi-bin/show_log.pl?nm=turaco&dt=2026-09-24%2019%3A58%3A27

Best regards,
Alexander
--------------IpAZYOvddcD4kX0VoI5BYV80--