Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1wzsp4-004JbB-1R for pgsql-hackers@arkaria.postgresql.org; Fri, 28 Aug 2026 09:24:30 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.96) (envelope-from ) id 1wzsp2-0070A8-2l for pgsql-hackers@arkaria.postgresql.org; Fri, 28 Aug 2026 09:24:28 +0000 Received: from magus.postgresql.org ([2a02:c0:301:0:ffff::29]) by malur.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1wzsp2-0070A0-1g for pgsql-hackers@lists.postgresql.org; Fri, 28 Aug 2026 09:24:28 +0000 Received: from mail-wr1-x435.google.com ([2a00:1450:4864:20::435]) by magus.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256 (Exim 4.98.2) (envelope-from ) id 1wzsp0-00000001hL7-0Bd5 for pgsql-hackers@lists.postgresql.org; Fri, 28 Aug 2026 09:24:28 +0000 Received: by mail-wr1-x435.google.com with SMTP id ffacd0b85a97d-47fd66a094eso272626f8f.3 for ; Fri, 28 Aug 2026 02:24:25 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787909060; x=1788513860; darn=lists.postgresql.org; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:from:to:cc:subject:date:message-id:reply-to:content-type; bh=6/TnWDhWmezgD7ntKow6OEeB07rxCvXW6ctjFDHuDMs=; b=Haw7mvtzo6Wn4s6sT0InTU3DsG73/AAMEEg4ERaGJWmjELkO8ER4rDl0X/x2X0ekA2 KjfWgu76zlWTYASg14yQlGTg/L8fSKKZwe+55RapFRtmPesu/T8ibTFuo2DVFZ6ugRJX yooievaXhojex/lH/H6QmfyQFNkbJXD7fpQGJ8qToIJ8m5Su8/2/yxQo4A3Bt0rIcwCT D5SsY8ig5aYzsVhkXaqEJdZY2uIszemFmcJvtZJee1Mk8FE24QvnGDyGZzGVz1IwMQVh llFv1AamiKqppYHiUqytUws2tCF8rAfIQ0f+OkFq5v2P1t/TTG1hhRO/ym1MzYyiay12 F9JQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787909060; x=1788513860; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to:content-type; bh=6/TnWDhWmezgD7ntKow6OEeB07rxCvXW6ctjFDHuDMs=; b=aqhr8OatTKqw281D2qzCllg1Z9/x2HnMUESJ4A+wOn2KaXwIYmk8JBaVHmg8OL4pnM QWtAdE9Xx4m9E9LjftdrQj1eUop6FAGGWKACxaByS4qIO92DsAy5iEegTyK8IlqxL51d 1UKwrLjRiDT60mZFos4KamxHfTN974Etlw34uQTJQJY8qHybe5tEfUfw+A8jZ5TvEy5c 4GiX247/LzzyiWXlokPbb6U0yV4r4VC+yLYg70kKoiHU11cTG0AilK/jdwYKCBHoxAtj PKM9rX1QeNZJNQ2xDpbQiH4dofDIwKoNt631O4vDyWUIYa/6U0NIzMiEVufLgHpdmgqp JBkw== X-Gm-Message-State: AFuF++lCok+HBZsAR9i7kumdJW3FEjtrGErJhOp7RMV7MoF9zYKeWEf2 t/fA2VK29lbuWoewkc7wkFcJuiwjZwowUKv+DcDmeTCI3AVR5S/7WWDf X-Gm-Gg: AR+sD112piQhtLgr4qMWr8p4trj+97yGu1cEw6+LKGY3J3jpadaUB4GVn2BroMyO8ML sIvhHo1OG/xdw0QiR7kdDWcsBidS6tGhRy7RDV+v3ZSB/758qvC4OPpSeScozkNEIH7VA2GGp+p gDSFfS7OlaVki+YZbl+qPIVYA6/fZ8EOgiuQMKK41GMSdJLcRsPdJrAw2NBeP4K9atrSgUqFlWF SY3SWzAMIyXnBSoCVrlud5B4TeQNHjAwrv4jIt/bxjigxdH9kHIrrvyf1qO8i+Tna0dIztvqreQ MofeAVoLc2/PLfrVG1jFIyTgjEdyvGZfcVQTUtb58NR7olgfvVkSy/92LDbW3g9a+w4mf0iJZ8s vMVptt01mAyWcCcErFNA04w74Z+YFliC0y5SpjSQRP7Ei07cDP5do3LkQRYPpM7K0Ju/LeUlJOY T+VLy5zaG/HpbksFHeLtjNC7XWx3v8xqVX5SI2nbhrEzc7u+nj3D8uJsk75STmP8OPnJGJDKKGI 6jmQ+eRuJ8cgG3xcF+0RsBgXelpyoL4AiToNd1Zw7Syzhb773GcMFrpx/k= X-Received: by 2002:a05:6000:2dc9:b0:47f:96e6:70a with SMTP id ffacd0b85a97d-482f79735c9mr8162824f8f.0.1787909059960; Fri, 28 Aug 2026 02:24:19 -0700 (PDT) Received: from bdtpg (ec2-15-237-197-144.eu-west-3.compute.amazonaws.com. [15.237.197.144]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-482fbac4d0asm2742485f8f.10.2026.08.28.02.24.19 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 28 Aug 2026 02:24:19 -0700 (PDT) Date: Fri, 28 Aug 2026 09:24:17 +0000 From: Bertrand Drouvot To: Daniel Gustafsson Cc: PostgreSQL Hackers , Zsolt Parragi Subject: Re: Offline data checksum changes can cause incorrect checksum state on standbys Message-ID: References: <1C2BC974-44FA-487C-8A3B-8135A316A89B@yesql.se> <188A1307-1A92-45CA-9DBE-FB0962D3756E@yesql.se> <6370C439-408D-40F6-B956-55DF19FFE41C@yesql.se> MIME-Version: 1.0 Content-Type: multipart/mixed; boundary="5BB69PG6qBWnL8Ha" Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: <6370C439-408D-40F6-B956-55DF19FFE41C@yesql.se> List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk --5BB69PG6qBWnL8Ha Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit Hi, On Fri, Aug 28, 2026 at 08:53:38AM +0200, Daniel Gustafsson wrote: > > On 28 Aug 2026, at 07:33, Bertrand Drouvot wrote: > > > Only looking at 0001 here, I've a few comments: > > Thanks, I've yet to dig into it completely but below are a few quick questions > to help me along the way. > > > === 1 > > > > @@ -9288,17 +9572,33 @@ xlog2_redo(XLogReaderState *record) > > > > SpinLockAcquire(&XLogCtl->info_lck); > > XLogCtl->data_checksum_version = state.new_checksum_state; > > + SetLocalDataChecksumState(state.new_checksum_state); > > SpinLockRelease(&XLogCtl->info_lck); > > > > This applies every XLOG2_CHECKSUMS record encountered during recovery, even when > > the same record was applied before. > > > > For example, a standby can replay the final "on" record and then stop cleanly > > without advancing its restartpoint beyond that record. If checksums are subsequently > > disabled offline, the next startup begins from the older restartpoint and replays > > the same on record again, overriding the offline disable. > > Do you mean that checksums are disabled offline across the cluster on all > nodes, or just on the standby? Disabling checksums offline on the standby is sufficient although that is not the intended procedure. Disabling offline on both the primary and standby also produce the issue. > > === 2 > > > > If this can happen then online checksums wouldn't work at all right? You’re right, my previous explanation was not fully accurate. The 0001-specific concern is that a checkpoint can capture checkPoint.dataChecksumState as inprogress-on, then insert XLOG_CHECKPOINT_REDO correctly carrying on. The delay protects the flush, but the earlier value remains stale. The equality check then does not persist on, and recovery no longer adopts it from the REDO record, so a crash before the following checkpoint completes can resolve the state back to off. > Have you been able to construct a repro (with injection points) where a > REDO record after a CHECKSUM record carries the wrong state? Not with an injection point, but you can repro that way: In xlog.c add 3 sleeps (see repro.txt attached): - In SetDataChecksumsOn() to hold the launcher at inprogress-on. - In SetDataChecksumsOn() to park it at on before its own checkpoint. - In CreateCheckPoint() sleep/spin until XLogCtl->data_checksum_version == on. Then: start a cluster with initdb --no-data-checksums Run SELECT pg_enable_data_checksums() Then within 60s run CHECKPOINT Once the checkpoint completes (SHOW data_checksums = on but pg_controldata still shows version 3) pkill -9 the cluster restart check SHOW data_checksums: it comes back off. Regards, -- Bertrand Drouvot PostgreSQL Contributors Team RDS Open Source Databases Amazon Web Services: https://aws.amazon.com --5BB69PG6qBWnL8Ha Content-Type: text/plain; charset=us-ascii Content-Disposition: attachment; filename="repro.txt" diff --git a/src/backend/access/transam/xlog.c b/src/backend/access/transam/xlog.c index fa24b00e7d4..672ea66cb0a 100644 --- a/src/backend/access/transam/xlog.c +++ b/src/backend/access/transam/xlog.c @@ -4858,6 +4858,10 @@ SetDataChecksumsOn(void) SpinLockRelease(&XLogCtl->info_lck); INJECTION_POINT("datachecksums-enable-checksums-delay", NULL); + ereport(LOG, (errmsg("BDTTESTHACK: launcher holding at inprogress-on for 60s"))); + pg_usleep(60 * 1000000L); + ereport(LOG, (errmsg("BDTTESTHACK: launcher releasing, about to write XLOG2_CHECKSUMS(on) and flip"))); + START_CRIT_SECTION(); MyProc->delayChkptFlags |= DELAY_CHKPT_START; @@ -4874,6 +4878,10 @@ SetDataChecksumsOn(void) INJECTION_POINT("datachecksums-on-before-checkpoint", NULL); + ereport(LOG, (errmsg("BDTTESTHACK: launcher holding at inprogress-on for 60s"))); + pg_usleep(60 * 1000000L); + ereport(LOG, (errmsg("BDTTESTHACK: launcher releasing, about to write XLOG2_CHECKSUMS(on) and flip"))); + RequestCheckpoint(CHECKPOINT_FORCE | CHECKPOINT_WAIT | CHECKPOINT_FAST); INJECTION_POINT("datachecksums-on-after-checkpoint", NULL); @@ -7795,6 +7803,22 @@ CreateCheckPoint(int flags) */ WALInsertLockRelease(); + if (!shutdown && checkPoint.dataChecksumState == PG_DATA_CHECKSUM_INPROGRESS_ON) + { + int i; + ereport(LOG, (errmsg("BDTTESTHACK: checkpoint sampled dataChecksumState=inprogress-on, waiting for flip to on"))); + for (i = 0; i < 600; i++) /* up to ~60s, then give up */ + { + uint32 v; + SpinLockAcquire(&XLogCtl->info_lck); + v = XLogCtl->data_checksum_version; + SpinLockRelease(&XLogCtl->info_lck); + if (v == PG_DATA_CHECKSUM_VERSION) /* "on" */ + break; + pg_usleep(100 * 1000L); /* 100ms */ + } + ereport(LOG, (errmsg("TESTHACK: checkpoint resuming after %d iterations; will sample redo_rec and write XLOG_CHECKPOINT_REDO", i))); + } /* * If this is an online checkpoint, we have not yet determined the redo * point. We do so now by inserting the special XLOG_CHECKPOINT_REDO --5BB69PG6qBWnL8Ha--