Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1wzpDj-004HLG-0h for pgsql-hackers@arkaria.postgresql.org; Fri, 28 Aug 2026 05:33:43 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.96) (envelope-from ) id 1wzpDg-005wOS-35 for pgsql-hackers@arkaria.postgresql.org; Fri, 28 Aug 2026 05:33:40 +0000 Received: from magus.postgresql.org ([2a02:c0:301:0:ffff::29]) by malur.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1wzpDg-005wO2-1z for pgsql-hackers@lists.postgresql.org; Fri, 28 Aug 2026 05:33:40 +0000 Received: from mail-wr1-x434.google.com ([2a00:1450:4864:20::434]) by magus.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256 (Exim 4.98.2) (envelope-from ) id 1wzpDe-00000001fhL-0KOL for pgsql-hackers@lists.postgresql.org; Fri, 28 Aug 2026 05:33:40 +0000 Received: by mail-wr1-x434.google.com with SMTP id ffacd0b85a97d-47f93b2fe4cso199201f8f.0 for ; Thu, 27 Aug 2026 22:33:37 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787895217; x=1788500017; darn=lists.postgresql.org; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:from:to:cc:subject:date:message-id:reply-to:content-type; bh=l0DSSJjCJ7v6OANNNurZMdnGEbWRMLdGLc6FFyxZ91Y=; b=RjFbl/Tteu7iJsuVJJAVROJ+ftd6wBz14WosauPFOuhaAMwDS+lLTvJIE9/6Caca6o 02wZ3gB6i6kz2zZkh+DBkcm6DCzO/iQf8C2+ClJLvjqpLUPE/ZlwF7Qa8NgQbRPWLBYQ 8L5/QHbfmfqVntRlrVGkGmKbjq1btBeim+R49oL4W6nPmFArATx+SVJlYMLAoxDBLlgK CBogYhX3D0U4n+t1cuul7DHlMQVXPYqTzvHKZoJDKf5xtj2Ysq1YrZ8ZS99qe0pLtKbx bJ7w2nTH9VO68OxHKXMMxIlrWUPIfkR78JbpT/3yJbvyl0eZokOqy19HdwDH7AeP89fn J/4g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787895217; x=1788500017; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to:content-type; bh=l0DSSJjCJ7v6OANNNurZMdnGEbWRMLdGLc6FFyxZ91Y=; b=kwHnmL8q7u6O6+DAWgjr9XJeSt2LTl0Cjdqik/J1UWMMvTzb8Y/psPOMLVRinEpemp K3iUSblZomqa0Cn+jvfLkYELmyozOm8HY1/q53nOWMNRQfwjOGoQPfwTlnjpJ6HnRoQv awfReUXw2BZ1Tg9sr3T82kGcLtW+O2HyWg9PhqvjXBLIv/24z3PApj0cFL8edcpEmKO+ 8XZoYgCvMb2ic2j8iC2H4LnzU2dHXqKn4O92xhQvTb9SdNP/Nlv78pPAPhmO7pSJtv9Q NEILJEanVwKbSIq2xcTQmYNl2mH7czH8aan4g0nc3F7rmQGO2GieDEmwsEzTQj3xD8tm B7sQ== X-Gm-Message-State: AFuF++lnad8Is7caOF61KVhUlsfzVr+Q513ACP30A2ftEedS796vDguW X6xOQ/8g+SLeWGniuPaMtp95fIkzMJMXWlL5gYIGdNK40VisjU2HbGgp5Va93g== X-Gm-Gg: AR+sD13L7THvoWzV09/J2HwXLP+aJ+mvzchMb4R9ZapjODYL+khehyI0UxcH9ttqodz YSkWebiTyvVIb70L+qXq6PTNWWTd0QEBT1rVqzdmZPWljr9eP/XFaUsJGm9ZMvqDGLn++kAaF56 gkB/nZHLVPinpshKushsGHS2LsmhdHFAjVhYFzGMbId4k0/ebErVCR/ltbuJHGBN+OX/+oa3JI4 jH9sKITHS5KXA9zic9/I84cGJWK7GhWeV74/UP33tHaq/W6ilxqC7Jz1eilh34NsIsX4UckMKi4 HZ8f2nRnV2Sc4avaa9O7JB+J+f7kFVA2Lg/N8Sh6H+BwuU8ZA/C2sFM84uJ+fYc4i+NB6B9T85g hKRVY0OTpTc/etlkEAQUT7nUWdS7qxhIciNoMFvMXSSHizW2ShG4/J6+QvSUsuldRZm+xsAb+Jm Bu5B38Q40CXvkXYD4pOupkoejJr8FlTh968kIB90xS75BKNIYjR5QEjl5bzqUlLOz/D8uoJTKg7 +aoFj4eTpp+Zic81uVfuCJDjSKtNmmrY9ix4IkWG0MkCjBt X-Received: by 2002:a05:6000:4a1e:b0:482:e658:bb8e with SMTP id ffacd0b85a97d-482f79bf4f4mr6012463f8f.12.1787895216589; Thu, 27 Aug 2026 22:33:36 -0700 (PDT) Received: from bdtpg (ec2-15-237-197-144.eu-west-3.compute.amazonaws.com. [15.237.197.144]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-482fbab3f16sm1424624f8f.1.2026.08.27.22.33.35 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 27 Aug 2026 22:33:36 -0700 (PDT) Date: Fri, 28 Aug 2026 05:33:34 +0000 From: Bertrand Drouvot To: Daniel Gustafsson Cc: PostgreSQL Hackers , Zsolt Parragi Subject: Re: Offline data checksum changes can cause incorrect checksum state on standbys Message-ID: References: <1C2BC974-44FA-487C-8A3B-8135A316A89B@yesql.se> <188A1307-1A92-45CA-9DBE-FB0962D3756E@yesql.se> MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: <188A1307-1A92-45CA-9DBE-FB0962D3756E@yesql.se> List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk Hi, On Wed, Aug 26, 2026 at 10:37:09PM +0200, Daniel Gustafsson wrote: > > On 14 Aug 2026, at 17:27, Bertrand Drouvot wrote: > > > Looking forward to seeing your and Zsolt's proposals. > > This has now been worked on quite extensively by Zsolt, myself and Tomas Vondra > and a number of patchrevisions have been created and rewritten. Thanks for looking at it! > The regression in a correctly done offline change is due to combining > pg_checksums which rewrite data on risk without WAL logging (or any logging at > all) the transformation, with online checksums which WAL log the state change. > With online checksums, the local state on the standby was overwritten during > replay by the dataChecksumState in the checkpoint. Agreed. > The fix in 0001 is to not adopt the > state change from the replay of checkpoints, only from XLOG2_CHECKSUMS records, > and to alert the user with a log entry if the states mismatch. That makes sense to me and matches the intent of my v1 for this regression: keep offline changes local while still applying WAL online transitions, with the useful addition of a warning on mismatch. > The 0001 patch is the least invasive patch to solve the regression that either > of us has managed to come up with, but it's still far from trivial. Only looking at 0001 here, I've a few comments: === 1 @@ -9288,17 +9572,33 @@ xlog2_redo(XLogReaderState *record) SpinLockAcquire(&XLogCtl->info_lck); XLogCtl->data_checksum_version = state.new_checksum_state; + SetLocalDataChecksumState(state.new_checksum_state); SpinLockRelease(&XLogCtl->info_lck); This applies every XLOG2_CHECKSUMS record encountered during recovery, even when the same record was applied before. For example, a standby can replay the final "on" record and then stop cleanly without advancing its restartpoint beyond that record. If checksums are subsequently disabled offline, the next startup begins from the older restartpoint and replays the same on record again, overriding the offline disable. === 2 + * checkPoint.dataChecksumState was sampled while holding the WAL insert + * locks, so it is the state in effect at the redo point. . . . + SpinLockAcquire(&XLogCtl->info_lck); + if (checkPoint.dataChecksumState == XLogCtl->data_checksum_version) + ControlFile->data_checksum_version = checkPoint.dataChecksumState; + SpinLockRelease(&XLogCtl->info_lck); I’m not sure the "state in effect at the redo point" is always correct. There is a window between inserting the checksum transition record and updating XLogCtl->data_checksum_version. XLOG2_CHECKSUMS(on) records the target state. The transition can therefore proceed as follows: 1. The current shared state is inprogress-on. 2. XLogChecksums() inserts XLOG2_CHECKSUMS(on) and releases its WAL insertion lock. 3. Before the shared state is updated to on, the checkpoint still reads inprogress-on and inserts XLOG_CHECKPOINT_REDO. 4. The transition then updates the shared state to on. The WAL order is then: XLOG2_CHECKSUMS(on) XLOG_CHECKPOINT_REDO(inprogress-on) If the server crashes after the concurrent checkpoint from step 3 completes, but before the enabling operation’s later checkpoint completes, recovery starts from that redo point and does not replay the preceding on record. It can therefore resolve inprogress-on back to off. The equality check above does not repair this ordering. === 3 + if ((haveBackupLabel || XLogRecPtrIsValid(ControlFile->backupStartPoint)) && + !(XLogRecPtrIsValid(ControlFile->backupEndPoint) && + ControlFile->backupEndRequired)) + { + if (wasShutdown) + AdoptReplayedDataChecksumState(checkPoint.dataChecksumState); + else + adoptChecksumStateFromNextCheckpoint = true; + } IIUC, this can overwrite the checksum state copied from the source during pg_rewind with the state from the last common checkpoint. For example, if the common checkpoint says off, but both source and target were enabled offline after divergence, recovery adopts off. Since the offline enable has no XLOG2_CHECKSUMS record, nothing restores on. I have only looked at 0001 for this point, so I don't know whether one of the following patches handles this case. === 4 + /* + * Note the data checksum state the flush below starts under. Replay runs + * concurrently and can change the state while the flush is in progress, + * in which case the flush covers pages written under both states; see + * where the state is persisted further down. + */ + SpinLockAcquire(&XLogCtl->info_lck); + checksum_state = XLogCtl->data_checksum_version; + SpinLockRelease(&XLogCtl->info_lck); + CheckPointGuts(lastCheckPoint.redo, flags); + * Persist only if the flush above ran under one state throughout; see + * CreateCheckPoint() for why. + */ + SpinLockAcquire(&XLogCtl->info_lck); + if (checksum_state == XLogCtl->data_checksum_version) + ControlFile->data_checksum_version = checksum_state; + SpinLockRelease(&XLogCtl->info_lck); I think comparing only data_checksum_version cannot detect a complete on->off->on transition during CheckPointGuts(). The initial and final values match even though the flush ran under multiple states. FWIW, while v4-0001 may address other issues present in v1, v1 would avoid the specific cases described in === 1 through === 3. Some parts of it may therefore be worth considering here. Regards, -- Bertrand Drouvot PostgreSQL Contributors Team RDS Open Source Databases Amazon Web Services: https://aws.amazon.com