agora inbox for pgsql-bugs@postgresql.org  
help / color / mirror / Atom feed
PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation
7+ messages / 4 participants
[nested] [flat]

* PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation
@ 2026-09-27 06:07  Jacky Nguyen <nktpro@gmail.com>
  0 siblings, 1 reply; 7+ messages in thread

From: Jacky Nguyen @ 2026-09-27 06:07 UTC (permalink / raw)
  To: pgsql-bugs@lists.postgresql.org

Hi PostgreSQL team,

We encountered a standby startup PANIC after a switchover on PostgreSQL
18.6. A standalone reproducer and recovery TAP test reproduce it on 18.6,
17.11, REL_18_STABLE, and master; the same test passes on 18.4 and 17.10.

*Expected:* a standby that has replayed a heap/visibility-map truncation
restarts recovery successfully.
*Actual on restart:*

LOG:  redo starts at 0/3059738
WARNING:  page 0 of relation base/5/16396_vm does not exist
PANIC:  WAL contains references to invalid pages
LOG:  startup process (PID 55286) was terminated by signal 6: Abort trap: 6

*Reproduction:* Build PostgreSQL with --enable-tap-tests
--enable-injection-points and install injection_points. Copy the attached t/
058_vm_truncate_invalid_pages.pl into src/test/recovery/t/, then run:

make -C src/test/recovery check PROVE_TESTS=t/058_vm_truncate_invalid_pages.pl

The test holds a standby restartpoint, promotes that standby, deletes and
vacuums an all-visible table (truncating its heap and VM forks), rejoins
the old primary as a standby, and restarts it. On 18.6, 20/20 TAP runs
failed with this PANIC; on 18.4, 20/20 passed. A shell reproducer, exact
steps, WAL excerpts, version matrix, and two candidate patches are included
in the attachment.

The affected 18.6 build is REL_18_6 (724edf9); 17.11 is REL_17_11
(083ac03). The runs were on macOS 15 / Darwin 24.6.0, arm64, clang 21.1.8,
with debug and injection points enabled and without cassert. The original
symptom was on Linux in a three-node streaming setup. The analysis in
REPORT.md points to VM block reads without an FPI during redo after an
already replayed VM truncation; that diagnosis and the candidate fixes are
provided for review.

Regards,
Jacky Nguyen

Attachments:

  [application/zip] pg18-vm-invalid-pages-reporting.zip (24.4K, ../../CAL4mQLAp562c1rCgg2Dqx6TBSdOk6vOvaLFjwFG9E6k4uwvJJw@mail.gmail.com/3-pg18-vm-invalid-pages-reporting.zip)
  download

^ permalink  raw  reply  [nested|flat] 7+ messages in thread

* Re: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation
@ 2026-09-27 10:33  Kirill Reshke <reshkekirill@gmail.com>
  parent: Jacky Nguyen <nktpro@gmail.com>
  0 siblings, 1 reply; 7+ messages in thread

From: Kirill Reshke @ 2026-09-27 10:33 UTC (permalink / raw)
  To: nktpro@gmail.com; +Cc: pgsql-bugs@lists.postgresql.org

On Sun, 27 Sept 2026 at 11:08, Jacky Nguyen <nktpro@gmail.com> wrote:
>
> Hi PostgreSQL team,
>
> We encountered a standby startup PANIC after a switchover on PostgreSQL 18.6. A standalone reproducer and recovery TAP test reproduce it on 18.6, 17.11, REL_18_STABLE, and master; the same test passes on 18.4 and 17.10.
>
> Expected: a standby that has replayed a heap/visibility-map truncation restarts recovery successfully.
> Actual on restart:
>
> LOG:  redo starts at 0/3059738
> WARNING:  page 0 of relation base/5/16396_vm does not exist
> PANIC:  WAL contains references to invalid pages
> LOG:  startup process (PID 55286) was terminated by signal 6: Abort trap: 6
>
> Reproduction: Build PostgreSQL with --enable-tap-tests --enable-injection-points and install injection_points. Copy the attached t/058_vm_truncate_invalid_pages.pl into src/test/recovery/t/, then run:
>
> make -C src/test/recovery check PROVE_TESTS=t/058_vm_truncate_invalid_pages.pl
>
> The test holds a standby restartpoint, promotes that standby, deletes and vacuums an all-visible table (truncating its heap and VM forks), rejoins the old primary as a standby, and restarts it. On 18.6, 20/20 TAP runs failed with this PANIC; on 18.4, 20/20 passed. A shell reproducer, exact steps, WAL excerpts, version matrix, and two candidate patches are included in the attachment.
>
> The affected 18.6 build is REL_18_6 (724edf9); 17.11 is REL_17_11 (083ac03). The runs were on macOS 15 / Darwin 24.6.0, arm64, clang 21.1.8, with debug and injection points enabled and without cassert. The original symptom was on Linux in a three-node streaming setup. The analysis in REPORT.md points to VM block reads without an FPI during redo after an already replayed VM truncation; that diagnosis and the candidate fixes are provided for review.
>
> Regards,
> Jacky Nguyen

Hi!
Thanks for report. Looks like this bug existed since first VM commit,
which was already missing invalid page guard from [0].
So, case from [0] reintroduces one we started to register VM pages.

Also I prefer candidate fix b from your email

[0] https://github.com/postgres/postgres/commit/defe93463c69f8e0bb717294a34d67c34ac0b03f


-- 
Best regards,
Kirill Reshke






^ permalink  raw  reply  [nested|flat] 7+ messages in thread

* Re: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation
@ 2026-09-28 06:37  shihao zhong <zhong950419@gmail.com>
  parent: Kirill Reshke <reshkekirill@gmail.com>
  0 siblings, 1 reply; 7+ messages in thread

From: shihao zhong @ 2026-09-28 06:37 UTC (permalink / raw)
  To: Kirill Reshke <reshkekirill@gmail.com>; +Cc: nktpro@gmail.com, pgsql-bugs@lists.postgresql.org, Melanie Plageman <melanieplageman@gmail.com>

Hi Jacky and Kirill,

Melanie's v2-0002 in the VM clear thread [1] is the same change as
candidate (a), and it fixes this report too. On master the attached test
fails without it and passes with it. The same change also passes on 17,
18 and 19.

This does not need a switchover. With full_page_writes = off, a plain
standby restart hits it if the restartpoint is before a DELETE or UPDATE
that cleared VM bits and a later VACUUM truncated the VM. So it would be
good to get v2-0002 in before the next minor release.

I asked Fable to create a test for me.

The attached clears VM bits by delete, HOT update
and cross-page update, and fails if any one of the three reads still uses
RBM_NORMAL.

Candidate (b) would also work, but once the read no longer uses
RBM_NORMAL, no invalid page is logged for the VM, so it is not needed.

Until the fix ships, ignore_invalid_pages = on lets the standby start. It
can be turned off after the next restartpoint.

[1]
https://postgr.es/m/CAAKRu_bApoksLDb-HX0GYciU3uLWqA1JagntaV8GP0%3D%2BidehHw%40mail.gmail.com

Thanks,
Shihao

Attachments:

  [application/octet-stream] v1-0001-Add-test-for-standby-restart-after-VM-truncation.patch (4.1K, ../../CAGRkXqRZva+Wiw6jb9-owhLjGOtoDn-btoFjJmejmJYOMSEbzg@mail.gmail.com/3-v1-0001-Add-test-for-standby-restart-after-VM-truncation.patch)
  download | inline diff:
From d4175f7ffc623c8d4acc283e51d8865017fa14e3 Mon Sep 17 00:00:00 2001
From: Shihao <zhong950419@gmail.com>
Date: Sun, 27 Sep 2026 09:58:16 -0700
Subject: [PATCH v1] Add test for standby restart after VM truncation

With full_page_writes off, records that clear VM bits carry no image of
the VM page. The test clears VM bits by delete, same-page update and
cross-page update, truncates the tables, and restarts the standby from a
restartpoint taken before those changes.

Discussion: https://postgr.es/m/CAL4mQLAp562c1rCgg2Dqx6TBSdOk6vOvaLFjwFG9E6k4uwvJJw@mail.gmail.com
---
 src/test/recovery/meson.build                |  1 +
 src/test/recovery/t/058_vm_clear_truncate.pl | 81 ++++++++++++++++++++
 2 files changed, 82 insertions(+)
 create mode 100644 src/test/recovery/t/058_vm_clear_truncate.pl

diff --git a/src/test/recovery/meson.build b/src/test/recovery/meson.build
index ebb12dd8766..bb28f9cfb81 100644
--- a/src/test/recovery/meson.build
+++ b/src/test/recovery/meson.build
@@ -66,6 +66,7 @@ tests += {
       't/055_cascade_reconnect.pl',
       't/056_standby_snapshot_export.pl',
       't/057_snapshot_commit_race.pl',
+      't/058_vm_clear_truncate.pl',
     ],
   },
 }
diff --git a/src/test/recovery/t/058_vm_clear_truncate.pl b/src/test/recovery/t/058_vm_clear_truncate.pl
new file mode 100644
index 00000000000..ee1f805b8eb
--- /dev/null
+++ b/src/test/recovery/t/058_vm_clear_truncate.pl
@@ -0,0 +1,81 @@
+
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A standby must be able to restart when WAL it replays again clears
+# visibility map bits on a VM page that a later, already replayed,
+# truncation removed.  With full_page_writes off, the clearing records
+# carry no image of the VM page, so redo has to cope with the page not
+# existing.
+use strict;
+use warnings FATAL => 'all';
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+	'postgresql.conf', qq{
+full_page_writes = off
+autovacuum = off
+});
+$primary->start;
+$primary->backup('bkp');
+
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'bkp', has_streaming => 1);
+$standby->start;
+
+# Make every heap page all-visible, then make the standby create a
+# restartpoint, so that a restart replays the changes below again.
+$primary->safe_psql(
+	'postgres', q{
+CREATE TABLE vm_del (a int);
+CREATE TABLE vm_hot (a int) WITH (fillfactor = 50);
+CREATE TABLE vm_upd (a int);
+INSERT INTO vm_del SELECT generate_series(1, 1000);
+INSERT INTO vm_hot SELECT generate_series(1, 1000);
+INSERT INTO vm_upd SELECT generate_series(1, 1000);
+VACUUM (FREEZE) vm_del, vm_hot, vm_upd;
+CHECKPOINT;
+});
+$primary->wait_for_replay_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT');
+
+my $start_lsn =
+  $primary->safe_psql('postgres', 'SELECT pg_current_wal_insert_lsn()');
+
+# Clear VM bits through delete, same-page update (old VM block) and
+# cross-page update (new VM block), then truncate all three tables to
+# zero blocks.
+$primary->safe_psql(
+	'postgres', q{
+DELETE FROM vm_del;
+UPDATE vm_hot SET a = -a WHERE a = 1;
+UPDATE vm_upd SET a = -a WHERE a = 1;
+DELETE FROM vm_hot;
+DELETE FROM vm_upd;
+VACUUM vm_del, vm_hot, vm_upd;
+});
+is( $primary->safe_psql(
+		'postgres',
+		"SELECT sum(pg_relation_size(c, 'vm')) FROM unnest('{vm_del,vm_hot,vm_upd}'::regclass[]) c"
+	),
+	'0',
+	'VMs truncated on primary');
+$primary->wait_for_replay_catchup($standby);
+
+$standby->stop;
+my $log_offset = -s $standby->logfile;
+my $ret = $standby->start(fail_ok => 1);
+
+my $log = slurp_file($standby->logfile, $log_offset);
+my ($redo_lsn) = $log =~ /redo starts at ([0-9A-F]+\/[0-9A-F]+)/;
+ok( defined($redo_lsn)
+	  && $primary->safe_psql('postgres',
+		"SELECT '$redo_lsn'::pg_lsn < '$start_lsn'::pg_lsn") eq 't',
+	'redo after restart starts before the VM bits were cleared');
+ok($ret, 'standby restarts after replaying VM truncation');
+unlike($log, qr/invalid pages/, 'no invalid page references in standby log');
+
+done_testing();
-- 
2.37.1 (Apple Git-137.1)



^ permalink  raw  reply  [nested|flat] 7+ messages in thread

* Re: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation
@ 2026-10-01 15:07  rahul@rhyadav.dev
  parent: shihao zhong <zhong950419@gmail.com>
  0 siblings, 2 replies; 7+ messages in thread

From: rahul@rhyadav.dev @ 2026-10-01 15:07 UTC (permalink / raw)
  To: shihao zhong <zhong950419@gmail.com>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Nktpro <nktpro@gmail.com>; Pgsql Bugs <pgsql-bugs@lists.postgresql.org>; Melanie Plageman <melanieplageman@gmail.com>

Hi,

On Mon, 28 Sep 2026, shihao zhong wrote:
> Melanie's v2-0002 in the VM clear thread [1] is the same change as
> candidate (a), and it fixes this report too.

I tested Melanie's v2 series as posted (0001-0003, which apply to
REL_19_STABLE) and candidate (a) on REL_18_STABLE, using the steps
from Shihao's test: full_page_writes = off and a plain standby
restart.  Both were debug builds with assertions, and the fix works:

- Unpatched, 18 and 19 fail to restart with "WAL contains references
  to invalid pages" after "page 0 of relation ..._vm does not exist"
  for each table.
- Patched, the standby restarts and its VM forks end up truncated
  again.
- On 18, reverting any one of the three RBM_ZERO_ON_ERROR changes
  brings the PANIC back, so the test covers each of them.

However, with wal_consistency_checking = all on the primary, the
patched standby still fails to restart:

  FATAL:  invalid page pd_lower 0 pd_upper 0 pd_special 0
  CONTEXT:  WAL redo at 0/03A80068 for Heap/DELETE: ...;
  blkref #1: rel 1663/5/16384, fork 2, blk 0 FPW

RBM_ZERO_ON_ERROR recreates the truncated VM page as all zeros.
verifyBackupPageConsistency() then masks it with heap_mask(), and
mask_unused_space() rejects a page with pd_lower 0.
heap_xlog_prune_freeze() and heap_xlog_multi_insert() already
initialize a VM page that was read as zeros, and doing the same at
the three VM clear sites fixes it.  The attached diff applies on top
of v2 on REL_19_STABLE; with it, the standby restarts with and
without wal_consistency_checking.  On 18, candidate (b) alone also
passes with wal_consistency_checking, since it never recreates the
page.

This only matters with wal_consistency_checking, but buildfarm
animals that use it could trip over Shihao's test once it's in.

Also, v2-0002 applies only on top of v2-0001 on REL_19_STABLE.  On
master it doesn't apply after master-v2-0001, 18 needs its own
version (candidate (a) applies there as is), and in 17 the redo code
is in heapam.c.

[1] https://postgr.es/m/CAAKRu_bApoksLDb-HX0GYciU3uLWqA1JagntaV8GP0%3D%2BidehHw%40mail.gmail.com

Regards,
Rahul Yadav
diff --git a/src/backend/access/heap/heapam_xlog.c b/src/backend/access/heap/heapam_xlog.c
index 4d99d99080..b87f38befc 100644
--- a/src/backend/access/heap/heapam_xlog.c
+++ b/src/backend/access/heap/heapam_xlog.c
@@ -89,6 +89,10 @@ heap_xlog_vm_clear(XLogReaderState *record,
 									  RBM_ZERO_ON_ERROR, false,
 									  &vmbuffer) == BLK_NEEDS_REDO)
 	{
+		/* initialize the page if it was read as zeros */
+		if (PageIsNew(BufferGetPage(vmbuffer)))
+			PageInit(BufferGetPage(vmbuffer), BLCKSZ, 0);
+
 		if (visibilitymap_clear(reln, heap_blkno, vmbuffer, flags))
 			PageSetLSN(BufferGetPage(vmbuffer), lsn);
 	}
@@ -855,6 +859,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 										  RBM_ZERO_ON_ERROR, false,
 										  &vmbuffer_new) == BLK_NEEDS_REDO)
 		{
+			/* initialize the page if it was read as zeros */
+			if (PageIsNew(BufferGetPage(vmbuffer_new)))
+				PageInit(BufferGetPage(vmbuffer_new), BLCKSZ, 0);
+
 			/*
 			 * If both the old and new heap pages were all-visible and their
 			 * VM bits are on the same VM page, that single VM page is
@@ -894,6 +902,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 										  RBM_ZERO_ON_ERROR, false,
 										  &vmbuffer_old) == BLK_NEEDS_REDO)
 		{
+			/* initialize the page if it was read as zeros */
+			if (PageIsNew(BufferGetPage(vmbuffer_old)))
+				PageInit(BufferGetPage(vmbuffer_old), BLCKSZ, 0);
+
 			if (visibilitymap_clear(reln, oldblk, vmbuffer_old,
 									VISIBILITYMAP_VALID_BITS))
 				PageSetLSN(BufferGetPage(vmbuffer_old), lsn);

Attachments:

  [text/plain] initialize-zeroed-vm-pages-on-v2.txt (1.5K, ../../P2s-NV0--F-9@rhyadav.dev/2-initialize-zeroed-vm-pages-on-v2.txt)
  download | inline diff:
diff --git a/src/backend/access/heap/heapam_xlog.c b/src/backend/access/heap/heapam_xlog.c
index 4d99d99080..b87f38befc 100644
--- a/src/backend/access/heap/heapam_xlog.c
+++ b/src/backend/access/heap/heapam_xlog.c
@@ -89,6 +89,10 @@ heap_xlog_vm_clear(XLogReaderState *record,
 									  RBM_ZERO_ON_ERROR, false,
 									  &vmbuffer) == BLK_NEEDS_REDO)
 	{
+		/* initialize the page if it was read as zeros */
+		if (PageIsNew(BufferGetPage(vmbuffer)))
+			PageInit(BufferGetPage(vmbuffer), BLCKSZ, 0);
+
 		if (visibilitymap_clear(reln, heap_blkno, vmbuffer, flags))
 			PageSetLSN(BufferGetPage(vmbuffer), lsn);
 	}
@@ -855,6 +859,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 										  RBM_ZERO_ON_ERROR, false,
 										  &vmbuffer_new) == BLK_NEEDS_REDO)
 		{
+			/* initialize the page if it was read as zeros */
+			if (PageIsNew(BufferGetPage(vmbuffer_new)))
+				PageInit(BufferGetPage(vmbuffer_new), BLCKSZ, 0);
+
 			/*
 			 * If both the old and new heap pages were all-visible and their
 			 * VM bits are on the same VM page, that single VM page is
@@ -894,6 +902,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 										  RBM_ZERO_ON_ERROR, false,
 										  &vmbuffer_old) == BLK_NEEDS_REDO)
 		{
+			/* initialize the page if it was read as zeros */
+			if (PageIsNew(BufferGetPage(vmbuffer_old)))
+				PageInit(BufferGetPage(vmbuffer_old), BLCKSZ, 0);
+
 			if (visibilitymap_clear(reln, oldblk, vmbuffer_old,
 									VISIBILITYMAP_VALID_BITS))
 				PageSetLSN(BufferGetPage(vmbuffer_old), lsn);

^ permalink  raw  reply  [nested|flat] 7+ messages in thread

* Re: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation
@ 2026-10-02 05:15  shihao zhong <zhong950419@gmail.com>
  parent: rahul@rhyadav.dev
  1 sibling, 0 replies; 7+ messages in thread

From: shihao zhong @ 2026-10-02 05:15 UTC (permalink / raw)
  To: rahul@rhyadav.dev; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Nktpro <nktpro@gmail.com>; Pgsql Bugs <pgsql-bugs@lists.postgresql.org>; Melanie Plageman <melanieplageman@gmail.com>

Hi Rahul,

On Thu, Oct 1, 2026, rahul@rhyadav.dev wrote:
> However, with wal_consistency_checking = all on the primary, the
> patched standby still fails to restart:

I can reproduce this on master, and your change fixes it. It matches
heap_xlog_prune_freeze(), and vm_readbuf() on the primary.

No posted version of the fix applies to master on its own, so here
is one as v3. I will add a CommitFest entry so this does not get lost
before the November releases.

0001 is Melanie's v2-0002 from [1]. The changed lines and the commit
message are hers, only the context is adjusted for master. I left out
her v2-0001 and v2-0003, this report does not need them.

0002 is your diff. It had no commit message, so I put one together
from your mail. Please correct it if it is wrong.

0003 is my test. It now sets wal_consistency_checking = all, and it
takes a row lock first. Without the lock, the VM_NEW read in
heap_xlog_update() is not covered once 0002 is in.

The test fails without 0001. It also fails if any one of the three
RBM_ZERO_ON_ERROR or the three PageInit() calls is removed. That
holds on master, 19, 18 and 17. The nocfbot files hold the same three
patches for the back branches.

Melanie, if you would rather handle this in your thread, tell me and
I will close the entry.

[1]
https://postgr.es/m/CAAKRu_bApoksLDb-HX0GYciU3uLWqA1JagntaV8GP0%3D%2BidehHw%40mail.gmail.com

Thanks,
Shihao

Attachments:

  [application/octet-stream] nocfbot-v3-REL_18_STABLE.patch (10.3K, ../../CAGRkXqSZXoazu1x48pz2SuW_xHYQDdRcbp2HNCfRNm_WLF8aYA@mail.gmail.com/3-nocfbot-v3-REL_18_STABLE.patch)
  download | inline diff:
From 9d4a603d5d5017b37a508e0f9c0f022052b32c88 Mon Sep 17 00:00:00 2001
From: Melanie Plageman <melanieplageman@gmail.com>
Date: Wed, 23 Sep 2026 11:59:53 -0400
Subject: [PATCH v3 1/3] Read visibility map pages with RBM_ZERO_ON_ERROR in VM
 clear redo

ed62d26caca started registering VM blocks when clearing the VM which is
required for protection against torn pages as well as for correct
incremental backups. However, it read the VM pages in recovery with
RBM_NORMAL which errors out when it encounters a corrupt page. This is
usually desirable, however, we still retain code paths that modify the
VM in recovery without the block having been registered. A crash while
modifying the VM page could lead to a corrupt page and no FPI to recover
it. As long as we can trivially produce corrupt pages during recovery
through our own redo mechanism, we shouldn't error out when reading a
corrupt VM page.

Make clearing the VM read the page with RBM_ZERO_ON_ERROR. This is
consistent with the VM's other redo paths which set the VM bit
(heap_xlog_prune_freeze() and heap_xlog_multi_insert()).

Backpatch-through: 17
---

Notes:
    This is v2-0002 from
    https://postgr.es/m/CAAKRu_bApoksLDb-HX0GYciU3uLWqA1JagntaV8GP0%3D%2BidehHw%40mail.gmail.com
    The commit message is the same.  The code is adjusted to apply to
    REL_18_STABLE without v2-0001.
    
    It fixes the standby PANIC that Jacky Nguyen reported in
    https://postgr.es/m/CAL4mQLAp562c1rCgg2Dqx6TBSdOk6vOvaLFjwFG9E6k4uwvJJw@mail.gmail.com

 src/backend/access/heap/heapam_xlog.c | 15 +++++++++------
 1 file changed, 9 insertions(+), 6 deletions(-)

diff --git a/src/backend/access/heap/heapam_xlog.c b/src/backend/access/heap/heapam_xlog.c
index 0ce23dedfe4..3551cbe1683 100644
--- a/src/backend/access/heap/heapam_xlog.c
+++ b/src/backend/access/heap/heapam_xlog.c
@@ -57,8 +57,9 @@ heap_xlog_vm_clear(XLogReaderState *record,
 	 */
 	if (XLogRecHasBlockRef(record, wal_vm_block_id))
 	{
-		if (XLogReadBufferForRedo(record, wal_vm_block_id,
-								  &vmbuffer) == BLK_NEEDS_REDO)
+		if (XLogReadBufferForRedoExtended(record, wal_vm_block_id,
+										  RBM_ZERO_ON_ERROR, false,
+										  &vmbuffer) == BLK_NEEDS_REDO)
 		{
 			if (visibilitymap_clear_locked(reln,
 										   heap_blkno, vmbuffer,
@@ -803,8 +804,9 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 
 			Assert(xlrec->flags & XLH_UPDATE_NEW_ALL_VISIBLE_CLEARED);
 
-			if (XLogReadBufferForRedo(record, HEAP_UPDATE_BLKREF_VM_NEW,
-									  &vmbuffer_new) == BLK_NEEDS_REDO)
+			if (XLogReadBufferForRedoExtended(record, HEAP_UPDATE_BLKREF_VM_NEW,
+											  RBM_ZERO_ON_ERROR, false,
+											  &vmbuffer_new) == BLK_NEEDS_REDO)
 			{
 				/*
 				 * If both the old and new heap pages were all-visible and
@@ -841,8 +843,9 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 
 			Assert(xlrec->flags & XLH_UPDATE_OLD_ALL_VISIBLE_CLEARED);
 
-			if (XLogReadBufferForRedo(record, HEAP_UPDATE_BLKREF_VM_OLD, &vmbuffer_old) ==
-				BLK_NEEDS_REDO)
+			if (XLogReadBufferForRedoExtended(record, HEAP_UPDATE_BLKREF_VM_OLD,
+											  RBM_ZERO_ON_ERROR, false,
+											  &vmbuffer_old) == BLK_NEEDS_REDO)
 			{
 				if (visibilitymap_clear_locked(reln, oldblk, vmbuffer_old,
 											   VISIBILITYMAP_VALID_BITS))
-- 
2.37.1 (Apple Git-137.1)


From 98b07d18fcaf7ff2d3702fa9d1c3bb734e85a264 Mon Sep 17 00:00:00 2001
From: Rahul Yadav <rahul@rhyadav.dev>
Date: Thu, 1 Oct 2026 15:07:40 +0000
Subject: [PATCH v3 2/3] Initialize zeroed VM pages in VM clear redo

RBM_ZERO_ON_ERROR recreates a truncated VM page as all zeros.  With
wal_consistency_checking, verifyBackupPageConsistency() then masks it
with heap_mask(), and mask_unused_space() rejects a page with pd_lower
0.  heap_xlog_prune_freeze() and heap_xlog_multi_insert() already
initialize a VM page that was read as zeros.  Do the same at the three
VM clear sites.

Discussion: https://postgr.es/m/P2s-NV0--F-9@rhyadav.dev
---

Notes:
    Rahul posted this as a diff on top of v2, without a commit message.
    The message above is put together from his mail.

 src/backend/access/heap/heapam_xlog.c | 12 ++++++++++++
 1 file changed, 12 insertions(+)

diff --git a/src/backend/access/heap/heapam_xlog.c b/src/backend/access/heap/heapam_xlog.c
index 3551cbe1683..20b01076eab 100644
--- a/src/backend/access/heap/heapam_xlog.c
+++ b/src/backend/access/heap/heapam_xlog.c
@@ -61,6 +61,10 @@ heap_xlog_vm_clear(XLogReaderState *record,
 										  RBM_ZERO_ON_ERROR, false,
 										  &vmbuffer) == BLK_NEEDS_REDO)
 		{
+			/* initialize the page if it was read as zeros */
+			if (PageIsNew(BufferGetPage(vmbuffer)))
+				PageInit(BufferGetPage(vmbuffer), BLCKSZ, 0);
+
 			if (visibilitymap_clear_locked(reln,
 										   heap_blkno, vmbuffer,
 										   flags))
@@ -808,6 +812,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 											  RBM_ZERO_ON_ERROR, false,
 											  &vmbuffer_new) == BLK_NEEDS_REDO)
 			{
+				/* initialize the page if it was read as zeros */
+				if (PageIsNew(BufferGetPage(vmbuffer_new)))
+					PageInit(BufferGetPage(vmbuffer_new), BLCKSZ, 0);
+
 				/*
 				 * If both the old and new heap pages were all-visible and
 				 * their VM bits are on the same VM page, that single VM page
@@ -847,6 +855,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 											  RBM_ZERO_ON_ERROR, false,
 											  &vmbuffer_old) == BLK_NEEDS_REDO)
 			{
+				/* initialize the page if it was read as zeros */
+				if (PageIsNew(BufferGetPage(vmbuffer_old)))
+					PageInit(BufferGetPage(vmbuffer_old), BLCKSZ, 0);
+
 				if (visibilitymap_clear_locked(reln, oldblk, vmbuffer_old,
 											   VISIBILITYMAP_VALID_BITS))
 					PageSetLSN(BufferGetPage(vmbuffer_old), lsn);
-- 
2.37.1 (Apple Git-137.1)


From d6b97e09be8ef6fe1b3c15fdbd4400179f7dc0a2 Mon Sep 17 00:00:00 2001
From: Shihao <zhong950419@gmail.com>
Date: Thu, 1 Oct 2026 22:41:19 -0600
Subject: [PATCH v3 3/3] Add test for standby restart after VM truncation

With full_page_writes off, redo of a record that clears VM bits does
not restore the VM page from an image.  The test clears VM bits by
delete, same-page update and cross-page update, truncates the tables,
and restarts the standby from a restartpoint taken before those
changes.  wal_consistency_checking is on, so the VM page that redo
recreates is checked too.

Author: Shihao Zhong <zhong950419@gmail.com>
Discussion: https://postgr.es/m/CAL4mQLAp562c1rCgg2Dqx6TBSdOk6vOvaLFjwFG9E6k4uwvJJw@mail.gmail.com
---
 src/test/recovery/meson.build                |  1 +
 src/test/recovery/t/058_vm_clear_truncate.pl | 88 ++++++++++++++++++++
 2 files changed, 89 insertions(+)
 create mode 100644 src/test/recovery/t/058_vm_clear_truncate.pl

diff --git a/src/test/recovery/meson.build b/src/test/recovery/meson.build
index 3f03a706a2f..efebb761cf9 100644
--- a/src/test/recovery/meson.build
+++ b/src/test/recovery/meson.build
@@ -62,6 +62,7 @@ tests += {
       't/055_cascade_reconnect.pl',
       't/056_standby_snapshot_export.pl',
       't/057_snapshot_commit_race.pl',
+      't/058_vm_clear_truncate.pl',
     ],
   },
 }
diff --git a/src/test/recovery/t/058_vm_clear_truncate.pl b/src/test/recovery/t/058_vm_clear_truncate.pl
new file mode 100644
index 00000000000..294eb7de95e
--- /dev/null
+++ b/src/test/recovery/t/058_vm_clear_truncate.pl
@@ -0,0 +1,88 @@
+
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A standby must be able to restart when WAL it replays again clears
+# visibility map bits on a VM page that a later, already replayed,
+# truncation removed.  With full_page_writes off, redo cannot restore the
+# VM page from an image in the clearing record, so it has to cope with the
+# page not existing.  wal_consistency_checking is on so that the page redo
+# recreates is checked as well.
+use strict;
+use warnings FATAL => 'all';
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+	'postgresql.conf', qq{
+full_page_writes = off
+wal_consistency_checking = all
+autovacuum = off
+});
+$primary->start;
+$primary->backup('bkp');
+
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'bkp', has_streaming => 1);
+$standby->start;
+
+# Make every heap page all-visible, then make the standby create a
+# restartpoint, so that a restart replays the changes below again.  The
+# row lock clears the all-frozen bit of vm_upd's first page.  Otherwise
+# the cross-page update below would first log a lock record that clears
+# that bit, and the update would not be the first record to read the VM
+# page.
+$primary->safe_psql(
+	'postgres', q{
+CREATE TABLE vm_del (a int);
+CREATE TABLE vm_hot (a int) WITH (fillfactor = 50);
+CREATE TABLE vm_upd (a int);
+INSERT INTO vm_del SELECT generate_series(1, 1000);
+INSERT INTO vm_hot SELECT generate_series(1, 1000);
+INSERT INTO vm_upd SELECT generate_series(1, 1000);
+VACUUM (FREEZE) vm_del, vm_hot, vm_upd;
+SELECT a FROM vm_upd WHERE a = 1 FOR UPDATE;
+CHECKPOINT;
+});
+$primary->wait_for_replay_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT');
+
+my $start_lsn =
+  $primary->safe_psql('postgres', 'SELECT pg_current_wal_insert_lsn()');
+
+# Clear VM bits through delete, same-page update (old VM block) and
+# cross-page update (new VM block), then truncate all three tables to
+# zero blocks.
+$primary->safe_psql(
+	'postgres', q{
+DELETE FROM vm_del;
+UPDATE vm_hot SET a = -a WHERE a = 1;
+UPDATE vm_upd SET a = -a WHERE a = 1;
+DELETE FROM vm_hot;
+DELETE FROM vm_upd;
+VACUUM vm_del, vm_hot, vm_upd;
+});
+is( $primary->safe_psql(
+		'postgres',
+		"SELECT sum(pg_relation_size(c, 'vm')) FROM unnest('{vm_del,vm_hot,vm_upd}'::regclass[]) c"
+	),
+	'0',
+	'VMs truncated on primary');
+$primary->wait_for_replay_catchup($standby);
+
+$standby->stop;
+my $log_offset = -s $standby->logfile;
+my $ret = $standby->start(fail_ok => 1);
+
+my $log = slurp_file($standby->logfile, $log_offset);
+my ($redo_lsn) = $log =~ /redo starts at ([0-9A-F]+\/[0-9A-F]+)/;
+ok( defined($redo_lsn)
+	  && $primary->safe_psql('postgres',
+		"SELECT '$redo_lsn'::pg_lsn < '$start_lsn'::pg_lsn") eq 't',
+	'redo after restart starts before the VM bits were cleared');
+ok($ret, 'standby restarts after replaying VM truncation');
+unlike($log, qr/invalid pages/, 'no invalid page references in standby log');
+
+done_testing();
-- 
2.37.1 (Apple Git-137.1)



  [application/octet-stream] nocfbot-v3-REL_19_STABLE.patch (10.3K, ../../CAGRkXqSZXoazu1x48pz2SuW_xHYQDdRcbp2HNCfRNm_WLF8aYA@mail.gmail.com/4-nocfbot-v3-REL_19_STABLE.patch)
  download | inline diff:
From 924d5bbb4698ceba38f6468f1bcdc0c069a6a6c0 Mon Sep 17 00:00:00 2001
From: Melanie Plageman <melanieplageman@gmail.com>
Date: Wed, 23 Sep 2026 11:59:53 -0400
Subject: [PATCH v3 1/3] Read visibility map pages with RBM_ZERO_ON_ERROR in VM
 clear redo

ed62d26caca started registering VM blocks when clearing the VM which is
required for protection against torn pages as well as for correct
incremental backups. However, it read the VM pages in recovery with
RBM_NORMAL which errors out when it encounters a corrupt page. This is
usually desirable, however, we still retain code paths that modify the
VM in recovery without the block having been registered. A crash while
modifying the VM page could lead to a corrupt page and no FPI to recover
it. As long as we can trivially produce corrupt pages during recovery
through our own redo mechanism, we shouldn't error out when reading a
corrupt VM page.

Make clearing the VM read the page with RBM_ZERO_ON_ERROR. This is
consistent with the VM's other redo paths which set the VM bit
(heap_xlog_prune_freeze() and heap_xlog_multi_insert()).

Backpatch-through: 17
---

Notes:
    This is v2-0002 from
    https://postgr.es/m/CAAKRu_bApoksLDb-HX0GYciU3uLWqA1JagntaV8GP0%3D%2BidehHw%40mail.gmail.com
    The commit message is the same.  The code is adjusted to apply to
    REL_19_STABLE without v2-0001.
    
    It fixes the standby PANIC that Jacky Nguyen reported in
    https://postgr.es/m/CAL4mQLAp562c1rCgg2Dqx6TBSdOk6vOvaLFjwFG9E6k4uwvJJw@mail.gmail.com

 src/backend/access/heap/heapam_xlog.c | 15 +++++++++------
 1 file changed, 9 insertions(+), 6 deletions(-)

diff --git a/src/backend/access/heap/heapam_xlog.c b/src/backend/access/heap/heapam_xlog.c
index fae3b477c09..288b4952579 100644
--- a/src/backend/access/heap/heapam_xlog.c
+++ b/src/backend/access/heap/heapam_xlog.c
@@ -53,8 +53,9 @@ heap_xlog_vm_clear(XLogReaderState *record,
 	 */
 	if (XLogRecHasBlockRef(record, wal_vm_block_id))
 	{
-		if (XLogReadBufferForRedo(record, wal_vm_block_id,
-								  &vmbuffer) == BLK_NEEDS_REDO)
+		if (XLogReadBufferForRedoExtended(record, wal_vm_block_id,
+										  RBM_ZERO_ON_ERROR, false,
+										  &vmbuffer) == BLK_NEEDS_REDO)
 		{
 			if (visibilitymap_clear(reln, heap_blkno, vmbuffer, flags))
 				PageSetLSN(BufferGetPage(vmbuffer), lsn);
@@ -819,8 +820,9 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 
 		Assert(xlrec->flags & XLH_UPDATE_NEW_ALL_VISIBLE_CLEARED);
 
-		if (XLogReadBufferForRedo(record, HEAP_UPDATE_BLKREF_VM_NEW,
-								  &vmbuffer_new) == BLK_NEEDS_REDO)
+		if (XLogReadBufferForRedoExtended(record, HEAP_UPDATE_BLKREF_VM_NEW,
+										  RBM_ZERO_ON_ERROR, false,
+										  &vmbuffer_new) == BLK_NEEDS_REDO)
 		{
 			/*
 			 * If both the old and new heap pages were all-visible and their
@@ -857,8 +859,9 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 
 		Assert(xlrec->flags & XLH_UPDATE_OLD_ALL_VISIBLE_CLEARED);
 
-		if (XLogReadBufferForRedo(record, HEAP_UPDATE_BLKREF_VM_OLD,
-								  &vmbuffer_old) == BLK_NEEDS_REDO)
+		if (XLogReadBufferForRedoExtended(record, HEAP_UPDATE_BLKREF_VM_OLD,
+										  RBM_ZERO_ON_ERROR, false,
+										  &vmbuffer_old) == BLK_NEEDS_REDO)
 		{
 			if (visibilitymap_clear(reln, oldblk, vmbuffer_old,
 									VISIBILITYMAP_VALID_BITS))
-- 
2.37.1 (Apple Git-137.1)


From 0b36adcefe72b5c7197f95559df5364e51827388 Mon Sep 17 00:00:00 2001
From: Rahul Yadav <rahul@rhyadav.dev>
Date: Thu, 1 Oct 2026 15:07:40 +0000
Subject: [PATCH v3 2/3] Initialize zeroed VM pages in VM clear redo

RBM_ZERO_ON_ERROR recreates a truncated VM page as all zeros.  With
wal_consistency_checking, verifyBackupPageConsistency() then masks it
with heap_mask(), and mask_unused_space() rejects a page with pd_lower
0.  heap_xlog_prune_freeze() and heap_xlog_multi_insert() already
initialize a VM page that was read as zeros.  Do the same at the three
VM clear sites.

Discussion: https://postgr.es/m/P2s-NV0--F-9@rhyadav.dev
---

Notes:
    Rahul posted this as a diff on top of v2, without a commit message.
    The message above is put together from his mail.

 src/backend/access/heap/heapam_xlog.c | 12 ++++++++++++
 1 file changed, 12 insertions(+)

diff --git a/src/backend/access/heap/heapam_xlog.c b/src/backend/access/heap/heapam_xlog.c
index 288b4952579..1ce0c09a083 100644
--- a/src/backend/access/heap/heapam_xlog.c
+++ b/src/backend/access/heap/heapam_xlog.c
@@ -57,6 +57,10 @@ heap_xlog_vm_clear(XLogReaderState *record,
 										  RBM_ZERO_ON_ERROR, false,
 										  &vmbuffer) == BLK_NEEDS_REDO)
 		{
+			/* initialize the page if it was read as zeros */
+			if (PageIsNew(BufferGetPage(vmbuffer)))
+				PageInit(BufferGetPage(vmbuffer), BLCKSZ, 0);
+
 			if (visibilitymap_clear(reln, heap_blkno, vmbuffer, flags))
 				PageSetLSN(BufferGetPage(vmbuffer), lsn);
 		}
@@ -824,6 +828,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 										  RBM_ZERO_ON_ERROR, false,
 										  &vmbuffer_new) == BLK_NEEDS_REDO)
 		{
+			/* initialize the page if it was read as zeros */
+			if (PageIsNew(BufferGetPage(vmbuffer_new)))
+				PageInit(BufferGetPage(vmbuffer_new), BLCKSZ, 0);
+
 			/*
 			 * If both the old and new heap pages were all-visible and their
 			 * VM bits are on the same VM page, that single VM page is
@@ -863,6 +871,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 										  RBM_ZERO_ON_ERROR, false,
 										  &vmbuffer_old) == BLK_NEEDS_REDO)
 		{
+			/* initialize the page if it was read as zeros */
+			if (PageIsNew(BufferGetPage(vmbuffer_old)))
+				PageInit(BufferGetPage(vmbuffer_old), BLCKSZ, 0);
+
 			if (visibilitymap_clear(reln, oldblk, vmbuffer_old,
 									VISIBILITYMAP_VALID_BITS))
 				PageSetLSN(BufferGetPage(vmbuffer_old), lsn);
-- 
2.37.1 (Apple Git-137.1)


From 364c4c54a128f698fff8c2b9b7ae26b55378dcf4 Mon Sep 17 00:00:00 2001
From: Shihao <zhong950419@gmail.com>
Date: Thu, 1 Oct 2026 22:38:25 -0600
Subject: [PATCH v3 3/3] Add test for standby restart after VM truncation

With full_page_writes off, redo of a record that clears VM bits does
not restore the VM page from an image.  The test clears VM bits by
delete, same-page update and cross-page update, truncates the tables,
and restarts the standby from a restartpoint taken before those
changes.  wal_consistency_checking is on, so the VM page that redo
recreates is checked too.

Author: Shihao Zhong <zhong950419@gmail.com>
Discussion: https://postgr.es/m/CAL4mQLAp562c1rCgg2Dqx6TBSdOk6vOvaLFjwFG9E6k4uwvJJw@mail.gmail.com
---
 src/test/recovery/meson.build                |  1 +
 src/test/recovery/t/058_vm_clear_truncate.pl | 88 ++++++++++++++++++++
 2 files changed, 89 insertions(+)
 create mode 100644 src/test/recovery/t/058_vm_clear_truncate.pl

diff --git a/src/test/recovery/meson.build b/src/test/recovery/meson.build
index ebb12dd8766..bb28f9cfb81 100644
--- a/src/test/recovery/meson.build
+++ b/src/test/recovery/meson.build
@@ -66,6 +66,7 @@ tests += {
       't/055_cascade_reconnect.pl',
       't/056_standby_snapshot_export.pl',
       't/057_snapshot_commit_race.pl',
+      't/058_vm_clear_truncate.pl',
     ],
   },
 }
diff --git a/src/test/recovery/t/058_vm_clear_truncate.pl b/src/test/recovery/t/058_vm_clear_truncate.pl
new file mode 100644
index 00000000000..294eb7de95e
--- /dev/null
+++ b/src/test/recovery/t/058_vm_clear_truncate.pl
@@ -0,0 +1,88 @@
+
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A standby must be able to restart when WAL it replays again clears
+# visibility map bits on a VM page that a later, already replayed,
+# truncation removed.  With full_page_writes off, redo cannot restore the
+# VM page from an image in the clearing record, so it has to cope with the
+# page not existing.  wal_consistency_checking is on so that the page redo
+# recreates is checked as well.
+use strict;
+use warnings FATAL => 'all';
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+	'postgresql.conf', qq{
+full_page_writes = off
+wal_consistency_checking = all
+autovacuum = off
+});
+$primary->start;
+$primary->backup('bkp');
+
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'bkp', has_streaming => 1);
+$standby->start;
+
+# Make every heap page all-visible, then make the standby create a
+# restartpoint, so that a restart replays the changes below again.  The
+# row lock clears the all-frozen bit of vm_upd's first page.  Otherwise
+# the cross-page update below would first log a lock record that clears
+# that bit, and the update would not be the first record to read the VM
+# page.
+$primary->safe_psql(
+	'postgres', q{
+CREATE TABLE vm_del (a int);
+CREATE TABLE vm_hot (a int) WITH (fillfactor = 50);
+CREATE TABLE vm_upd (a int);
+INSERT INTO vm_del SELECT generate_series(1, 1000);
+INSERT INTO vm_hot SELECT generate_series(1, 1000);
+INSERT INTO vm_upd SELECT generate_series(1, 1000);
+VACUUM (FREEZE) vm_del, vm_hot, vm_upd;
+SELECT a FROM vm_upd WHERE a = 1 FOR UPDATE;
+CHECKPOINT;
+});
+$primary->wait_for_replay_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT');
+
+my $start_lsn =
+  $primary->safe_psql('postgres', 'SELECT pg_current_wal_insert_lsn()');
+
+# Clear VM bits through delete, same-page update (old VM block) and
+# cross-page update (new VM block), then truncate all three tables to
+# zero blocks.
+$primary->safe_psql(
+	'postgres', q{
+DELETE FROM vm_del;
+UPDATE vm_hot SET a = -a WHERE a = 1;
+UPDATE vm_upd SET a = -a WHERE a = 1;
+DELETE FROM vm_hot;
+DELETE FROM vm_upd;
+VACUUM vm_del, vm_hot, vm_upd;
+});
+is( $primary->safe_psql(
+		'postgres',
+		"SELECT sum(pg_relation_size(c, 'vm')) FROM unnest('{vm_del,vm_hot,vm_upd}'::regclass[]) c"
+	),
+	'0',
+	'VMs truncated on primary');
+$primary->wait_for_replay_catchup($standby);
+
+$standby->stop;
+my $log_offset = -s $standby->logfile;
+my $ret = $standby->start(fail_ok => 1);
+
+my $log = slurp_file($standby->logfile, $log_offset);
+my ($redo_lsn) = $log =~ /redo starts at ([0-9A-F]+\/[0-9A-F]+)/;
+ok( defined($redo_lsn)
+	  && $primary->safe_psql('postgres',
+		"SELECT '$redo_lsn'::pg_lsn < '$start_lsn'::pg_lsn") eq 't',
+	'redo after restart starts before the VM bits were cleared');
+ok($ret, 'standby restarts after replaying VM truncation');
+unlike($log, qr/invalid pages/, 'no invalid page references in standby log');
+
+done_testing();
-- 
2.37.1 (Apple Git-137.1)



  [application/octet-stream] nocfbot-v3-REL_17_STABLE.patch (10.3K, ../../CAGRkXqSZXoazu1x48pz2SuW_xHYQDdRcbp2HNCfRNm_WLF8aYA@mail.gmail.com/5-nocfbot-v3-REL_17_STABLE.patch)
  download | inline diff:
From 1863d54b172435f5b804208626c888afc3af19ce Mon Sep 17 00:00:00 2001
From: Melanie Plageman <melanieplageman@gmail.com>
Date: Wed, 23 Sep 2026 11:59:53 -0400
Subject: [PATCH v3 1/3] Read visibility map pages with RBM_ZERO_ON_ERROR in VM
 clear redo

ed62d26caca started registering VM blocks when clearing the VM which is
required for protection against torn pages as well as for correct
incremental backups. However, it read the VM pages in recovery with
RBM_NORMAL which errors out when it encounters a corrupt page. This is
usually desirable, however, we still retain code paths that modify the
VM in recovery without the block having been registered. A crash while
modifying the VM page could lead to a corrupt page and no FPI to recover
it. As long as we can trivially produce corrupt pages during recovery
through our own redo mechanism, we shouldn't error out when reading a
corrupt VM page.

Make clearing the VM read the page with RBM_ZERO_ON_ERROR. This is
consistent with the VM's other redo paths which set the VM bit
(heap_xlog_prune_freeze() and heap_xlog_multi_insert()).

Backpatch-through: 17
---

Notes:
    This is v2-0002 from
    https://postgr.es/m/CAAKRu_bApoksLDb-HX0GYciU3uLWqA1JagntaV8GP0%3D%2BidehHw%40mail.gmail.com
    The commit message is the same.  The code is adjusted to apply to
    REL_17_STABLE without v2-0001.
    
    It fixes the standby PANIC that Jacky Nguyen reported in
    https://postgr.es/m/CAL4mQLAp562c1rCgg2Dqx6TBSdOk6vOvaLFjwFG9E6k4uwvJJw@mail.gmail.com

 src/backend/access/heap/heapam.c | 15 +++++++++------
 1 file changed, 9 insertions(+), 6 deletions(-)

diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index 5b9f3a2670e..4d061a4bc2e 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -303,8 +303,9 @@ heap_xlog_vm_clear(XLogReaderState *record,
 	 */
 	if (XLogRecHasBlockRef(record, wal_vm_block_id))
 	{
-		if (XLogReadBufferForRedo(record, wal_vm_block_id,
-								  &vmbuffer) == BLK_NEEDS_REDO)
+		if (XLogReadBufferForRedoExtended(record, wal_vm_block_id,
+										  RBM_ZERO_ON_ERROR, false,
+										  &vmbuffer) == BLK_NEEDS_REDO)
 		{
 			if (visibilitymap_clear_locked(reln,
 										   heap_blkno, vmbuffer,
@@ -10337,8 +10338,9 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 
 			Assert(xlrec->flags & XLH_UPDATE_NEW_ALL_VISIBLE_CLEARED);
 
-			if (XLogReadBufferForRedo(record, HEAP_UPDATE_BLKREF_VM_NEW,
-									  &vmbuffer_new) == BLK_NEEDS_REDO)
+			if (XLogReadBufferForRedoExtended(record, HEAP_UPDATE_BLKREF_VM_NEW,
+											  RBM_ZERO_ON_ERROR, false,
+											  &vmbuffer_new) == BLK_NEEDS_REDO)
 			{
 				/*
 				 * If both the old and new heap pages were all-visible and
@@ -10375,8 +10377,9 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 
 			Assert(xlrec->flags & XLH_UPDATE_OLD_ALL_VISIBLE_CLEARED);
 
-			if (XLogReadBufferForRedo(record, HEAP_UPDATE_BLKREF_VM_OLD, &vmbuffer_old) ==
-				BLK_NEEDS_REDO)
+			if (XLogReadBufferForRedoExtended(record, HEAP_UPDATE_BLKREF_VM_OLD,
+											  RBM_ZERO_ON_ERROR, false,
+											  &vmbuffer_old) == BLK_NEEDS_REDO)
 			{
 				if (visibilitymap_clear_locked(reln, oldblk, vmbuffer_old,
 											   VISIBILITYMAP_VALID_BITS))
-- 
2.37.1 (Apple Git-137.1)


From ddb43c81eb55d4a6b2d2a304e4071316cfc7b51b Mon Sep 17 00:00:00 2001
From: Rahul Yadav <rahul@rhyadav.dev>
Date: Thu, 1 Oct 2026 15:07:40 +0000
Subject: [PATCH v3 2/3] Initialize zeroed VM pages in VM clear redo

RBM_ZERO_ON_ERROR recreates a truncated VM page as all zeros.  With
wal_consistency_checking, verifyBackupPageConsistency() then masks it
with heap_mask(), and mask_unused_space() rejects a page with pd_lower
0.  heap_xlog_prune_freeze() and heap_xlog_multi_insert() already
initialize a VM page that was read as zeros.  Do the same at the three
VM clear sites.

Discussion: https://postgr.es/m/P2s-NV0--F-9@rhyadav.dev
---

Notes:
    Rahul posted this as a diff on top of v2, without a commit message.
    The message above is put together from his mail.

 src/backend/access/heap/heapam.c | 12 ++++++++++++
 1 file changed, 12 insertions(+)

diff --git a/src/backend/access/heap/heapam.c b/src/backend/access/heap/heapam.c
index 4d061a4bc2e..27560bb2862 100644
--- a/src/backend/access/heap/heapam.c
+++ b/src/backend/access/heap/heapam.c
@@ -307,6 +307,10 @@ heap_xlog_vm_clear(XLogReaderState *record,
 										  RBM_ZERO_ON_ERROR, false,
 										  &vmbuffer) == BLK_NEEDS_REDO)
 		{
+			/* initialize the page if it was read as zeros */
+			if (PageIsNew(BufferGetPage(vmbuffer)))
+				PageInit(BufferGetPage(vmbuffer), BLCKSZ, 0);
+
 			if (visibilitymap_clear_locked(reln,
 										   heap_blkno, vmbuffer,
 										   flags))
@@ -10342,6 +10346,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 											  RBM_ZERO_ON_ERROR, false,
 											  &vmbuffer_new) == BLK_NEEDS_REDO)
 			{
+				/* initialize the page if it was read as zeros */
+				if (PageIsNew(BufferGetPage(vmbuffer_new)))
+					PageInit(BufferGetPage(vmbuffer_new), BLCKSZ, 0);
+
 				/*
 				 * If both the old and new heap pages were all-visible and
 				 * their VM bits are on the same VM page, that single VM page
@@ -10381,6 +10389,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 											  RBM_ZERO_ON_ERROR, false,
 											  &vmbuffer_old) == BLK_NEEDS_REDO)
 			{
+				/* initialize the page if it was read as zeros */
+				if (PageIsNew(BufferGetPage(vmbuffer_old)))
+					PageInit(BufferGetPage(vmbuffer_old), BLCKSZ, 0);
+
 				if (visibilitymap_clear_locked(reln, oldblk, vmbuffer_old,
 											   VISIBILITYMAP_VALID_BITS))
 					PageSetLSN(BufferGetPage(vmbuffer_old), lsn);
-- 
2.37.1 (Apple Git-137.1)


From 79ed7be58cd3fb1bd7f933e1c8a709b391cec73b Mon Sep 17 00:00:00 2001
From: Shihao <zhong950419@gmail.com>
Date: Thu, 1 Oct 2026 22:43:56 -0600
Subject: [PATCH v3 3/3] Add test for standby restart after VM truncation

With full_page_writes off, redo of a record that clears VM bits does
not restore the VM page from an image.  The test clears VM bits by
delete, same-page update and cross-page update, truncates the tables,
and restarts the standby from a restartpoint taken before those
changes.  wal_consistency_checking is on, so the VM page that redo
recreates is checked too.

Author: Shihao Zhong <zhong950419@gmail.com>
Discussion: https://postgr.es/m/CAL4mQLAp562c1rCgg2Dqx6TBSdOk6vOvaLFjwFG9E6k4uwvJJw@mail.gmail.com
---
 src/test/recovery/meson.build                |  1 +
 src/test/recovery/t/058_vm_clear_truncate.pl | 88 ++++++++++++++++++++
 2 files changed, 89 insertions(+)
 create mode 100644 src/test/recovery/t/058_vm_clear_truncate.pl

diff --git a/src/test/recovery/meson.build b/src/test/recovery/meson.build
index eb716d39d5c..924545844c6 100644
--- a/src/test/recovery/meson.build
+++ b/src/test/recovery/meson.build
@@ -60,6 +60,7 @@ tests += {
       't/054_unlogged_sequence_promotion.pl',
       't/055_cascade_reconnect.pl',
       't/056_standby_snapshot_export.pl',
+      't/058_vm_clear_truncate.pl',
     ],
   },
 }
diff --git a/src/test/recovery/t/058_vm_clear_truncate.pl b/src/test/recovery/t/058_vm_clear_truncate.pl
new file mode 100644
index 00000000000..294eb7de95e
--- /dev/null
+++ b/src/test/recovery/t/058_vm_clear_truncate.pl
@@ -0,0 +1,88 @@
+
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A standby must be able to restart when WAL it replays again clears
+# visibility map bits on a VM page that a later, already replayed,
+# truncation removed.  With full_page_writes off, redo cannot restore the
+# VM page from an image in the clearing record, so it has to cope with the
+# page not existing.  wal_consistency_checking is on so that the page redo
+# recreates is checked as well.
+use strict;
+use warnings FATAL => 'all';
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+	'postgresql.conf', qq{
+full_page_writes = off
+wal_consistency_checking = all
+autovacuum = off
+});
+$primary->start;
+$primary->backup('bkp');
+
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'bkp', has_streaming => 1);
+$standby->start;
+
+# Make every heap page all-visible, then make the standby create a
+# restartpoint, so that a restart replays the changes below again.  The
+# row lock clears the all-frozen bit of vm_upd's first page.  Otherwise
+# the cross-page update below would first log a lock record that clears
+# that bit, and the update would not be the first record to read the VM
+# page.
+$primary->safe_psql(
+	'postgres', q{
+CREATE TABLE vm_del (a int);
+CREATE TABLE vm_hot (a int) WITH (fillfactor = 50);
+CREATE TABLE vm_upd (a int);
+INSERT INTO vm_del SELECT generate_series(1, 1000);
+INSERT INTO vm_hot SELECT generate_series(1, 1000);
+INSERT INTO vm_upd SELECT generate_series(1, 1000);
+VACUUM (FREEZE) vm_del, vm_hot, vm_upd;
+SELECT a FROM vm_upd WHERE a = 1 FOR UPDATE;
+CHECKPOINT;
+});
+$primary->wait_for_replay_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT');
+
+my $start_lsn =
+  $primary->safe_psql('postgres', 'SELECT pg_current_wal_insert_lsn()');
+
+# Clear VM bits through delete, same-page update (old VM block) and
+# cross-page update (new VM block), then truncate all three tables to
+# zero blocks.
+$primary->safe_psql(
+	'postgres', q{
+DELETE FROM vm_del;
+UPDATE vm_hot SET a = -a WHERE a = 1;
+UPDATE vm_upd SET a = -a WHERE a = 1;
+DELETE FROM vm_hot;
+DELETE FROM vm_upd;
+VACUUM vm_del, vm_hot, vm_upd;
+});
+is( $primary->safe_psql(
+		'postgres',
+		"SELECT sum(pg_relation_size(c, 'vm')) FROM unnest('{vm_del,vm_hot,vm_upd}'::regclass[]) c"
+	),
+	'0',
+	'VMs truncated on primary');
+$primary->wait_for_replay_catchup($standby);
+
+$standby->stop;
+my $log_offset = -s $standby->logfile;
+my $ret = $standby->start(fail_ok => 1);
+
+my $log = slurp_file($standby->logfile, $log_offset);
+my ($redo_lsn) = $log =~ /redo starts at ([0-9A-F]+\/[0-9A-F]+)/;
+ok( defined($redo_lsn)
+	  && $primary->safe_psql('postgres',
+		"SELECT '$redo_lsn'::pg_lsn < '$start_lsn'::pg_lsn") eq 't',
+	'redo after restart starts before the VM bits were cleared');
+ok($ret, 'standby restarts after replaying VM truncation');
+unlike($log, qr/invalid pages/, 'no invalid page references in standby log');
+
+done_testing();
-- 
2.37.1 (Apple Git-137.1)



  [application/octet-stream] v3-0002-Initialize-zeroed-VM-pages-in-VM-clear-redo.patch (2.4K, ../../CAGRkXqSZXoazu1x48pz2SuW_xHYQDdRcbp2HNCfRNm_WLF8aYA@mail.gmail.com/6-v3-0002-Initialize-zeroed-VM-pages-in-VM-clear-redo.patch)
  download | inline diff:
From 9e36031d684efaede04ebbfebed59a594808d2a4 Mon Sep 17 00:00:00 2001
From: Rahul Yadav <rahul@rhyadav.dev>
Date: Thu, 1 Oct 2026 15:07:40 +0000
Subject: [PATCH v3 2/3] Initialize zeroed VM pages in VM clear redo

RBM_ZERO_ON_ERROR recreates a truncated VM page as all zeros.  With
wal_consistency_checking, verifyBackupPageConsistency() then masks it
with heap_mask(), and mask_unused_space() rejects a page with pd_lower
0.  heap_xlog_prune_freeze() and heap_xlog_multi_insert() already
initialize a VM page that was read as zeros.  Do the same at the three
VM clear sites.

Discussion: https://postgr.es/m/P2s-NV0--F-9@rhyadav.dev
---

Notes:
    Rahul posted this as a diff on top of v2, without a commit message.
    The message above is put together from his mail.

 src/backend/access/heap/heapam_xlog.c | 12 ++++++++++++
 1 file changed, 12 insertions(+)

diff --git a/src/backend/access/heap/heapam_xlog.c b/src/backend/access/heap/heapam_xlog.c
index b82f388b4be..6ffd2fa6180 100644
--- a/src/backend/access/heap/heapam_xlog.c
+++ b/src/backend/access/heap/heapam_xlog.c
@@ -57,6 +57,10 @@ heap_xlog_vm_clear(XLogReaderState *record,
 									  RBM_ZERO_ON_ERROR, false,
 									  &vmbuffer) == BLK_NEEDS_REDO)
 	{
+		/* initialize the page if it was read as zeros */
+		if (PageIsNew(BufferGetPage(vmbuffer)))
+			PageInit(BufferGetPage(vmbuffer), BLCKSZ, 0);
+
 		if (visibilitymap_clear(target_locator, heap_blkno, vmbuffer, flags))
 			PageSetLSN(BufferGetPage(vmbuffer), lsn);
 	}
@@ -817,6 +821,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 										  RBM_ZERO_ON_ERROR, false,
 										  &vmbuffer_new) == BLK_NEEDS_REDO)
 		{
+			/* initialize the page if it was read as zeros */
+			if (PageIsNew(BufferGetPage(vmbuffer_new)))
+				PageInit(BufferGetPage(vmbuffer_new), BLCKSZ, 0);
+
 			/*
 			 * If both the old and new heap pages were all-visible and their
 			 * VM bits are on the same VM page, that single VM page is
@@ -856,6 +864,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 										  RBM_ZERO_ON_ERROR, false,
 										  &vmbuffer_old) == BLK_NEEDS_REDO)
 		{
+			/* initialize the page if it was read as zeros */
+			if (PageIsNew(BufferGetPage(vmbuffer_old)))
+				PageInit(BufferGetPage(vmbuffer_old), BLCKSZ, 0);
+
 			if (visibilitymap_clear(rlocator, oldblk, vmbuffer_old,
 									VISIBILITYMAP_VALID_BITS))
 				PageSetLSN(BufferGetPage(vmbuffer_old), lsn);
-- 
2.37.1 (Apple Git-137.1)



  [application/octet-stream] v3-0001-Read-visibility-map-pages-with-RBM_ZERO_ON_ERROR-.patch (3.4K, ../../CAGRkXqSZXoazu1x48pz2SuW_xHYQDdRcbp2HNCfRNm_WLF8aYA@mail.gmail.com/7-v3-0001-Read-visibility-map-pages-with-RBM_ZERO_ON_ERROR-.patch)
  download | inline diff:
From 45824b79b53ee3257d037d3ac9801d8eda3c35d6 Mon Sep 17 00:00:00 2001
From: Melanie Plageman <melanieplageman@gmail.com>
Date: Wed, 23 Sep 2026 11:59:53 -0400
Subject: [PATCH v3 1/3] Read visibility map pages with RBM_ZERO_ON_ERROR in VM
 clear redo

ed62d26caca started registering VM blocks when clearing the VM which is
required for protection against torn pages as well as for correct
incremental backups. However, it read the VM pages in recovery with
RBM_NORMAL which errors out when it encounters a corrupt page. This is
usually desirable, however, we still retain code paths that modify the
VM in recovery without the block having been registered. A crash while
modifying the VM page could lead to a corrupt page and no FPI to recover
it. As long as we can trivially produce corrupt pages during recovery
through our own redo mechanism, we shouldn't error out when reading a
corrupt VM page.

Make clearing the VM read the page with RBM_ZERO_ON_ERROR. This is
consistent with the VM's other redo paths which set the VM bit
(heap_xlog_prune_freeze() and heap_xlog_multi_insert()).

Backpatch-through: 17
---

Notes:
    This is v2-0002 from
    https://postgr.es/m/CAAKRu_bApoksLDb-HX0GYciU3uLWqA1JagntaV8GP0%3D%2BidehHw%40mail.gmail.com
    The changed lines and the commit message are the same.  Only the
    context differs, so that it applies to master without v2-0001.
    
    It fixes the standby PANIC that Jacky Nguyen reported in
    https://postgr.es/m/CAL4mQLAp562c1rCgg2Dqx6TBSdOk6vOvaLFjwFG9E6k4uwvJJw@mail.gmail.com

 src/backend/access/heap/heapam_xlog.c | 15 +++++++++------
 1 file changed, 9 insertions(+), 6 deletions(-)

diff --git a/src/backend/access/heap/heapam_xlog.c b/src/backend/access/heap/heapam_xlog.c
index 5fa1de09cfb..b82f388b4be 100644
--- a/src/backend/access/heap/heapam_xlog.c
+++ b/src/backend/access/heap/heapam_xlog.c
@@ -53,8 +53,9 @@ heap_xlog_vm_clear(XLogReaderState *record,
 	 * read it. These will either apply an FPI or indicate that we should
 	 * clear the requested bits ourselves.
 	 */
-	if (XLogReadBufferForRedo(record, wal_vm_block_id,
-							  &vmbuffer) == BLK_NEEDS_REDO)
+	if (XLogReadBufferForRedoExtended(record, wal_vm_block_id,
+									  RBM_ZERO_ON_ERROR, false,
+									  &vmbuffer) == BLK_NEEDS_REDO)
 	{
 		if (visibilitymap_clear(target_locator, heap_blkno, vmbuffer, flags))
 			PageSetLSN(BufferGetPage(vmbuffer), lsn);
@@ -812,8 +813,9 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 
 		Assert(xlrec->flags & XLH_UPDATE_NEW_ALL_VISIBLE_CLEARED);
 
-		if (XLogReadBufferForRedo(record, HEAP_UPDATE_BLKREF_VM_NEW,
-								  &vmbuffer_new) == BLK_NEEDS_REDO)
+		if (XLogReadBufferForRedoExtended(record, HEAP_UPDATE_BLKREF_VM_NEW,
+										  RBM_ZERO_ON_ERROR, false,
+										  &vmbuffer_new) == BLK_NEEDS_REDO)
 		{
 			/*
 			 * If both the old and new heap pages were all-visible and their
@@ -850,8 +852,9 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 
 		Assert(xlrec->flags & XLH_UPDATE_OLD_ALL_VISIBLE_CLEARED);
 
-		if (XLogReadBufferForRedo(record, HEAP_UPDATE_BLKREF_VM_OLD,
-								  &vmbuffer_old) == BLK_NEEDS_REDO)
+		if (XLogReadBufferForRedoExtended(record, HEAP_UPDATE_BLKREF_VM_OLD,
+										  RBM_ZERO_ON_ERROR, false,
+										  &vmbuffer_old) == BLK_NEEDS_REDO)
 		{
 			if (visibilitymap_clear(rlocator, oldblk, vmbuffer_old,
 									VISIBILITYMAP_VALID_BITS))
-- 
2.37.1 (Apple Git-137.1)



  [application/octet-stream] v3-0003-Add-test-for-standby-restart-after-VM-truncation.patch (4.7K, ../../CAGRkXqSZXoazu1x48pz2SuW_xHYQDdRcbp2HNCfRNm_WLF8aYA@mail.gmail.com/8-v3-0003-Add-test-for-standby-restart-after-VM-truncation.patch)
  download | inline diff:
From 406b48c564e67ea692e663df03e4429e6d322d25 Mon Sep 17 00:00:00 2001
From: Shihao <zhong950419@gmail.com>
Date: Thu, 1 Oct 2026 22:31:12 -0600
Subject: [PATCH v3 3/3] Add test for standby restart after VM truncation

With full_page_writes off, redo of a record that clears VM bits does
not restore the VM page from an image.  The test clears VM bits by
delete, same-page update and cross-page update, truncates the tables,
and restarts the standby from a restartpoint taken before those
changes.  wal_consistency_checking is on, so the VM page that redo
recreates is checked too.

Author: Shihao Zhong <zhong950419@gmail.com>
Discussion: https://postgr.es/m/CAL4mQLAp562c1rCgg2Dqx6TBSdOk6vOvaLFjwFG9E6k4uwvJJw@mail.gmail.com
---
 src/test/recovery/meson.build                |  1 +
 src/test/recovery/t/058_vm_clear_truncate.pl | 88 ++++++++++++++++++++
 2 files changed, 89 insertions(+)
 create mode 100644 src/test/recovery/t/058_vm_clear_truncate.pl

diff --git a/src/test/recovery/meson.build b/src/test/recovery/meson.build
index ebb12dd8766..bb28f9cfb81 100644
--- a/src/test/recovery/meson.build
+++ b/src/test/recovery/meson.build
@@ -66,6 +66,7 @@ tests += {
       't/055_cascade_reconnect.pl',
       't/056_standby_snapshot_export.pl',
       't/057_snapshot_commit_race.pl',
+      't/058_vm_clear_truncate.pl',
     ],
   },
 }
diff --git a/src/test/recovery/t/058_vm_clear_truncate.pl b/src/test/recovery/t/058_vm_clear_truncate.pl
new file mode 100644
index 00000000000..294eb7de95e
--- /dev/null
+++ b/src/test/recovery/t/058_vm_clear_truncate.pl
@@ -0,0 +1,88 @@
+
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# A standby must be able to restart when WAL it replays again clears
+# visibility map bits on a VM page that a later, already replayed,
+# truncation removed.  With full_page_writes off, redo cannot restore the
+# VM page from an image in the clearing record, so it has to cope with the
+# page not existing.  wal_consistency_checking is on so that the page redo
+# recreates is checked as well.
+use strict;
+use warnings FATAL => 'all';
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf(
+	'postgresql.conf', qq{
+full_page_writes = off
+wal_consistency_checking = all
+autovacuum = off
+});
+$primary->start;
+$primary->backup('bkp');
+
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'bkp', has_streaming => 1);
+$standby->start;
+
+# Make every heap page all-visible, then make the standby create a
+# restartpoint, so that a restart replays the changes below again.  The
+# row lock clears the all-frozen bit of vm_upd's first page.  Otherwise
+# the cross-page update below would first log a lock record that clears
+# that bit, and the update would not be the first record to read the VM
+# page.
+$primary->safe_psql(
+	'postgres', q{
+CREATE TABLE vm_del (a int);
+CREATE TABLE vm_hot (a int) WITH (fillfactor = 50);
+CREATE TABLE vm_upd (a int);
+INSERT INTO vm_del SELECT generate_series(1, 1000);
+INSERT INTO vm_hot SELECT generate_series(1, 1000);
+INSERT INTO vm_upd SELECT generate_series(1, 1000);
+VACUUM (FREEZE) vm_del, vm_hot, vm_upd;
+SELECT a FROM vm_upd WHERE a = 1 FOR UPDATE;
+CHECKPOINT;
+});
+$primary->wait_for_replay_catchup($standby);
+$standby->safe_psql('postgres', 'CHECKPOINT');
+
+my $start_lsn =
+  $primary->safe_psql('postgres', 'SELECT pg_current_wal_insert_lsn()');
+
+# Clear VM bits through delete, same-page update (old VM block) and
+# cross-page update (new VM block), then truncate all three tables to
+# zero blocks.
+$primary->safe_psql(
+	'postgres', q{
+DELETE FROM vm_del;
+UPDATE vm_hot SET a = -a WHERE a = 1;
+UPDATE vm_upd SET a = -a WHERE a = 1;
+DELETE FROM vm_hot;
+DELETE FROM vm_upd;
+VACUUM vm_del, vm_hot, vm_upd;
+});
+is( $primary->safe_psql(
+		'postgres',
+		"SELECT sum(pg_relation_size(c, 'vm')) FROM unnest('{vm_del,vm_hot,vm_upd}'::regclass[]) c"
+	),
+	'0',
+	'VMs truncated on primary');
+$primary->wait_for_replay_catchup($standby);
+
+$standby->stop;
+my $log_offset = -s $standby->logfile;
+my $ret = $standby->start(fail_ok => 1);
+
+my $log = slurp_file($standby->logfile, $log_offset);
+my ($redo_lsn) = $log =~ /redo starts at ([0-9A-F]+\/[0-9A-F]+)/;
+ok( defined($redo_lsn)
+	  && $primary->safe_psql('postgres',
+		"SELECT '$redo_lsn'::pg_lsn < '$start_lsn'::pg_lsn") eq 't',
+	'redo after restart starts before the VM bits were cleared');
+ok($ret, 'standby restarts after replaying VM truncation');
+unlike($log, qr/invalid pages/, 'no invalid page references in standby log');
+
+done_testing();
-- 
2.37.1 (Apple Git-137.1)



^ permalink  raw  reply  [nested|flat] 7+ messages in thread

* Re: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation
@ 2026-10-02 07:00  rahul@rhyadav.dev
  parent: rahul@rhyadav.dev
  1 sibling, 1 reply; 7+ messages in thread

From: rahul@rhyadav.dev @ 2026-10-02 07:00 UTC (permalink / raw)
  To: shihao zhong <zhong950419@gmail.com>; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Nktpro <nktpro@gmail.com>; Pgsql Bugs <pgsql-bugs@lists.postgresql.org>; Melanie Plageman <melanieplageman@gmail.com>

Hi Shihao,

On Fri, 2 Oct 2026, shihao zhong wrote:
> 0002 is your diff. It had no commit message, so I put one together
> from your mail. Please correct it if it is wrong.

Thanks for putting v3 together.  The message is right for master and
19.  On 17 and 18 the function that already does this is
heap_xlog_visible(), so in those versions the sentence would be:

  heap_xlog_visible() already initializes a VM page that was read as
  zeros.

0002 could also carry the same "Backpatch-through: 17" as 0001.

I ran my reproducer, which follows the same steps as 0003, against v3
on master and the nocfbot versions on 19, 18 and 17.  Unpatched
master and 17 fail to restart with the invalid pages PANIC.  With v3
the standby restarts cleanly on all four branches, in each of these
setups:

  - full_page_writes = off
  - full_page_writes = off, wal_consistency_checking = all
  - full_page_writes = on, wal_consistency_checking = all

0003 itself passes here too on all four branches, and fails on
master without 0001 and 0002.

Regards,
Rahul Yadav








^ permalink  raw  reply  [nested|flat] 7+ messages in thread

* Re: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation
@ 2026-10-03 04:12  shihao zhong <zhong950419@gmail.com>
  parent: rahul@rhyadav.dev
  0 siblings, 0 replies; 7+ messages in thread

From: shihao zhong @ 2026-10-03 04:12 UTC (permalink / raw)
  To: rahul@rhyadav.dev; +Cc: Kirill Reshke <reshkekirill@gmail.com>; Nktpro <nktpro@gmail.com>; Pgsql Bugs <pgsql-bugs@lists.postgresql.org>; Melanie Plageman <melanieplageman@gmail.com>

Hi Rahul,

Thanks for testing v3 on all four branches.

You are right about 17 and 18. I will change that sentence to
heap_xlog_visible() in those two versions, and add
"Backpatch-through: 17" to 0002. Both are commit message changes
only, so I will fold them into the next version and not post a v4
just for that.

I added a CommitFest entry so this is tracked for the November
releases.

https://commitfest.postgresql.org/patch/7386/

Thanks,
Shihao

^ permalink  raw  reply  [nested|flat] 7+ messages in thread


end of thread, other threads:[~2026-10-03 04:12 UTC | newest]

Thread overview: 7+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2026-09-27 06:07 PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation Jacky Nguyen <nktpro@gmail.com>
2026-09-27 10:33 ` Kirill Reshke <reshkekirill@gmail.com>
2026-09-28 06:37   ` shihao zhong <zhong950419@gmail.com>
2026-10-01 15:07     ` rahul@rhyadav.dev
2026-10-02 05:15       ` shihao zhong <zhong950419@gmail.com>
2026-10-02 07:00       ` rahul@rhyadav.dev
2026-10-03 04:12         ` shihao zhong <zhong950419@gmail.com>

This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox