agora inbox for pgsql-bugs@postgresql.org  
help / color / mirror / Atom feed
From: rahul@rhyadav.dev
To: shihao zhong <zhong950419@gmail.com>
Cc: Kirill Reshke <reshkekirill@gmail.com>
Cc: Nktpro <nktpro@gmail.com>
Cc: Pgsql Bugs <pgsql-bugs@lists.postgresql.org>
Cc: Melanie Plageman <melanieplageman@gmail.com>
Subject: Re: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation
Date: Thu, 1 Oct 2026 17:07:40 +0200 (CEST)
Message-ID: <P2s-NV0--F-9@rhyadav.dev> (raw)
In-Reply-To: <CAGRkXqRZva+Wiw6jb9-owhLjGOtoDn-btoFjJmejmJYOMSEbzg@mail.gmail.com>
References: <CAL4mQLAp562c1rCgg2Dqx6TBSdOk6vOvaLFjwFG9E6k4uwvJJw@mail.gmail.com>
	<CALdSSPjEM90Y_6RQGox7OFJC+oAurWy5YSuXaH8Mr7m-TMOyWA@mail.gmail.com>
	<CAGRkXqRZva+Wiw6jb9-owhLjGOtoDn-btoFjJmejmJYOMSEbzg@mail.gmail.com>

Hi,

On Mon, 28 Sep 2026, shihao zhong wrote:
> Melanie's v2-0002 in the VM clear thread [1] is the same change as
> candidate (a), and it fixes this report too.

I tested Melanie's v2 series as posted (0001-0003, which apply to
REL_19_STABLE) and candidate (a) on REL_18_STABLE, using the steps
from Shihao's test: full_page_writes = off and a plain standby
restart.  Both were debug builds with assertions, and the fix works:

- Unpatched, 18 and 19 fail to restart with "WAL contains references
  to invalid pages" after "page 0 of relation ..._vm does not exist"
  for each table.
- Patched, the standby restarts and its VM forks end up truncated
  again.
- On 18, reverting any one of the three RBM_ZERO_ON_ERROR changes
  brings the PANIC back, so the test covers each of them.

However, with wal_consistency_checking = all on the primary, the
patched standby still fails to restart:

  FATAL:  invalid page pd_lower 0 pd_upper 0 pd_special 0
  CONTEXT:  WAL redo at 0/03A80068 for Heap/DELETE: ...;
  blkref #1: rel 1663/5/16384, fork 2, blk 0 FPW

RBM_ZERO_ON_ERROR recreates the truncated VM page as all zeros.
verifyBackupPageConsistency() then masks it with heap_mask(), and
mask_unused_space() rejects a page with pd_lower 0.
heap_xlog_prune_freeze() and heap_xlog_multi_insert() already
initialize a VM page that was read as zeros, and doing the same at
the three VM clear sites fixes it.  The attached diff applies on top
of v2 on REL_19_STABLE; with it, the standby restarts with and
without wal_consistency_checking.  On 18, candidate (b) alone also
passes with wal_consistency_checking, since it never recreates the
page.

This only matters with wal_consistency_checking, but buildfarm
animals that use it could trip over Shihao's test once it's in.

Also, v2-0002 applies only on top of v2-0001 on REL_19_STABLE.  On
master it doesn't apply after master-v2-0001, 18 needs its own
version (candidate (a) applies there as is), and in 17 the redo code
is in heapam.c.

[1] https://postgr.es/m/CAAKRu_bApoksLDb-HX0GYciU3uLWqA1JagntaV8GP0%3D%2BidehHw%40mail.gmail.com

Regards,
Rahul Yadav
diff --git a/src/backend/access/heap/heapam_xlog.c b/src/backend/access/heap/heapam_xlog.c
index 4d99d99080..b87f38befc 100644
--- a/src/backend/access/heap/heapam_xlog.c
+++ b/src/backend/access/heap/heapam_xlog.c
@@ -89,6 +89,10 @@ heap_xlog_vm_clear(XLogReaderState *record,
 									  RBM_ZERO_ON_ERROR, false,
 									  &vmbuffer) == BLK_NEEDS_REDO)
 	{
+		/* initialize the page if it was read as zeros */
+		if (PageIsNew(BufferGetPage(vmbuffer)))
+			PageInit(BufferGetPage(vmbuffer), BLCKSZ, 0);
+
 		if (visibilitymap_clear(reln, heap_blkno, vmbuffer, flags))
 			PageSetLSN(BufferGetPage(vmbuffer), lsn);
 	}
@@ -855,6 +859,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 										  RBM_ZERO_ON_ERROR, false,
 										  &vmbuffer_new) == BLK_NEEDS_REDO)
 		{
+			/* initialize the page if it was read as zeros */
+			if (PageIsNew(BufferGetPage(vmbuffer_new)))
+				PageInit(BufferGetPage(vmbuffer_new), BLCKSZ, 0);
+
 			/*
 			 * If both the old and new heap pages were all-visible and their
 			 * VM bits are on the same VM page, that single VM page is
@@ -894,6 +902,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 										  RBM_ZERO_ON_ERROR, false,
 										  &vmbuffer_old) == BLK_NEEDS_REDO)
 		{
+			/* initialize the page if it was read as zeros */
+			if (PageIsNew(BufferGetPage(vmbuffer_old)))
+				PageInit(BufferGetPage(vmbuffer_old), BLCKSZ, 0);
+
 			if (visibilitymap_clear(reln, oldblk, vmbuffer_old,
 									VISIBILITYMAP_VALID_BITS))
 				PageSetLSN(BufferGetPage(vmbuffer_old), lsn);

Attachments:

  [text/plain] initialize-zeroed-vm-pages-on-v2.txt (1.5K, ../P2s-NV0--F-9@rhyadav.dev/2-initialize-zeroed-vm-pages-on-v2.txt)
  download | inline diff:
diff --git a/src/backend/access/heap/heapam_xlog.c b/src/backend/access/heap/heapam_xlog.c
index 4d99d99080..b87f38befc 100644
--- a/src/backend/access/heap/heapam_xlog.c
+++ b/src/backend/access/heap/heapam_xlog.c
@@ -89,6 +89,10 @@ heap_xlog_vm_clear(XLogReaderState *record,
 									  RBM_ZERO_ON_ERROR, false,
 									  &vmbuffer) == BLK_NEEDS_REDO)
 	{
+		/* initialize the page if it was read as zeros */
+		if (PageIsNew(BufferGetPage(vmbuffer)))
+			PageInit(BufferGetPage(vmbuffer), BLCKSZ, 0);
+
 		if (visibilitymap_clear(reln, heap_blkno, vmbuffer, flags))
 			PageSetLSN(BufferGetPage(vmbuffer), lsn);
 	}
@@ -855,6 +859,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 										  RBM_ZERO_ON_ERROR, false,
 										  &vmbuffer_new) == BLK_NEEDS_REDO)
 		{
+			/* initialize the page if it was read as zeros */
+			if (PageIsNew(BufferGetPage(vmbuffer_new)))
+				PageInit(BufferGetPage(vmbuffer_new), BLCKSZ, 0);
+
 			/*
 			 * If both the old and new heap pages were all-visible and their
 			 * VM bits are on the same VM page, that single VM page is
@@ -894,6 +902,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
 										  RBM_ZERO_ON_ERROR, false,
 										  &vmbuffer_old) == BLK_NEEDS_REDO)
 		{
+			/* initialize the page if it was read as zeros */
+			if (PageIsNew(BufferGetPage(vmbuffer_old)))
+				PageInit(BufferGetPage(vmbuffer_old), BLCKSZ, 0);
+
 			if (visibilitymap_clear(reln, oldblk, vmbuffer_old,
 									VISIBILITYMAP_VALID_BITS))
 				PageSetLSN(BufferGetPage(vmbuffer_old), lsn);

view thread (7+ messages)  latest in thread

Message-ID: <P2s-NV0--F-9@rhyadav.dev>
Permalink:  ../P2s-NV0--F-9@rhyadav.dev/
Also on:    postgresql.org/message-id/P2s-NV0--F-9@rhyadav.dev

reply

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Reply to all the recipients using the --to and --cc options:
  reply via email

  To: pgsql-bugs@postgresql.org
  Cc: rahul@rhyadav.dev, zhong950419@gmail.com, reshkekirill@gmail.com, nktpro@gmail.com, pgsql-bugs@lists.postgresql.org, melanieplageman@gmail.com
  Subject: Re: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation
  In-Reply-To: <P2s-NV0--F-9@rhyadav.dev>

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox