agora inbox for pgsql-bugs@postgresql.org
help / color / mirror / Atom feedFrom: rahul@rhyadav.dev
To: shihao zhong <zhong950419@gmail.com>
Cc: Kirill Reshke <reshkekirill@gmail.com>
Cc: Nktpro <nktpro@gmail.com>
Cc: Pgsql Bugs <pgsql-bugs@lists.postgresql.org>
Cc: Melanie Plageman <melanieplageman@gmail.com>
Subject: Re: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation
Date: Thu, 1 Oct 2026 17:07:40 +0200 (CEST)
Message-ID: <P2s-NV0--F-9@rhyadav.dev> (raw)
In-Reply-To: <CAGRkXqRZva+Wiw6jb9-owhLjGOtoDn-btoFjJmejmJYOMSEbzg@mail.gmail.com>
References: <CAL4mQLAp562c1rCgg2Dqx6TBSdOk6vOvaLFjwFG9E6k4uwvJJw@mail.gmail.com>
<CALdSSPjEM90Y_6RQGox7OFJC+oAurWy5YSuXaH8Mr7m-TMOyWA@mail.gmail.com>
<CAGRkXqRZva+Wiw6jb9-owhLjGOtoDn-btoFjJmejmJYOMSEbzg@mail.gmail.com>
Hi,
On Mon, 28 Sep 2026, shihao zhong wrote:
> Melanie's v2-0002 in the VM clear thread [1] is the same change as
> candidate (a), and it fixes this report too.
I tested Melanie's v2 series as posted (0001-0003, which apply to
REL_19_STABLE) and candidate (a) on REL_18_STABLE, using the steps
from Shihao's test: full_page_writes = off and a plain standby
restart. Both were debug builds with assertions, and the fix works:
- Unpatched, 18 and 19 fail to restart with "WAL contains references
to invalid pages" after "page 0 of relation ..._vm does not exist"
for each table.
- Patched, the standby restarts and its VM forks end up truncated
again.
- On 18, reverting any one of the three RBM_ZERO_ON_ERROR changes
brings the PANIC back, so the test covers each of them.
However, with wal_consistency_checking = all on the primary, the
patched standby still fails to restart:
FATAL: invalid page pd_lower 0 pd_upper 0 pd_special 0
CONTEXT: WAL redo at 0/03A80068 for Heap/DELETE: ...;
blkref #1: rel 1663/5/16384, fork 2, blk 0 FPW
RBM_ZERO_ON_ERROR recreates the truncated VM page as all zeros.
verifyBackupPageConsistency() then masks it with heap_mask(), and
mask_unused_space() rejects a page with pd_lower 0.
heap_xlog_prune_freeze() and heap_xlog_multi_insert() already
initialize a VM page that was read as zeros, and doing the same at
the three VM clear sites fixes it. The attached diff applies on top
of v2 on REL_19_STABLE; with it, the standby restarts with and
without wal_consistency_checking. On 18, candidate (b) alone also
passes with wal_consistency_checking, since it never recreates the
page.
This only matters with wal_consistency_checking, but buildfarm
animals that use it could trip over Shihao's test once it's in.
Also, v2-0002 applies only on top of v2-0001 on REL_19_STABLE. On
master it doesn't apply after master-v2-0001, 18 needs its own
version (candidate (a) applies there as is), and in 17 the redo code
is in heapam.c.
[1] https://postgr.es/m/CAAKRu_bApoksLDb-HX0GYciU3uLWqA1JagntaV8GP0%3D%2BidehHw%40mail.gmail.com
Regards,
Rahul Yadav
diff --git a/src/backend/access/heap/heapam_xlog.c b/src/backend/access/heap/heapam_xlog.c
index 4d99d99080..b87f38befc 100644
--- a/src/backend/access/heap/heapam_xlog.c
+++ b/src/backend/access/heap/heapam_xlog.c
@@ -89,6 +89,10 @@ heap_xlog_vm_clear(XLogReaderState *record,
RBM_ZERO_ON_ERROR, false,
&vmbuffer) == BLK_NEEDS_REDO)
{
+ /* initialize the page if it was read as zeros */
+ if (PageIsNew(BufferGetPage(vmbuffer)))
+ PageInit(BufferGetPage(vmbuffer), BLCKSZ, 0);
+
if (visibilitymap_clear(reln, heap_blkno, vmbuffer, flags))
PageSetLSN(BufferGetPage(vmbuffer), lsn);
}
@@ -855,6 +859,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
RBM_ZERO_ON_ERROR, false,
&vmbuffer_new) == BLK_NEEDS_REDO)
{
+ /* initialize the page if it was read as zeros */
+ if (PageIsNew(BufferGetPage(vmbuffer_new)))
+ PageInit(BufferGetPage(vmbuffer_new), BLCKSZ, 0);
+
/*
* If both the old and new heap pages were all-visible and their
* VM bits are on the same VM page, that single VM page is
@@ -894,6 +902,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
RBM_ZERO_ON_ERROR, false,
&vmbuffer_old) == BLK_NEEDS_REDO)
{
+ /* initialize the page if it was read as zeros */
+ if (PageIsNew(BufferGetPage(vmbuffer_old)))
+ PageInit(BufferGetPage(vmbuffer_old), BLCKSZ, 0);
+
if (visibilitymap_clear(reln, oldblk, vmbuffer_old,
VISIBILITYMAP_VALID_BITS))
PageSetLSN(BufferGetPage(vmbuffer_old), lsn);
Attachments:
[text/plain] initialize-zeroed-vm-pages-on-v2.txt (1.5K, ../P2s-NV0--F-9@rhyadav.dev/2-initialize-zeroed-vm-pages-on-v2.txt)
download | inline diff:
diff --git a/src/backend/access/heap/heapam_xlog.c b/src/backend/access/heap/heapam_xlog.c
index 4d99d99080..b87f38befc 100644
--- a/src/backend/access/heap/heapam_xlog.c
+++ b/src/backend/access/heap/heapam_xlog.c
@@ -89,6 +89,10 @@ heap_xlog_vm_clear(XLogReaderState *record,
RBM_ZERO_ON_ERROR, false,
&vmbuffer) == BLK_NEEDS_REDO)
{
+ /* initialize the page if it was read as zeros */
+ if (PageIsNew(BufferGetPage(vmbuffer)))
+ PageInit(BufferGetPage(vmbuffer), BLCKSZ, 0);
+
if (visibilitymap_clear(reln, heap_blkno, vmbuffer, flags))
PageSetLSN(BufferGetPage(vmbuffer), lsn);
}
@@ -855,6 +859,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
RBM_ZERO_ON_ERROR, false,
&vmbuffer_new) == BLK_NEEDS_REDO)
{
+ /* initialize the page if it was read as zeros */
+ if (PageIsNew(BufferGetPage(vmbuffer_new)))
+ PageInit(BufferGetPage(vmbuffer_new), BLCKSZ, 0);
+
/*
* If both the old and new heap pages were all-visible and their
* VM bits are on the same VM page, that single VM page is
@@ -894,6 +902,10 @@ heap_xlog_update(XLogReaderState *record, bool hot_update)
RBM_ZERO_ON_ERROR, false,
&vmbuffer_old) == BLK_NEEDS_REDO)
{
+ /* initialize the page if it was read as zeros */
+ if (PageIsNew(BufferGetPage(vmbuffer_old)))
+ PageInit(BufferGetPage(vmbuffer_old), BLCKSZ, 0);
+
if (visibilitymap_clear(reln, oldblk, vmbuffer_old,
VISIBILITYMAP_VALID_BITS))
PageSetLSN(BufferGetPage(vmbuffer_old), lsn);
view thread (7+ messages) latest in thread
Message-ID: <P2s-NV0--F-9@rhyadav.dev>
Permalink: ../P2s-NV0--F-9@rhyadav.dev/
Also on: postgresql.org/message-id/P2s-NV0--F-9@rhyadav.dev
reply
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Reply to all the recipients using the --to and --cc options:
reply via email
To: pgsql-bugs@postgresql.org
Cc: rahul@rhyadav.dev, zhong950419@gmail.com, reshkekirill@gmail.com, nktpro@gmail.com, pgsql-bugs@lists.postgresql.org, melanieplageman@gmail.com
Subject: Re: PostgreSQL 18.6/17.11: standby PANIC on restart after VM truncation
In-Reply-To: <P2s-NV0--F-9@rhyadav.dev>
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox