agora inbox for pgsql-hackers@postgresql.org
help / color / mirror / Atom feedSnapshot export on a standby corrupts hint bits on subxact overflow
6+ messages / 2 participants
[nested] [flat]
* Snapshot export on a standby corrupts hint bits on subxact overflow
@ 2026-07-27 02:17 Peter Geoghegan <pg@bowt.ie>
2026-07-28 06:07 ` Re: Snapshot export on a standby corrupts hint bits on subxact overflow Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 1 reply; 6+ messages in thread
From: Peter Geoghegan @ 2026-07-27 02:17 UTC (permalink / raw)
To: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; +Cc: Andres Freund <andres@anarazel.de>; scott@scottray.io
I have heard periodic reports that indicate that pg_export_snapshot
can cause a standby to incorrectly set hint bits, which I suspected
was tied to subxact overflow. For example, Francisco Reinolds 2023 bug
report [1] describes symptoms that match what I'd heard elsewhere. I
was reminded of this by a recent bug report from Scott Ray [2] that
also involves snapshot import/export. That turned out to be an
unrelated issue, though it looks legit.
I decided to reinvestigate the problem today, with help from Claude
code. I found a bug that exactly matches the known symptoms. Attached
patch 0001 has a reproducer + draft bug fix. This is likely a bug in
2017 commit 6c2003f8. There's also a second patch 0002 that fixes
another bug found along the way (though that's much less serious than
the one that 0001 deals with).
The test case in 0001 shows a scenario where pg_export_snapshot on a
standby hands out a snapshot that claims no transaction is running,
which is wrong. A backend that imports it writes wrong hint bits,
which are then seen by every other session on that standby --
including sessions that never touched the exported snapshot.
User-visible symptoms include rows reappearing after deletion, rows
vanishing after insertion, and duplicate entries in unique indexes
(all symptoms that I've personally seen in the wild). pg_dump -j can
hit this when run on a replica.
Recall that HeapTupleSatisfiesMVCC treats an XID that the snapshot
says isn't running, and that clog doesn't show as committed, as
aborted. That's how it's possible for a snapshot that shows no running
xacts to corrupt hint bits instead of just giving temporary wrong
answers (of course an FPI is also likely to mask the problem in
practice, I saw that myself when I investigated the problem a few
years back).
0002 is a separate bug of the same general nature.
pg_current_snapshot() copies the active snapshot's xip array, so on a
standby it returns a pg_snapshot claiming nothing is in progress, and
pg_visible_in_snapshot answers "visible" for transactions that are
still running:
primary pg_current_snapshot() = 696:698:696 visible(696) = f
standby pg_current_snapshot() = 696:698: visible(696) = t
This is what we see for transaction 696, which is still running (and
whose deleted row is correctly still visible on both nodes).
Both bug fixes were entirely authored by Claude, so treat them
skeptically. But I'm sure that at least the first one is a real bug.
[1] https://postgr.es/m/17846-1a0e5ce976f4c01a@postgresql.org
[2] https://postgr.es/m/QpAansP4iVg_ttSs9x81PFAptL2sqR3AS06u8Jksm3_bHJvUwQjHOocRajbxBc3iiLdf9ZMC6gtXjZsx...
--
Peter Geoghegan
Attachments:
[application/x-patch] v1-0002-Read-the-in-progress-set-from-subxip-in-pg_curren.patch (3.9K, ../../CAH2-WzmHVeYY=pjz9x8DhhxVjXHX0pvoQ-MdiB1Tt6=o2GTiKg@mail.gmail.com/2-v1-0002-Read-the-in-progress-set-from-subxip-in-pg_curren.patch)
download | inline diff:
From 8fc6659411760c19ea3ebaea41d3c30bacdef7d7 Mon Sep 17 00:00:00 2001
From: Peter Geoghegan <pg@bowt.ie>
Date: Sun, 26 Jul 2026 21:18:45 -0400
Subject: [PATCH v1 2/2] Read the in-progress set from subxip in
pg_current_snapshot()
pg_current_snapshot() copied the active snapshot's xip array. A snapshot
taken during recovery keeps all of its running XIDs in subxip and leaves xip
empty, so on a standby the function returned a pg_snapshot claiming that
nothing was in progress, and pg_visible_in_snapshot() answered "visible" for
transactions that were still running:
primary pg_current_snapshot() = 696:698:696 visible(696) = f
standby pg_current_snapshot() = 696:698: visible(696) = t
while the row that transaction 696 had deleted was, correctly, still visible
on both nodes. The documented caveat about subtransaction IDs does not cover
this: 696 is a top-level XID. Unlike the export path this needs no subxid
overflow, since xip is empty on a standby either way.
Read the array the snapshot actually filled. XIDs of subtransactions do
appear in it on a standby, because recovery does not distinguish them, but
that is within the existing caveat and reporting a running subtransaction as
still running is the better error.
pg_current_snapshot() and txid_current_snapshot() share this implementation.
---
src/backend/utils/adt/xid8funcs.c | 20 +++++++++++++++++--
.../recovery/t/055_standby_snapshot_export.pl | 12 +++++++++++
2 files changed, 30 insertions(+), 2 deletions(-)
diff --git a/src/backend/utils/adt/xid8funcs.c b/src/backend/utils/adt/xid8funcs.c
index c607e78d9..03053996b 100644
--- a/src/backend/utils/adt/xid8funcs.c
+++ b/src/backend/utils/adt/xid8funcs.c
@@ -374,14 +374,30 @@ pg_current_snapshot(PG_FUNCTION_ARGS)
uint32 nxip,
i;
Snapshot cur;
+ TransactionId *xip;
FullTransactionId next_fxid = ReadNextFullTransactionId();
cur = GetActiveSnapshot();
if (cur == NULL)
elog(ERROR, "no active snapshot set");
+ /*
+ * A snapshot taken during recovery stores all of its XIDs in subxip and
+ * leaves xip empty, so read the in-progress set from there. Reading xip
+ * would report every running transaction as already completed.
+ */
+ if (cur->takenDuringRecovery)
+ {
+ nxip = cur->subxcnt;
+ xip = cur->subxip;
+ }
+ else
+ {
+ nxip = cur->xcnt;
+ xip = cur->xip;
+ }
+
/* allocate */
- nxip = cur->xcnt;
snap = palloc(PG_SNAPSHOT_SIZE(nxip));
/*
@@ -395,7 +411,7 @@ pg_current_snapshot(PG_FUNCTION_ARGS)
snap->nxip = nxip;
for (i = 0; i < nxip; i++)
snap->xip[i] =
- FullTransactionIdFromAllowableAt(next_fxid, cur->xip[i]);
+ FullTransactionIdFromAllowableAt(next_fxid, xip[i]);
/*
* We want them guaranteed to be in ascending order. This also removes
diff --git a/src/test/recovery/t/055_standby_snapshot_export.pl b/src/test/recovery/t/055_standby_snapshot_export.pl
index 139190cf2..975b434df 100644
--- a/src/test/recovery/t/055_standby_snapshot_export.pl
+++ b/src/test/recovery/t/055_standby_snapshot_export.pl
@@ -82,6 +82,18 @@ $o->query_safe(
$primary->safe_psql('postgres', 'INSERT INTO burner VALUES (0)');
$primary->wait_for_replay_catchup($standby);
+# pg_current_snapshot() reads the same pair of arrays and has to make the same
+# distinction. Reporting a running transaction as visible would contradict
+# what the very snapshot it came from shows. Unlike the export path this
+# needs no overflow: on a standby xip is always empty.
+is($standby->safe_psql(
+ 'postgres',
+ "SELECT pg_visible_in_snapshot('$u_xid'::xid8, pg_current_snapshot())"),
+ 'f', 'pg_visible_in_snapshot agrees a running XID is not visible');
+
+is($standby->safe_psql('postgres', 'SELECT count(*) FROM victim WHERE k = 7'),
+ 1, 'and its delete has not taken effect');
+
my $s1 = $standby->background_psql('postgres');
$s1->query_safe('BEGIN ISOLATION LEVEL REPEATABLE READ');
my $snap = $s1->query_safe('SELECT pg_export_snapshot()');
--
2.53.0
[application/x-patch] v1-0001-Don-t-discard-subxip-when-exporting-a-snapshot-ta.patch (15.8K, ../../CAH2-WzmHVeYY=pjz9x8DhhxVjXHX0pvoQ-MdiB1Tt6=o2GTiKg@mail.gmail.com/3-v1-0001-Don-t-discard-subxip-when-exporting-a-snapshot-ta.patch)
download | inline diff:
From 2c17ded222408a12caf80093aad5c745eeabcacc Mon Sep 17 00:00:00 2001
From: Peter Geoghegan <pg@bowt.ie>
Date: Sun, 26 Jul 2026 21:17:34 -0400
Subject: [PATCH v1 1/2] Don't discard subxip when exporting a snapshot taken
during recovery
A snapshot taken during recovery stores every running XID -- top-level ones
included -- in subxip, leaving xip empty. CopySnapshot(),
EstimateSnapshotSpace() and SerializeSnapshot() each make an exception for
such a snapshot when subxip has overflowed, since discarding the array would
lose the only record of what is running. ExportSnapshot() never got that
exception: it wrote "sof:1" and dropped the array whenever the snapshot was
suboverflowed, which on a standby happens as soon as some transaction on the
primary reports 64 subtransactions.
A backend importing such a file then held a snapshot whose in-progress set
was empty, and XidInMVCCSnapshot() reported every XID between xmin and xmax
as no longer running. HeapTupleSatisfiesMVCC() took live transactions for
aborted ones and stamped HEAP_XMIN_INVALID or HEAP_XMAX_INVALID on their
tuples. Those hint bits are set with InvalidTransactionId, which skips the
commit-LSN interlock, so they were applied immediately and were seen by every
other session on the standby, not just by the importer: rows that a committed
transaction had deleted came back to life, rows it had inserted disappeared,
and a unique key could be returned twice. Only a full page image from the
primary put the page right again.
Export the array for a snapshot taken during recovery, keeping "sof:1" so
that the importer still consults pg_subtrans for the subtransactions that
KnownAssignedXidsRemoveTree() has already dropped. "rec:" now has to be
written before "sof:", because ImportSnapshot() parses strictly front to back
and must know whether an overflowed array is still meaningful before reading
it. Nothing outside these two functions reads these files and pg_snapshots
is emptied at startup, so the format change needs no compatibility handling.
Such a snapshot usually belongs to a transaction that can have no XID, and so
no subcommitted children, but it outlives promotion: a transaction on the
promoted server can import one, write, and export again, so the children are
added here as they are in the non-overflowed case. The combined count can in
principle exceed what the importer will accept, since a recovery snapshot's
subxip is bounded by the same value; there is no way to represent such a
snapshot, so refuse to export it rather than write an array that is quietly
missing entries.
Also assert in SetTransactionSnapshot() that a snapshot taken during recovery
whose XID range is not empty has a non-empty subxip, which is what makes this
class of mistake visible rather than silent.
Reachable through pg_export_snapshot() on a standby since 6c2003f8a1b, and
pg_dump -j is enough to hit it.
---
src/backend/utils/time/snapmgr.c | 62 +++++-
src/test/recovery/meson.build | 1 +
.../recovery/t/055_standby_snapshot_export.pl | 191 ++++++++++++++++++
3 files changed, 250 insertions(+), 4 deletions(-)
create mode 100644 src/test/recovery/t/055_standby_snapshot_export.pl
diff --git a/src/backend/utils/time/snapmgr.c b/src/backend/utils/time/snapmgr.c
index bc98a4361..beb1a6391 100644
--- a/src/backend/utils/time/snapmgr.c
+++ b/src/backend/utils/time/snapmgr.c
@@ -548,6 +548,16 @@ SetTransactionSnapshot(Snapshot sourcesnap, VirtualTransactionId *sourcevxid,
CurrentSnapshot->takenDuringRecovery = sourcesnap->takenDuringRecovery;
/* NB: curcid should NOT be copied, it's a local matter */
+ /*
+ * A snapshot taken during recovery keeps every running XID in subxip, and
+ * its xmin is the oldest of them, so an empty subxip means that nothing
+ * was running. Catch a source that lost the array along the way: such a
+ * snapshot silently reports running transactions as no longer running.
+ */
+ Assert(!CurrentSnapshot->takenDuringRecovery ||
+ CurrentSnapshot->subxcnt > 0 ||
+ CurrentSnapshot->xmin == CurrentSnapshot->xmax);
+
CurrentSnapshot->snapXactCompletionCount = 0;
/*
@@ -1222,13 +1232,54 @@ ExportSnapshot(Snapshot snapshot)
if (addTopXid)
appendStringInfo(&buf, "xip:%u\n", topXid);
+ /*
+ * The importer has to know whether the snapshot was taken during recovery
+ * before it reads the subxid data, since that determines whether an
+ * overflowed subxip array is still meaningful. Emit it first.
+ */
+ appendStringInfo(&buf, "rec:%u\n", snapshot->takenDuringRecovery);
+
/*
* Similarly, we add our subcommitted child XIDs to the subxid data. Here,
* we have to cope with possible overflow.
+ *
+ * Ignore the subxid array if it has overflowed, unless the snapshot was
+ * taken during recovery - in that case, top-level XIDs are in subxip as
+ * well, and we mustn't lose them.
+ *
+ * Such a snapshot usually belongs to a transaction that can have no XID,
+ * and hence no subcommitted children, since it was taken while the server
+ * was still in recovery. It can outlive promotion, though: a transaction
+ * on the promoted server can import one and then write.
*/
if (snapshot->suboverflowed ||
snapshot->subxcnt + nchildren > GetMaxSnapshotSubxidCount())
+ {
appendStringInfoString(&buf, "sof:1\n");
+ if (snapshot->takenDuringRecovery)
+ {
+ /*
+ * The importer's array is bounded by the same value, so a snapshot
+ * that does not fit cannot be represented at all. Refusing to
+ * export it is the only honest answer: an array quietly missing
+ * entries tells the importer that transactions which are still
+ * running have finished.
+ */
+ if (snapshot->subxcnt + nchildren > GetMaxSnapshotSubxidCount())
+ ereport(ERROR,
+ (errcode(ERRCODE_PROGRAM_LIMIT_EXCEEDED),
+ errmsg("cannot export a snapshot containing %d transaction IDs",
+ snapshot->subxcnt + nchildren),
+ errdetail("Snapshots taken during recovery record every running transaction ID, and this one exceeds the limit of %d.",
+ GetMaxSnapshotSubxidCount())));
+
+ appendStringInfo(&buf, "sxcnt:%d\n", snapshot->subxcnt + nchildren);
+ for (int32 i = 0; i < snapshot->subxcnt; i++)
+ appendStringInfo(&buf, "sxp:%u\n", snapshot->subxip[i]);
+ for (int32 i = 0; i < nchildren; i++)
+ appendStringInfo(&buf, "sxp:%u\n", children[i]);
+ }
+ }
else
{
appendStringInfoString(&buf, "sof:0\n");
@@ -1238,7 +1289,6 @@ ExportSnapshot(Snapshot snapshot)
for (int32 i = 0; i < nchildren; i++)
appendStringInfo(&buf, "sxp:%u\n", children[i]);
}
- appendStringInfo(&buf, "rec:%u\n", snapshot->takenDuringRecovery);
/*
* Now write the text representation into a file. We first write to a
@@ -1492,9 +1542,15 @@ ImportSnapshot(const char *idstr)
for (i = 0; i < xcnt; i++)
snapshot.xip[i] = parseXidFromText("xip:", &filebuf, path);
+ snapshot.takenDuringRecovery = parseIntFromText("rec:", &filebuf, path);
snapshot.suboverflowed = parseIntFromText("sof:", &filebuf, path);
- if (!snapshot.suboverflowed)
+ /*
+ * An overflowed subxip array carries no information and is not written
+ * out, except for a snapshot taken during recovery: that keeps all running
+ * XIDs there, so it is written out and must be read back.
+ */
+ if (!snapshot.suboverflowed || snapshot.takenDuringRecovery)
{
snapshot.subxcnt = xcnt = parseIntFromText("sxcnt:", &filebuf, path);
@@ -1514,8 +1570,6 @@ ImportSnapshot(const char *idstr)
snapshot.subxip = NULL;
}
- snapshot.takenDuringRecovery = parseIntFromText("rec:", &filebuf, path);
-
/*
* Do some additional sanity checking, just to protect ourselves. We
* don't trouble to check the array elements, just the most critical
diff --git a/src/test/recovery/meson.build b/src/test/recovery/meson.build
index ad0d85f41..8f24a1686 100644
--- a/src/test/recovery/meson.build
+++ b/src/test/recovery/meson.build
@@ -63,6 +63,7 @@ tests += {
't/052_checkpoint_segment_missing.pl',
't/053_standby_login_event_trigger.pl',
't/054_unlogged_sequence_promotion.pl',
+ 't/055_standby_snapshot_export.pl',
],
},
}
diff --git a/src/test/recovery/t/055_standby_snapshot_export.pl b/src/test/recovery/t/055_standby_snapshot_export.pl
new file mode 100644
index 000000000..139190cf2
--- /dev/null
+++ b/src/test/recovery/t/055_standby_snapshot_export.pl
@@ -0,0 +1,191 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+#
+# Test snapshot export and import on a standby.
+#
+# A snapshot taken during recovery stores every running XID -- top-level ones
+# included -- in subxip, leaving xip empty. Once such a snapshot is also
+# marked suboverflowed, which happens as soon as some transaction on the
+# primary reports 64 subtransactions, exporting it must still write subxip
+# out: an importer that loses it believes nothing at all is running between
+# xmin and xmax, treats live transactions as aborted, and stamps
+# HEAP_XMIN_INVALID / HEAP_XMAX_INVALID on their tuples. Those hint bits are
+# then seen by every other session on the standby.
+
+use strict;
+use warnings FATAL => 'all';
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Enough XID-acquiring subtransactions to arm lastOverflowedXid on the
+# standby, which happens at 64, and to overflow the primary's own subxid
+# cache at 65 so that later xl_running_xacts records keep it armed.
+my $nsubxacts = 80;
+
+# Checksums are off on purpose. With XLogHintBitIsNeeded(),
+# MarkSharedBufferDirtyHint() declines to dirty a page for a hint bit set
+# during recovery, so a wrong hint would live in shared buffers only.
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', q[
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+]);
+$primary->start;
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', q[
+hot_standby_feedback = off
+max_standby_streaming_delay = -1
+]);
+$standby->start;
+
+$primary->safe_psql(
+ 'postgres', q[
+CREATE TABLE victim(k int PRIMARY KEY, pad text);
+INSERT INTO victim SELECT g, repeat('x', 1200) FROM generate_series(1, 40) g;
+CREATE TABLE burner(i int);
+]);
+
+# Hold the primary's removal horizon below U so that the version U deletes
+# stays RECENTLY_DEAD. Otherwise the re-INSERT prunes it on the primary and
+# replay of the prune record removes the evidence from the standby.
+my $guard = $primary->background_psql('postgres');
+$guard->query_safe(
+ 'BEGIN ISOLATION LEVEL REPEATABLE READ; SELECT count(*) FROM victim');
+
+# U deletes a row and stays open, so every snapshot taken from here on must
+# report its XID as running.
+my $u = $primary->background_psql('postgres');
+$u->query_safe('BEGIN');
+$u->query_safe('DELETE FROM victim WHERE k = 7');
+my $u_xid = $u->query_safe('SELECT pg_current_xact_id()');
+
+# O overflows the subxid cache and stays open, keeping lastOverflowedXid
+# armed on the standby for as long as it lives.
+my $o = $primary->background_psql('postgres');
+$o->query_safe('BEGIN');
+$o->query_safe(
+ qq[DO \$\$ BEGIN
+ FOR i IN 1..$nsubxacts LOOP
+ BEGIN INSERT INTO burner VALUES (i); EXCEPTION WHEN OTHERS THEN NULL; END;
+ END LOOP; END \$\$]);
+
+# Commit something so that the standby's xmax ends up past U's XID. Without
+# this, XidInMVCCSnapshot() answers via its xid >= xmax range test and the
+# test would pass without exercising anything.
+$primary->safe_psql('postgres', 'INSERT INTO burner VALUES (0)');
+$primary->wait_for_replay_catchup($standby);
+
+my $s1 = $standby->background_psql('postgres');
+$s1->query_safe('BEGIN ISOLATION LEVEL REPEATABLE READ');
+my $snap = $s1->query_safe('SELECT pg_export_snapshot()');
+
+my $file = slurp_file($standby->data_dir . "/pg_snapshots/$snap");
+note("exported snapshot $snap:\n$file");
+
+like($file, qr/^rec:1$/m, 'snapshot was taken during recovery');
+like($file, qr/^sof:1$/m, 'snapshot is suboverflowed');
+
+# Without these the test could pass while exercising nothing: an xmax below
+# U's XID answers "not running" through the plain range test instead.
+my ($xmin) = $file =~ /^xmin:(\d+)$/m;
+my ($xmax) = $file =~ /^xmax:(\d+)$/m;
+cmp_ok($xmin, '<=', $u_xid, 'exported xmin does not follow the running XID');
+cmp_ok($u_xid, '<', $xmax, 'running XID precedes exported xmax');
+
+like($file, qr/^sxcnt:[1-9]/m,
+ 'suboverflowed recovery snapshot exports its subxip array');
+like($file, qr/^sxp:$u_xid$/m, 'running XID is exported');
+
+# Import the snapshot and read the victim page. This answers correctly --
+# a transaction that is still running and one that aborted are
+# indistinguishable here -- but it is what writes the hint bits.
+my $s2 = $standby->background_psql('postgres');
+$s2->query_safe(
+ qq[BEGIN ISOLATION LEVEL REPEATABLE READ;
+ SET TRANSACTION SNAPSHOT '$snap';
+ SET enable_indexscan = off;
+ SET enable_bitmapscan = off;
+ SET enable_indexonlyscan = off]);
+is($s2->query_safe('SELECT count(*) FROM victim'),
+ 40, 'importing backend sees its own snapshot');
+$s2->query_safe('COMMIT');
+$s1->query_safe('COMMIT');
+
+# U commits and the freed key is used again.
+$u->query_safe('COMMIT');
+$primary->safe_psql('postgres',
+ q[INSERT INTO victim VALUES (7, repeat('y', 1200))]);
+$o->query_safe('COMMIT');
+$primary->wait_for_replay_catchup($standby);
+
+# Sessions that never touched the exported snapshot must agree with the
+# primary. A stale HEAP_XMAX_INVALID on the version U deleted brings the old
+# row back to life, so the key is returned twice.
+is($standby->safe_psql('postgres', 'SELECT count(*) FROM victim WHERE k = 7'),
+ 1, 'standby sees the re-inserted row once');
+
+is( $standby->safe_psql(
+ 'postgres',
+ 'SELECT count(*) FROM (SELECT k FROM victim GROUP BY k HAVING count(*) > 1) d'
+ ),
+ 0, 'no duplicate keys on standby');
+
+my $digest = q[SELECT md5(string_agg(k::text, ',' ORDER BY k)) FROM victim];
+is($standby->safe_psql('postgres', $digest),
+ $primary->safe_psql('postgres', $digest),
+ 'primary and standby agree on table contents');
+
+# pg_snapshots is only emptied at startup and ImportSnapshot() takes rec: from
+# the file rather than from RecoveryInProgress(), so a snapshot exported by a
+# standby outlives promotion and keeps taking XidInMVCCSnapshot()'s recovery
+# branch on a server that is no longer in recovery.
+my $s3 = $standby->background_psql('postgres');
+$s3->query_safe('BEGIN ISOLATION LEVEL REPEATABLE READ');
+my $snap2 = $s3->query_safe('SELECT pg_export_snapshot()');
+my $before = $s3->query_safe('SELECT count(*) FROM victim');
+
+$standby->promote;
+$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery()')
+ or die "standby never finished promotion";
+
+my $s4 = $standby->background_psql('postgres');
+$s4->query_safe(
+ qq[BEGIN ISOLATION LEVEL REPEATABLE READ;
+ SET TRANSACTION SNAPSHOT '$snap2']);
+is($s4->query_safe('SELECT count(*) FROM victim'),
+ $before, 'recovery-taken snapshot still imports after promotion');
+
+# That importing transaction is read-write, since the read-only requirement
+# binds only a SERIALIZABLE source. So it can acquire subcommitted children
+# and export again, which hands ExportSnapshot() a snapshot taken during
+# recovery together with a non-empty children array.
+$s4->query_safe('SAVEPOINT sp');
+$s4->query_safe(q[INSERT INTO victim VALUES (5000, repeat('z', 10))]);
+$s4->query_safe('RELEASE sp');
+my $snap3 = $s4->query_safe('SELECT pg_export_snapshot()');
+
+my $file3 = slurp_file($standby->data_dir . "/pg_snapshots/$snap3");
+like($file3, qr/^sxcnt:[1-9]/m,
+ 'a recovery-taken snapshot re-exports its subxip array after a write');
+
+$s4->query_safe('COMMIT');
+$s3->query_safe('COMMIT');
+
+$guard->quit;
+$u->quit;
+$o->quit;
+$s1->quit;
+$s2->quit;
+$s3->quit;
+$s4->quit;
+$standby->stop;
+$primary->stop;
+
+done_testing();
--
2.53.0
^ permalink raw reply [nested|flat] 6+ messages in thread
* Re: Snapshot export on a standby corrupts hint bits on subxact overflow
2026-07-27 02:17 Snapshot export on a standby corrupts hint bits on subxact overflow Peter Geoghegan <pg@bowt.ie>
@ 2026-07-28 06:07 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-07-29 09:36 ` Re: Snapshot export on a standby corrupts hint bits on subxact overflow Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 1 reply; 6+ messages in thread
From: Bertrand Drouvot @ 2026-07-28 06:07 UTC (permalink / raw)
To: Peter Geoghegan <pg@bowt.ie>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; Andres Freund <andres@anarazel.de>; scott@scottray.io
Hi,
On Sun, Jul 26, 2026 at 10:17:21PM -0400, Peter Geoghegan wrote:
> I decided to reinvestigate the problem today, with help from Claude
> code. I found a bug that exactly matches the known symptoms. Attached
> patch 0001 has a reproducer + draft bug fix. This is likely a bug in
> 2017 commit 6c2003f8. There's also a second patch 0002 that fixes
> another bug found along the way (though that's much less serious than
> the one that 0001 deals with).
Thanks for having looked at it! The first problem is also something I worked
on in the past without success.
> The test case in 0001 shows a scenario where pg_export_snapshot on a
> standby hands out a snapshot that claims no transaction is running,
> which is wrong. A backend that imports it writes wrong hint bits,
> which are then seen by every other session on that standby --
> including sessions that never touched the exported snapshot.
> User-visible symptoms include rows reappearing after deletion, rows
> vanishing after insertion, and duplicate entries in unique indexes
> (all symptoms that I've personally seen in the wild).
The fix in 0001 makes sense to me.
Some comments:
=== 1
+ if (snapshot->subxcnt + nchildren > GetMaxSnapshotSubxidCount())
+ ereport(ERROR,
+ (errcode(ERRCODE_PROGRAM_LIMIT_EXCEEDED),
Yes, this check is needed here.
+ for (int32 i = 0; i < snapshot->subxcnt; i++)
+ appendStringInfo(&buf, "sxp:%u\n", snapshot->subxip[i]);
+ for (int32 i = 0; i < nchildren; i++)
+ appendStringInfo(&buf, "sxp:%u\n", children[i]);
We count every existing subxip entry and every committed child without filtering
against snapshot->xmax. A committed child created after the imported snapshot
was taken can have an XID >= snapshot->xmax. I think that such an entry is
unnecessary for visibility, since XidInMVCCSnapshot() classifies every XID >= xmax
as still in progress before looking at subxip.
It can be observed by applying the attached post-xmax.txt on top of 0001, which
produces:
# BDT snapshot saved for post-promotion re-export 00000007-00000002-1: sof=1
# BDT re-exported snapshot 00000000-00000004-1: sof=1; xmax=781; sxp=[698, 763, 764, 765, 766, 767, 768, 769, 770, 771, 772, 773, 774, 775, 776, 777, 778, 782]; sxp >= xmax=[782]
We can see that xmax=781, and that 782 has been exported.
That's not a visibility issue, however, those entries count toward GetMaxSnapshotSubxidCount()
and could therefore cause an avoidable PROGRAM_LIMIT_EXCEEDED.
=== 2
> 0002 is a separate bug of the same general nature.
+ if (cur->takenDuringRecovery)
+ {
+ nxip = cur->subxcnt;
+ xip = cur->subxip;
+ }
+ else
SnapshotData.subxip permits entries at or above xmax, whereas pg_snapshot.xip
requires every entry to satisfy xmin <= xip[i] < xmax.
So, I think filtering is appropriate in both places:
1/ In ExportSnapshot(), do not include recovery subxip entries and committed
child XIDs at or above xmax when counting and serializing them, so unnecessary
entries do not consume the limited recovery subxip capacity.
2/ In pg_current_snapshot(), do not include source XIDs outside [xmin, xmax),
so that it enforces the rule regardless of how the source snapshot was produced.
=== 3
0002 can now expose subtransaction IDs, so I think some comments:
"Note that only top-level transaction IDs are exposed to user sessions"
"Note that only top-transaction XIDs are included in the snapshot"
and the docs for pg_current_snapshot() and xip_list need updates to mention the
recovery specific exception.
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
diff --git a/src/test/recovery/t/055_standby_snapshot_export.pl b/src/test/recovery/t/055_standby_snapshot_export.pl
index 139190cf231..4cf3e78c766 100644
--- a/src/test/recovery/t/055_standby_snapshot_export.pl
+++ b/src/test/recovery/t/055_standby_snapshot_export.pl
@@ -122,7 +122,6 @@ $s1->query_safe('COMMIT');
$u->query_safe('COMMIT');
$primary->safe_psql('postgres',
q[INSERT INTO victim VALUES (7, repeat('y', 1200))]);
-$o->query_safe('COMMIT');
$primary->wait_for_replay_catchup($standby);
# Sessions that never touched the exported snapshot must agree with the
@@ -151,6 +150,13 @@ $s3->query_safe('BEGIN ISOLATION LEVEL REPEATABLE READ');
my $snap2 = $s3->query_safe('SELECT pg_export_snapshot()');
my $before = $s3->query_safe('SELECT count(*) FROM victim');
+my $file2 = slurp_file($standby->data_dir . "/pg_snapshots/$snap2");
+my ($sof2) = $file2 =~ /^sof:(\d+)$/m;
+note("BDT snapshot saved for post-promotion re-export $snap2: sof=$sof2");
+
+$o->query_safe('COMMIT');
+$primary->wait_for_replay_catchup($standby);
+
$standby->promote;
$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery()')
or die "standby never finished promotion";
@@ -172,6 +178,16 @@ $s4->query_safe('RELEASE sp');
my $snap3 = $s4->query_safe('SELECT pg_export_snapshot()');
my $file3 = slurp_file($standby->data_dir . "/pg_snapshots/$snap3");
+my ($sof3) = $file3 =~ /^sof:(\d+)$/m;
+my ($xmax3) = $file3 =~ /^xmax:(\d+)$/m;
+my @subxids3 = $file3 =~ /^sxp:(\d+)$/mg;
+my @post_xmax_subxids3 = grep { $_ >= $xmax3 } @subxids3;
+note(
+ "BDT re-exported snapshot $snap3: sof=$sof3; xmax=$xmax3; sxp=["
+ . join(', ', @subxids3)
+ . "]; sxp >= xmax=["
+ . join(', ', @post_xmax_subxids3)
+ . "]");
like($file3, qr/^sxcnt:[1-9]/m,
'a recovery-taken snapshot re-exports its subxip array after a write');
Attachments:
[text/plain] post-xmax.txt (1.8K, ../../amhHJiZyFA9s70Uq@bdtpg/2-post-xmax.txt)
download | inline diff:
diff --git a/src/test/recovery/t/055_standby_snapshot_export.pl b/src/test/recovery/t/055_standby_snapshot_export.pl
index 139190cf231..4cf3e78c766 100644
--- a/src/test/recovery/t/055_standby_snapshot_export.pl
+++ b/src/test/recovery/t/055_standby_snapshot_export.pl
@@ -122,7 +122,6 @@ $s1->query_safe('COMMIT');
$u->query_safe('COMMIT');
$primary->safe_psql('postgres',
q[INSERT INTO victim VALUES (7, repeat('y', 1200))]);
-$o->query_safe('COMMIT');
$primary->wait_for_replay_catchup($standby);
# Sessions that never touched the exported snapshot must agree with the
@@ -151,6 +150,13 @@ $s3->query_safe('BEGIN ISOLATION LEVEL REPEATABLE READ');
my $snap2 = $s3->query_safe('SELECT pg_export_snapshot()');
my $before = $s3->query_safe('SELECT count(*) FROM victim');
+my $file2 = slurp_file($standby->data_dir . "/pg_snapshots/$snap2");
+my ($sof2) = $file2 =~ /^sof:(\d+)$/m;
+note("BDT snapshot saved for post-promotion re-export $snap2: sof=$sof2");
+
+$o->query_safe('COMMIT');
+$primary->wait_for_replay_catchup($standby);
+
$standby->promote;
$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery()')
or die "standby never finished promotion";
@@ -172,6 +178,16 @@ $s4->query_safe('RELEASE sp');
my $snap3 = $s4->query_safe('SELECT pg_export_snapshot()');
my $file3 = slurp_file($standby->data_dir . "/pg_snapshots/$snap3");
+my ($sof3) = $file3 =~ /^sof:(\d+)$/m;
+my ($xmax3) = $file3 =~ /^xmax:(\d+)$/m;
+my @subxids3 = $file3 =~ /^sxp:(\d+)$/mg;
+my @post_xmax_subxids3 = grep { $_ >= $xmax3 } @subxids3;
+note(
+ "BDT re-exported snapshot $snap3: sof=$sof3; xmax=$xmax3; sxp=["
+ . join(', ', @subxids3)
+ . "]; sxp >= xmax=["
+ . join(', ', @post_xmax_subxids3)
+ . "]");
like($file3, qr/^sxcnt:[1-9]/m,
'a recovery-taken snapshot re-exports its subxip array after a write');
^ permalink raw reply [nested|flat] 6+ messages in thread
* Re: Snapshot export on a standby corrupts hint bits on subxact overflow
2026-07-27 02:17 Snapshot export on a standby corrupts hint bits on subxact overflow Peter Geoghegan <pg@bowt.ie>
2026-07-28 06:07 ` Re: Snapshot export on a standby corrupts hint bits on subxact overflow Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
@ 2026-07-29 09:36 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-08-24 23:07 ` Re: Snapshot export on a standby corrupts hint bits on subxact overflow Peter Geoghegan <pg@bowt.ie>
0 siblings, 1 reply; 6+ messages in thread
From: Bertrand Drouvot @ 2026-07-29 09:36 UTC (permalink / raw)
To: Peter Geoghegan <pg@bowt.ie>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; Andres Freund <andres@anarazel.de>; scott@scottray.io
Hi,
On Tue, Jul 28, 2026 at 06:07:34AM +0000, Bertrand Drouvot wrote:
> So, I think filtering is appropriate in both places:
>
> 1/ In ExportSnapshot(), do not include recovery subxip entries and committed
> child XIDs at or above xmax when counting and serializing them, so unnecessary
> entries do not consume the limited recovery subxip capacity.
>
> 2/ In pg_current_snapshot(), do not include source XIDs outside [xmin, xmax),
> so that it enforces the rule regardless of how the source snapshot was produced.
>
Following our off-list discussion, PFA v2 that addresses the comments above.
It adds xmax filtering, updates the documentation, and adds a few tests.
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
Attachments:
[text/x-diff] v2-0001-Don-t-discard-subxip-when-exporting-a-snapshot-ta.patch (18.1K, ../../amnJhkudengxKX+Y@bdtpg/2-v2-0001-Don-t-discard-subxip-when-exporting-a-snapshot-ta.patch)
download | inline diff:
From 3dc670619804794a6da2b5009cd58670e2082abb Mon Sep 17 00:00:00 2001
From: Peter Geoghegan <pg@bowt.ie>
Date: Sun, 26 Jul 2026 21:17:34 -0400
Subject: [PATCH v2 1/2] Don't discard subxip when exporting a snapshot taken
during recovery
A snapshot taken during recovery stores its in-progress set in subxip,
including every running top-level XID, and leaves xip empty.
ExportSnapshot() nevertheless discarded that array whenever suboverflowed
was set. An importer then treated live transactions as completed and
could set incorrect HEAP_XMIN_INVALID or HEAP_XMAX_INVALID hint bits on
the standby.
Export overflowed recovery subxip arrays and teach ImportSnapshot() to
read them. Keep the overflow flag so XidInMVCCSnapshot() still consults
pg_subtrans for children already removed from KnownAssignedXids.
A recovery snapshot can survive promotion, after which the importing
transaction can acquire committed children and export it again. Filter
recovery subxip entries and those children to [xmin, xmax) before
counting and writing them. Reject a snapshot that still exceeds the
import limit before pseudo-registering it, avoiding leftover cleanup
state on error.
Add a recovery TAP test covering the standby corruption scenario,
pg_subtrans fallback, and import/re-export across promotion.
Co-authored-by: Peter Geoghegan <pg@bowt.ie>
Co-authored-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Discussion: https://postgr.es/m/CAH2-WzmHVeYY%3Dpjz9x8DhhxVjXHX0pvoQ-MdiB1Tt6%3Do2GTiKg%40mail.gmail.com
---
src/backend/utils/time/snapmgr.c | 112 ++++++++-
src/test/recovery/meson.build | 1 +
.../recovery/t/055_standby_snapshot_export.pl | 227 ++++++++++++++++++
3 files changed, 330 insertions(+), 10 deletions(-)
32.3% src/backend/utils/time/
67.3% src/test/recovery/t/
diff --git a/src/backend/utils/time/snapmgr.c b/src/backend/utils/time/snapmgr.c
index bc98a4361bf..48c2f56b4d9 100644
--- a/src/backend/utils/time/snapmgr.c
+++ b/src/backend/utils/time/snapmgr.c
@@ -548,6 +548,17 @@ SetTransactionSnapshot(Snapshot sourcesnap, VirtualTransactionId *sourcevxid,
CurrentSnapshot->takenDuringRecovery = sourcesnap->takenDuringRecovery;
/* NB: curcid should NOT be copied, it's a local matter */
+ /*
+ * A snapshot taken during recovery keeps its in-progress set in subxip,
+ * including every running top-level XID. Its xmin is the oldest of them,
+ * so an empty subxip means that nothing was running. Catch a source that
+ * lost the array along the way: such a snapshot silently reports running
+ * transactions as no longer running.
+ */
+ Assert(!CurrentSnapshot->takenDuringRecovery ||
+ CurrentSnapshot->subxcnt > 0 ||
+ CurrentSnapshot->xmin == CurrentSnapshot->xmax);
+
CurrentSnapshot->snapXactCompletionCount = 0;
/*
@@ -1118,7 +1129,9 @@ ExportSnapshot(Snapshot snapshot)
TransactionId *children;
ExportedSnapshot *esnap;
int nchildren;
+ int nsubxids;
int addTopXid;
+ bool write_subxids;
StringInfoData buf;
FILE *f;
MemoryContext oldcxt;
@@ -1161,6 +1174,47 @@ ExportSnapshot(Snapshot snapshot)
*/
nchildren = xactGetCommittedChildren(&children);
+ /*
+ * SnapshotData allows subxip entries outside [xmin, xmax), but they carry
+ * no information in an exported snapshot and cannot be represented in a
+ * pg_snapshot. This matters for a recovery snapshot that survives
+ * promotion: committed children acquired afterwards are at or above the
+ * old xmax. Filter such entries before counting them.
+ */
+ if (snapshot->takenDuringRecovery)
+ {
+ nsubxids = 0;
+ for (int32 i = 0; i < snapshot->subxcnt; i++)
+ {
+ if (TransactionIdFollowsOrEquals(snapshot->subxip[i],
+ snapshot->xmin) &&
+ TransactionIdPrecedes(snapshot->subxip[i], snapshot->xmax))
+ nsubxids++;
+ }
+ for (int32 i = 0; i < nchildren; i++)
+ {
+ if (TransactionIdFollowsOrEquals(children[i], snapshot->xmin) &&
+ TransactionIdPrecedes(children[i], snapshot->xmax))
+ nsubxids++;
+ }
+
+ /*
+ * The importer's array is bounded by the same value, so a recovery
+ * snapshot that does not fit cannot be represented at all. Check
+ * before pseudo-registering the exported snapshot, so that an error
+ * cannot leave cleanup state behind.
+ */
+ if (nsubxids > GetMaxSnapshotSubxidCount())
+ ereport(ERROR,
+ (errcode(ERRCODE_PROGRAM_LIMIT_EXCEEDED),
+ errmsg("cannot export a snapshot containing %d transaction IDs",
+ nsubxids),
+ errdetail("Snapshots taken during recovery record running top-level transaction IDs in the subtransaction array, which is limited to %d entries.",
+ GetMaxSnapshotSubxidCount())));
+ }
+ else
+ nsubxids = snapshot->subxcnt + nchildren;
+
/*
* Generate file path for the snapshot. We start numbering of snapshots
* inside the transaction from 1.
@@ -1211,8 +1265,8 @@ ExportSnapshot(Snapshot snapshot)
*
* However, it could be that our topXid is after the xmax, in which case
* we shouldn't include it because xip[] members are expected to be before
- * xmax. (We need not make the same check for subxip[] members, see
- * snapshot.h.)
+ * xmax. SnapshotData does not require the same of subxip[] members (see
+ * snapshot.h), but the filtering above enforces it for the file format.
*/
addTopXid = (TransactionIdIsValid(topXid) &&
TransactionIdPrecedes(topXid, snapshot->xmax)) ? 1 : 0;
@@ -1222,23 +1276,57 @@ ExportSnapshot(Snapshot snapshot)
if (addTopXid)
appendStringInfo(&buf, "xip:%u\n", topXid);
+ /*
+ * The importer has to know whether the snapshot was taken during recovery
+ * before it reads the subxid data, since that determines whether an
+ * overflowed subxip array is still meaningful. Emit it first.
+ */
+ appendStringInfo(&buf, "rec:%u\n", snapshot->takenDuringRecovery);
+
/*
* Similarly, we add our subcommitted child XIDs to the subxid data. Here,
* we have to cope with possible overflow.
+ *
+ * Ignore the subxid array if it has overflowed, unless the snapshot was
+ * taken during recovery: in that case, top-level XIDs are in subxip as
+ * well, and we mustn't lose them.
+ *
+ * Such a snapshot usually belongs to a transaction that can have no XID,
+ * and hence no subcommitted children, since it was taken while the server
+ * was still in recovery. It can outlive promotion, though: a transaction
+ * on the promoted server can import one and then write.
*/
if (snapshot->suboverflowed ||
- snapshot->subxcnt + nchildren > GetMaxSnapshotSubxidCount())
+ nsubxids > GetMaxSnapshotSubxidCount())
+ {
appendStringInfoString(&buf, "sof:1\n");
+ write_subxids = snapshot->takenDuringRecovery;
+ }
else
{
appendStringInfoString(&buf, "sof:0\n");
- appendStringInfo(&buf, "sxcnt:%d\n", snapshot->subxcnt + nchildren);
+ write_subxids = true;
+ }
+
+ if (write_subxids)
+ {
+ appendStringInfo(&buf, "sxcnt:%d\n", nsubxids);
for (int32 i = 0; i < snapshot->subxcnt; i++)
- appendStringInfo(&buf, "sxp:%u\n", snapshot->subxip[i]);
+ {
+ if (!snapshot->takenDuringRecovery ||
+ (TransactionIdFollowsOrEquals(snapshot->subxip[i],
+ snapshot->xmin) &&
+ TransactionIdPrecedes(snapshot->subxip[i], snapshot->xmax)))
+ appendStringInfo(&buf, "sxp:%u\n", snapshot->subxip[i]);
+ }
for (int32 i = 0; i < nchildren; i++)
- appendStringInfo(&buf, "sxp:%u\n", children[i]);
+ {
+ if (!snapshot->takenDuringRecovery ||
+ (TransactionIdFollowsOrEquals(children[i], snapshot->xmin) &&
+ TransactionIdPrecedes(children[i], snapshot->xmax)))
+ appendStringInfo(&buf, "sxp:%u\n", children[i]);
+ }
}
- appendStringInfo(&buf, "rec:%u\n", snapshot->takenDuringRecovery);
/*
* Now write the text representation into a file. We first write to a
@@ -1492,9 +1580,15 @@ ImportSnapshot(const char *idstr)
for (i = 0; i < xcnt; i++)
snapshot.xip[i] = parseXidFromText("xip:", &filebuf, path);
+ snapshot.takenDuringRecovery = parseIntFromText("rec:", &filebuf, path);
snapshot.suboverflowed = parseIntFromText("sof:", &filebuf, path);
- if (!snapshot.suboverflowed)
+ /*
+ * An overflowed subxip array carries no information and is not written
+ * out, except for a snapshot taken during recovery: that keeps its
+ * in-progress XIDs there, so it is written out and must be read back.
+ */
+ if (!snapshot.suboverflowed || snapshot.takenDuringRecovery)
{
snapshot.subxcnt = xcnt = parseIntFromText("sxcnt:", &filebuf, path);
@@ -1514,8 +1608,6 @@ ImportSnapshot(const char *idstr)
snapshot.subxip = NULL;
}
- snapshot.takenDuringRecovery = parseIntFromText("rec:", &filebuf, path);
-
/*
* Do some additional sanity checking, just to protect ourselves. We
* don't trouble to check the array elements, just the most critical
diff --git a/src/test/recovery/meson.build b/src/test/recovery/meson.build
index ad0d85f4189..8f24a168614 100644
--- a/src/test/recovery/meson.build
+++ b/src/test/recovery/meson.build
@@ -63,6 +63,7 @@ tests += {
't/052_checkpoint_segment_missing.pl',
't/053_standby_login_event_trigger.pl',
't/054_unlogged_sequence_promotion.pl',
+ 't/055_standby_snapshot_export.pl',
],
},
}
diff --git a/src/test/recovery/t/055_standby_snapshot_export.pl b/src/test/recovery/t/055_standby_snapshot_export.pl
new file mode 100644
index 00000000000..704ac1141bc
--- /dev/null
+++ b/src/test/recovery/t/055_standby_snapshot_export.pl
@@ -0,0 +1,227 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+#
+# Test snapshot export and import on a standby.
+#
+# A snapshot taken during recovery stores its in-progress set, including
+# every running top-level XID, in subxip, leaving xip empty. Once it is also
+# marked suboverflowed, which happens as soon as some transaction on the
+# primary reports 64 subtransactions, exporting it must still write subxip
+# out: an importer that loses it believes nothing at all is running between
+# xmin and xmax, treats live transactions as aborted, and stamps
+# HEAP_XMIN_INVALID / HEAP_XMAX_INVALID on their tuples. Those hint bits are
+# then seen by every other session on the standby.
+
+use strict;
+use warnings FATAL => 'all';
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Enough XID-acquiring subtransactions to arm lastOverflowedXid on the
+# standby, which happens at 64, and to overflow the primary's own subxid
+# cache at 65 so that later xl_running_xacts records keep it armed.
+my $nsubxacts = 80;
+
+# Checksums are off on purpose. With XLogHintBitIsNeeded(),
+# MarkSharedBufferDirtyHint() declines to dirty a page for a hint bit set
+# during recovery, so a wrong hint would live in shared buffers only.
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1, no_data_checksums => 1);
+$primary->append_conf(
+ 'postgresql.conf', q[
+autovacuum = off
+checkpoint_timeout = 1h
+max_wal_size = 10GB
+]);
+$primary->start;
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf(
+ 'postgresql.conf', q[
+hot_standby_feedback = off
+max_standby_streaming_delay = -1
+]);
+$standby->start;
+
+$primary->safe_psql(
+ 'postgres', q[
+CREATE TABLE victim(k int PRIMARY KEY, pad text);
+INSERT INTO victim SELECT g, repeat('x', 1200) FROM generate_series(1, 40) g;
+CREATE TABLE burner(i int);
+]);
+
+# Hold the primary's removal horizon below U so that the version U deletes
+# stays RECENTLY_DEAD. Otherwise the re-INSERT prunes it on the primary and
+# replay of the prune record removes the evidence from the standby.
+my $guard = $primary->background_psql('postgres');
+$guard->query_safe(
+ 'BEGIN ISOLATION LEVEL REPEATABLE READ; SELECT count(*) FROM victim');
+
+# U deletes a row and stays open, so every snapshot taken from here on must
+# report its XID as running.
+my $u = $primary->background_psql('postgres');
+$u->query_safe('BEGIN');
+$u->query_safe('DELETE FROM victim WHERE k = 7');
+my $u_xid = $u->query_safe('SELECT pg_current_xact_id()');
+
+# O deletes another row in an early subtransaction, then overflows the subxid
+# cache and stays open. Recovery removes the deleting subtransaction's XID
+# from KnownAssignedXids, so a later visibility check of the tuple's xmax
+# must use pg_subtrans to map that child XID back to O. Keeping O open also
+# keeps lastOverflowedXid armed on the standby.
+my $o = $primary->background_psql('postgres');
+$o->query_safe('BEGIN');
+$o->query_safe('SAVEPOINT early');
+$o->query_safe('DELETE FROM victim WHERE k = 8');
+$o->query_safe('RELEASE early');
+$o->query_safe(
+ qq[DO \$\$ BEGIN
+ FOR i IN 1..$nsubxacts LOOP
+ BEGIN INSERT INTO burner VALUES (i); EXCEPTION WHEN OTHERS THEN NULL; END;
+ END LOOP; END \$\$]);
+
+# Commit something so that the standby's xmax ends up past U's XID. Without
+# this, XidInMVCCSnapshot() answers via its xid >= xmax range test and the
+# test would pass without exercising anything.
+$primary->safe_psql('postgres', 'INSERT INTO burner VALUES (0)');
+$primary->wait_for_replay_catchup($standby);
+
+my $s1 = $standby->background_psql('postgres');
+$s1->query_safe('BEGIN ISOLATION LEVEL REPEATABLE READ');
+my $snap = $s1->query_safe('SELECT pg_export_snapshot()');
+
+my $file = slurp_file($standby->data_dir . "/pg_snapshots/$snap");
+note("exported snapshot $snap:\n$file");
+
+like($file, qr/^rec:1$/m, 'snapshot was taken during recovery');
+like($file, qr/^sof:1$/m, 'snapshot is suboverflowed');
+
+# Without these the test could pass while exercising nothing: an xmax below
+# U's XID answers "not running" through the plain range test instead.
+my ($xmin) = $file =~ /^xmin:(\d+)$/m;
+my ($xmax) = $file =~ /^xmax:(\d+)$/m;
+cmp_ok($xmin, '<=', $u_xid, 'exported xmin does not follow the running XID');
+cmp_ok($u_xid, '<', $xmax, 'running XID precedes exported xmax');
+
+like($file, qr/^sxcnt:[1-9]/m,
+ 'suboverflowed recovery snapshot exports its subxip array');
+like($file, qr/^sxp:$u_xid$/m, 'running XID is exported');
+
+# Import the snapshot and read the victim page. This answers correctly,
+# a transaction that is still running and one that aborted are
+# indistinguishable here, but it is what writes the hint bits.
+my $s2 = $standby->background_psql('postgres');
+$s2->query_safe(
+ qq[BEGIN ISOLATION LEVEL REPEATABLE READ;
+ SET TRANSACTION SNAPSHOT '$snap';
+ SET enable_indexscan = off;
+ SET enable_bitmapscan = off;
+ SET enable_indexonlyscan = off]);
+is($s2->query_safe('SELECT count(*) FROM victim'),
+ 40, 'importing backend sees its own snapshot');
+$s2->query_safe('COMMIT');
+$s1->query_safe('COMMIT');
+
+# U commits and the freed key is used again.
+$u->query_safe('COMMIT');
+$primary->safe_psql('postgres',
+ q[INSERT INTO victim VALUES (7, repeat('y', 1200))]);
+$primary->wait_for_replay_catchup($standby);
+
+# Sessions that never touched the exported snapshot must agree with the
+# primary. A stale HEAP_XMAX_INVALID on the version U deleted brings the old
+# row back to life, so the key is returned twice.
+is( $standby->safe_psql(
+ 'postgres', 'SELECT count(*) FROM victim WHERE k = 7'),
+ 1,
+ 'standby sees the re-inserted row once');
+
+is( $standby->safe_psql(
+ 'postgres',
+ 'SELECT count(*) FROM (SELECT k FROM victim GROUP BY k HAVING count(*) > 1) d'
+ ),
+ 0,
+ 'no duplicate keys on standby');
+
+my $digest = q[SELECT md5(string_agg(k::text, ',' ORDER BY k)) FROM victim];
+is( $standby->safe_psql('postgres', $digest),
+ $primary->safe_psql('postgres', $digest),
+ 'primary and standby agree on table contents');
+
+# pg_snapshots is only emptied at startup and ImportSnapshot() takes rec: from
+# the file rather than from RecoveryInProgress(), so a snapshot exported by a
+# standby outlives promotion and keeps taking XidInMVCCSnapshot()'s recovery
+# branch on a server that is no longer in recovery.
+my $s3 = $standby->background_psql('postgres');
+$s3->query_safe('BEGIN ISOLATION LEVEL REPEATABLE READ');
+my $snap2 = $s3->query_safe('SELECT pg_export_snapshot()');
+my $before = $s3->query_safe('SELECT count(*) FROM victim');
+
+my $file2 = slurp_file($standby->data_dir . "/pg_snapshots/$snap2");
+like($file2, qr/^sof:1$/m,
+ 'snapshot saved for post-promotion re-export is suboverflowed');
+
+$o->query_safe('COMMIT');
+$primary->wait_for_replay_catchup($standby);
+
+is( $standby->safe_psql(
+ 'postgres', 'SELECT count(*) FROM victim WHERE k = 8'),
+ 0,
+ 'standby sees the committed subtransaction delete');
+
+$standby->promote;
+$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery()')
+ or die "standby never finished promotion";
+
+my $s4 = $standby->background_psql('postgres');
+$s4->query_safe(
+ qq[BEGIN ISOLATION LEVEL REPEATABLE READ;
+ SET TRANSACTION SNAPSHOT '$snap2']);
+is($s4->query_safe('SELECT count(*) FROM victim'),
+ $before, 'recovery-taken snapshot still imports after promotion');
+
+# That importing transaction is read-write, since the read-only requirement
+# binds only a SERIALIZABLE source. So it can acquire subcommitted children
+# and export again, which hands ExportSnapshot() a snapshot taken during
+# recovery together with a non-empty children array.
+$s4->query_safe('SAVEPOINT sp');
+$s4->query_safe(q[INSERT INTO victim VALUES (5000, repeat('z', 10))]);
+$s4->query_safe('RELEASE sp');
+my $snap3 = $s4->query_safe('SELECT pg_export_snapshot()');
+
+my $file3 = slurp_file($standby->data_dir . "/pg_snapshots/$snap3");
+my ($xmin3) = $file3 =~ /^xmin:(\d+)$/m;
+my ($xmax3) = $file3 =~ /^xmax:(\d+)$/m;
+my @subxids3 = $file3 =~ /^sxp:(\d+)$/mg;
+like($file3, qr/^sof:1$/m,
+ 're-exported recovery snapshot remains suboverflowed');
+like($file3, qr/^sxcnt:[1-9]/m,
+ 'a recovery-taken snapshot re-exports its subxip array after a write');
+is(scalar(grep { $_ < $xmin3 || $_ >= $xmax3 } @subxids3),
+ 0, 're-exported snapshot omits XIDs outside its range');
+
+my $s5 = $standby->background_psql('postgres');
+$s5->query_safe(
+ qq[BEGIN ISOLATION LEVEL REPEATABLE READ;
+ SET TRANSACTION SNAPSHOT '$snap3']);
+is($s5->query_safe('SELECT count(*) FROM victim'),
+ $before, 're-exported recovery snapshot can be imported');
+
+$s5->query_safe('COMMIT');
+$s4->query_safe('COMMIT');
+$s3->query_safe('COMMIT');
+
+$guard->quit;
+$u->quit;
+$o->quit;
+$s1->quit;
+$s2->quit;
+$s3->quit;
+$s4->quit;
+$s5->quit;
+$standby->stop;
+$primary->stop;
+
+done_testing();
--
2.34.1
[text/x-diff] v2-0002-Read-the-in-progress-set-from-subxip-in-pg_curren.patch (11.0K, ../../amnJhkudengxKX+Y@bdtpg/3-v2-0002-Read-the-in-progress-set-from-subxip-in-pg_curren.patch)
download | inline diff:
From 26eb08045d85821cdc662023b414775afae66702 Mon Sep 17 00:00:00 2001
From: Peter Geoghegan <pg@bowt.ie>
Date: Sun, 26 Jul 2026 21:18:45 -0400
Subject: [PATCH v2 2/2] Read the in-progress set from subxip in
pg_current_snapshot()
pg_current_snapshot() copied the active snapshot's xip array. A snapshot
taken during recovery keeps its in-progress XIDs in subxip and leaves xip
empty, so the function reported every transaction in [xmin, xmax) as
completed. pg_visible_in_snapshot() could therefore disagree with tuple
visibility on a standby. No subxid overflow is required.
Read the array populated by recovery and filter it to [xmin, xmax), as
required by pg_snapshot. Account for the larger recovery array in the
compile-time allocation bound. Recovery cannot distinguish top-level
and subtransaction XIDs, so update the source comments and documentation
to describe the subtransaction IDs that can appear in recovery snapshots.
Test the behavior without overflow through both pg_current_snapshot() and
the legacy txid_current_snapshot() interface. Also verify both interfaces
after importing a recovery snapshot across promotion, including range
validity and textual pg_snapshot input.
Co-authored-by: Peter Geoghegan <pg@bowt.ie>
Co-authored-by: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Discussion: https://postgr.es/m/CAH2-WzmHVeYY%3Dpjz9x8DhhxVjXHX0pvoQ-MdiB1Tt6%3Do2GTiKg%40mail.gmail.com
---
doc/src/sgml/func/func-info.sgml | 13 ++--
src/backend/utils/adt/xid8funcs.c | 66 ++++++++++++++-----
.../recovery/t/055_standby_snapshot_export.pl | 55 ++++++++++++++--
3 files changed, 107 insertions(+), 27 deletions(-)
14.5% doc/src/sgml/func/
49.5% src/backend/utils/adt/
35.8% src/test/recovery/t/
diff --git a/doc/src/sgml/func/func-info.sgml b/doc/src/sgml/func/func-info.sgml
index 122fc740f1a..bbde54ff0fa 100644
--- a/doc/src/sgml/func/func-info.sgml
+++ b/doc/src/sgml/func/func-info.sgml
@@ -2905,9 +2905,10 @@ acl | {postgres=arwdDxtm/postgres,foo=r/postgres}
<para>
Returns a current <firstterm>snapshot</firstterm>, a data structure
showing which transaction IDs are now in-progress.
- Only top-level transaction IDs are included in the snapshot;
- subtransaction IDs are not shown; see <xref linkend="subxacts"/>
- for details.
+ Normally, only top-level transaction IDs are included in the snapshot.
+ A snapshot taken during recovery can also include
+ subtransaction IDs because recovery cannot distinguish them from
+ top-level transaction IDs. See <xref linkend="subxacts"/> for details.
</para></entry>
</row>
@@ -3086,8 +3087,10 @@ acl | {postgres=arwdDxtm/postgres,foo=r/postgres}
ID that is <literal>xmin <= <replaceable>X</replaceable> <
xmax</literal> and not in this list was already completed at the time
of the snapshot, and thus is either visible or dead according to its
- commit status. This list does not include the transaction IDs of
- subtransactions (subxids).
+ commit status. This list normally does not include the transaction IDs
+ of subtransactions (subxids), but a snapshot taken during recovery can
+ include them because recovery cannot distinguish subtransaction IDs
+ from top-level transaction IDs.
</entry>
</row>
</tbody>
diff --git a/src/backend/utils/adt/xid8funcs.c b/src/backend/utils/adt/xid8funcs.c
index c607e78d9ac..0a398c584e6 100644
--- a/src/backend/utils/adt/xid8funcs.c
+++ b/src/backend/utils/adt/xid8funcs.c
@@ -3,11 +3,13 @@
*
* Export internal transaction IDs to user level.
*
- * Note that only top-level transaction IDs are exposed to user sessions.
- * This is important because xid8s frequently persist beyond the global
- * xmin horizon, or may even be shipped to other machines, so we cannot
- * rely on being able to correlate subtransaction IDs with their parents
- * via functions such as SubTransGetTopmostTransaction().
+ * Normally only top-level transaction IDs are exposed to user sessions.
+ * Snapshots taken during recovery can also expose subtransaction IDs, since
+ * recovery cannot distinguish them from top-level IDs. xid8s frequently
+ * persist beyond the global xmin horizon, or may even be shipped to other
+ * machines, so callers cannot rely on being able to correlate subtransaction
+ * IDs with their parents via functions such as
+ * SubTransGetTopmostTransaction().
*
* These functions are used to support the txid_XXX functions and the newer
* pg_current_xact_id, pg_current_snapshot and related fmgr functions, since
@@ -33,6 +35,7 @@
#include "libpq/pqformat.h"
#include "miscadmin.h"
#include "storage/lwlock.h"
+#include "storage/proc.h"
#include "storage/procarray.h"
#include "storage/procnumber.h"
#include "utils/builtins.h"
@@ -75,9 +78,11 @@ typedef struct
/*
* Compile-time limits on the procarray (MAX_BACKENDS processes plus
- * MAX_BACKENDS prepared transactions) guarantee nxip won't be too large.
+ * MAX_BACKENDS prepared transactions) and the number of subxids cached for
+ * each process guarantee nxip won't be too large.
*/
-StaticAssertDecl(MAX_BACKENDS * 2 <= PG_SNAPSHOT_MAX_NXIP,
+StaticAssertDecl((PGPROC_MAX_CACHED_SUBXIDS + 1) * MAX_BACKENDS * 2 <=
+ PG_SNAPSHOT_MAX_NXIP,
"possible overflow in pg_current_snapshot()");
@@ -365,37 +370,62 @@ pg_current_xact_id_if_assigned(PG_FUNCTION_ARGS)
*
* Return current snapshot
*
- * Note that only top-transaction XIDs are included in the snapshot.
+ * Snapshots taken during recovery can also include subtransaction XIDs because
+ * recovery cannot distinguish them from top-level XIDs.
*/
Datum
pg_current_snapshot(PG_FUNCTION_ARGS)
{
pg_snapshot *snap;
- uint32 nxip,
+ uint32 nxip = 0,
+ source_nxip,
i;
Snapshot cur;
+ TransactionId *xip;
FullTransactionId next_fxid = ReadNextFullTransactionId();
cur = GetActiveSnapshot();
if (cur == NULL)
elog(ERROR, "no active snapshot set");
+ /*
+ * A snapshot taken during recovery stores its in-progress XIDs in subxip
+ * and leaves xip empty, so read the in-progress set from there. Reading
+ * xip would report every running transaction as already completed.
+ */
+ if (cur->takenDuringRecovery)
+ {
+ source_nxip = cur->subxcnt;
+ xip = cur->subxip;
+ }
+ else
+ {
+ source_nxip = cur->xcnt;
+ xip = cur->xip;
+ }
+
/* allocate */
- nxip = cur->xcnt;
- snap = palloc(PG_SNAPSHOT_SIZE(nxip));
+ snap = palloc(PG_SNAPSHOT_SIZE(source_nxip));
/*
- * Fill. This is the current backend's active snapshot, so MyProc->xmin
- * is <= all these XIDs. As long as that remains so, oldestXid can't
- * advance past any of these XIDs. Hence, these XIDs remain allowable
- * relative to next_fxid.
+ * Fill. Unlike SnapshotData's subxip, pg_snapshot's xip cannot contain
+ * XIDs outside [xmin, xmax), so filter the source at this boundary. This
+ * is the current backend's active snapshot, so MyProc->xmin protects all
+ * retained XIDs from oldestXid. Hence, they remain allowable relative to
+ * next_fxid.
*/
snap->xmin = FullTransactionIdFromAllowableAt(next_fxid, cur->xmin);
snap->xmax = FullTransactionIdFromAllowableAt(next_fxid, cur->xmax);
+ for (i = 0; i < source_nxip; i++)
+ {
+ if (TransactionIdPrecedes(xip[i], cur->xmin) ||
+ TransactionIdFollowsOrEquals(xip[i], cur->xmax))
+ continue;
+
+ snap->xip[nxip++] =
+ FullTransactionIdFromAllowableAt(next_fxid, xip[i]);
+ }
snap->nxip = nxip;
- for (i = 0; i < nxip; i++)
- snap->xip[i] =
- FullTransactionIdFromAllowableAt(next_fxid, cur->xip[i]);
/*
* We want them guaranteed to be in ascending order. This also removes
diff --git a/src/test/recovery/t/055_standby_snapshot_export.pl b/src/test/recovery/t/055_standby_snapshot_export.pl
index 704ac1141bc..898bacefd8e 100644
--- a/src/test/recovery/t/055_standby_snapshot_export.pl
+++ b/src/test/recovery/t/055_standby_snapshot_export.pl
@@ -66,6 +66,33 @@ $u->query_safe('BEGIN');
$u->query_safe('DELETE FROM victim WHERE k = 7');
my $u_xid = $u->query_safe('SELECT pg_current_xact_id()');
+# Commit something so that the standby's xmax ends up past U's XID. Without
+# this, the visibility functions answer through their xid >= xmax range test
+# and the test would pass without examining the in-progress array.
+$primary->safe_psql('postgres', 'INSERT INTO burner VALUES (0)');
+$primary->wait_for_replay_catchup($standby);
+
+# A recovery snapshot stores its in-progress XIDs in subxip and leaves xip
+# empty. Verify the SQL interface before creating any subxid overflow.
+is( $standby->safe_psql(
+ 'postgres',
+ "SELECT pg_visible_in_snapshot('$u_xid'::xid8, pg_current_snapshot())"
+ ),
+ 'f',
+ 'pg_current_snapshot reports a running XID as not visible');
+
+is( $standby->safe_psql(
+ 'postgres',
+ "SELECT txid_visible_in_snapshot('$u_xid'::bigint, txid_current_snapshot())"
+ ),
+ 'f',
+ 'txid_current_snapshot reports a running XID as not visible');
+
+is( $standby->safe_psql(
+ 'postgres', 'SELECT count(*) FROM victim WHERE k = 7'),
+ 1,
+ 'the running transaction has not deleted its row');
+
# O deletes another row in an early subtransaction, then overflows the subxid
# cache and stays open. Recovery removes the deleting subtransaction's XID
# from KnownAssignedXids, so a later visibility check of the tuple's xmax
@@ -82,10 +109,10 @@ $o->query_safe(
BEGIN INSERT INTO burner VALUES (i); EXCEPTION WHEN OTHERS THEN NULL; END;
END LOOP; END \$\$]);
-# Commit something so that the standby's xmax ends up past U's XID. Without
-# this, XidInMVCCSnapshot() answers via its xid >= xmax range test and the
-# test would pass without exercising anything.
-$primary->safe_psql('postgres', 'INSERT INTO burner VALUES (0)');
+# Commit after O has overflowed to flush its preceding xid-assignment WAL, then
+# wait until the standby has removed those subxids and marked snapshots
+# overflowed.
+$primary->safe_psql('postgres', 'INSERT INTO burner VALUES (-1)');
$primary->wait_for_replay_catchup($standby);
my $s1 = $standby->background_psql('postgres');
@@ -209,6 +236,26 @@ $s5->query_safe(
is($s5->query_safe('SELECT count(*) FROM victim'),
$before, 're-exported recovery snapshot can be imported');
+is( $s5->query_safe(
+ 'SELECT count(*) '
+ . 'FROM pg_snapshot_xip(pg_current_snapshot()) AS x(xid) '
+ . 'WHERE xid < pg_snapshot_xmin(pg_current_snapshot()) '
+ . 'OR xid >= pg_snapshot_xmax(pg_current_snapshot())'),
+ 0,
+ 'pg_current_snapshot has no explicit XIDs outside its range');
+
+is( $s5->query_safe(
+ 'SELECT count(*) '
+ . 'FROM txid_snapshot_xip(txid_current_snapshot()) AS x(xid) '
+ . 'WHERE xid < txid_snapshot_xmin(txid_current_snapshot()) '
+ . 'OR xid >= txid_snapshot_xmax(txid_current_snapshot())'),
+ 0,
+ 'txid_current_snapshot has no explicit XIDs outside its range');
+
+my $current_snapshot = $s5->query_safe('SELECT pg_current_snapshot()::text');
+is($s5->query_safe("SELECT '$current_snapshot'::pg_snapshot IS NOT NULL"),
+ 't', 'pg_current_snapshot output is valid pg_snapshot input');
+
$s5->query_safe('COMMIT');
$s4->query_safe('COMMIT');
$s3->query_safe('COMMIT');
--
2.34.1
^ permalink raw reply [nested|flat] 6+ messages in thread
* Re: Snapshot export on a standby corrupts hint bits on subxact overflow
2026-07-27 02:17 Snapshot export on a standby corrupts hint bits on subxact overflow Peter Geoghegan <pg@bowt.ie>
2026-07-28 06:07 ` Re: Snapshot export on a standby corrupts hint bits on subxact overflow Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-07-29 09:36 ` Re: Snapshot export on a standby corrupts hint bits on subxact overflow Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
@ 2026-08-24 23:07 ` Peter Geoghegan <pg@bowt.ie>
2026-08-25 09:32 ` Re: Snapshot export on a standby corrupts hint bits on subxact overflow Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
0 siblings, 1 reply; 6+ messages in thread
From: Peter Geoghegan @ 2026-08-24 23:07 UTC (permalink / raw)
To: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; Andres Freund <andres@anarazel.de>; scott@scottray.io
On Wed, Jul 29, 2026 at 5:36 AM Bertrand Drouvot
<bertranddrouvot.pg@gmail.com> wrote:
> > 1/ In ExportSnapshot(), do not include recovery subxip entries and committed
> > child XIDs at or above xmax when counting and serializing them, so unnecessary
> > entries do not consume the limited recovery subxip capacity.
That is a valid issue, but I wonder if it's worth including in a
back-patchable fix. Is the special case worth the added risk?
Attached v3 simplifies 0001, partly by leaving that part out entirely.
It also simplifies the logic by always writing "sof:%u" and "sxcnt:%d"
to the temp file -- the idea is to make ImportSnapshot import any
subxacts it finds in the file (while still sanitizing the inputs).
The nchildren-won't-fit issue is extremely narrow in practice. The
worst that can happen is that the user sees an "invalid snapshot data
in file..." error on import. But that can only happen when:
1. A snapshot is taken on a standby.
2. The xact holding that snapshot outlives promotion.
3. The xact then writes, acquiring an XID. It accumulates enough
committed subxacts that they, plus the snapshot's own subxip entries,
exceed GetMaxSnapshotSubxidCount().
4. It then calls pg_export_snapshot().
When someone goes to import the snapshot exported in step 4, they see
the error. In practice, GetMaxSnapshotSubxidCount will fit about 15k
XIDs with max_connections=200. So I think that this will just never
happen.
Note also that it need not be the original xact that acquired the
snapshot that holds onto it in step 2. You could export + import a
snapshot during recovery, and then have the importing xact continue
after promotion -- at which point the snapshot is exported + imported
a second time. That's probably manageable in practice, but it makes me
nervous. Especially because there is evidently plenty of potential for
somebody to get something wrong in this area (we learned of three bugs
here in the past 6 weeks or so).
Maybe we could improve the error message, but I want the committed
solution to be as simple as possible.
> > 2/ In pg_current_snapshot(), do not include source XIDs outside [xmin, xmax),
> > so that it enforces the rule regardless of how the source snapshot was produced.
I'm not treating this one as a priority, so I haven't worked on it.
I'm focused on committing 0001 in the next few days, since it's a bug
that has caused users real harm.
--
Peter Geoghegan
Attachments:
[application/octet-stream] v3-0001-Export-subxip-for-snapshots-taken-during-recovery.patch (14.4K, ../../CAH2-WznimsP_QS-bP3xMtRsmMGDud7GDdYAvUU4MYHQiVN_Q1Q@mail.gmail.com/2-v3-0001-Export-subxip-for-snapshots-taken-during-recovery.patch)
download | inline diff:
From 3ebcc94322080611ca73f76dcc9be2a5b477559a Mon Sep 17 00:00:00 2001
From: Peter Geoghegan <pg@bowt.ie>
Date: Thu, 20 Aug 2026 16:14:03 -0400
Subject: [PATCH v3] Export subxip[] for snapshots taken during recovery.
A snapshot taken during recovery stores its whole in-progress set in
subxip, every running top-level XID included, and leaves xip empty. Its
suboverflowed flag therefore does not mean that subxip is redundant. We
nevertheless treated it that way, which allowed queries running on
standbys to see in-progress transactions as aborted. This could lead to
hint bits being incorrectly set on standbys, which could result in wrong
answers for sessions that never imported the snapshot.
To fix, teach snapshot export to include the subxip[] array regardless
of the overflow flag when its snapshot was taken during recovery.
Claude Code found this problem. The committed TAP test is a simplified
version of the one that it wrote to demonstrate this bug.
Author: Peter Geoghegan <pg@bowt.ie>
Author: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
Discussion: https://postgr.es/m/CAH2-WzmHVeYY%3Dpjz9x8DhhxVjXHX0pvoQ-MdiB1Tt6%3Do2GTiKg%40mail.gmail.com
Backpatch-through: 14
---
src/backend/utils/time/snapmgr.c | 42 ++--
src/test/recovery/meson.build | 1 +
.../recovery/t/056_standby_snapshot_export.pl | 222 ++++++++++++++++++
3 files changed, 249 insertions(+), 16 deletions(-)
create mode 100644 src/test/recovery/t/056_standby_snapshot_export.pl
diff --git a/src/backend/utils/time/snapmgr.c b/src/backend/utils/time/snapmgr.c
index 70aed370a..24f3ffb7a 100644
--- a/src/backend/utils/time/snapmgr.c
+++ b/src/backend/utils/time/snapmgr.c
@@ -1117,8 +1117,10 @@ ExportSnapshot(Snapshot snapshot)
TransactionId topXid;
TransactionId *children;
ExportedSnapshot *esnap;
+ int nsubxids;
int nchildren;
int addTopXid;
+ bool suboverflowed;
StringInfoData buf;
FILE *f;
MemoryContext oldcxt;
@@ -1223,16 +1225,29 @@ ExportSnapshot(Snapshot snapshot)
appendStringInfo(&buf, "xip:%u\n", topXid);
/*
- * Similarly, we add our subcommitted child XIDs to the subxid data. Here,
- * we have to cope with possible overflow.
+ * Similarly, we add our subcommitted child XIDs to the subxid data.
+ *
+ * Report overflow when the snapshot overflowed, and also when our subxids
+ * won't fit in what a snapshot can hold; claiming overflow is always
+ * safe, since it just makes importers fall back on pg_subtrans.
*/
- if (snapshot->suboverflowed ||
- snapshot->subxcnt + nchildren > GetMaxSnapshotSubxidCount())
- appendStringInfoString(&buf, "sof:1\n");
- else
+ nsubxids = snapshot->subxcnt + nchildren;
+ suboverflowed = snapshot->suboverflowed ||
+ nsubxids > GetMaxSnapshotSubxidCount();
+
+ /*
+ * Ignore the subxid array if it has overflowed, unless the snapshot was
+ * taken during recovery - in that case, top-level XIDs are in subxip as
+ * well, and we mustn't lose them. CopySnapshot() and SerializeSnapshot()
+ * make the same exception.
+ */
+ if (suboverflowed && !snapshot->takenDuringRecovery)
+ nsubxids = 0;
+
+ appendStringInfo(&buf, "sof:%u\n", suboverflowed);
+ appendStringInfo(&buf, "sxcnt:%d\n", nsubxids);
+ if (nsubxids > 0)
{
- appendStringInfoString(&buf, "sof:0\n");
- appendStringInfo(&buf, "sxcnt:%d\n", snapshot->subxcnt + nchildren);
for (int32 i = 0; i < snapshot->subxcnt; i++)
appendStringInfo(&buf, "sxp:%u\n", snapshot->subxip[i]);
for (int32 i = 0; i < nchildren; i++)
@@ -1493,11 +1508,11 @@ ImportSnapshot(const char *idstr)
snapshot.xip[i] = parseXidFromText("xip:", &filebuf, path);
snapshot.suboverflowed = parseIntFromText("sof:", &filebuf, path);
+ snapshot.subxcnt = xcnt = parseIntFromText("sxcnt:", &filebuf, path);
+ snapshot.subxip = NULL;
- if (!snapshot.suboverflowed)
+ if (snapshot.subxcnt)
{
- snapshot.subxcnt = xcnt = parseIntFromText("sxcnt:", &filebuf, path);
-
/* sanity-check the xid count before palloc */
if (xcnt < 0 || xcnt > GetMaxSnapshotSubxidCount())
ereport(ERROR,
@@ -1508,11 +1523,6 @@ ImportSnapshot(const char *idstr)
for (i = 0; i < xcnt; i++)
snapshot.subxip[i] = parseXidFromText("sxp:", &filebuf, path);
}
- else
- {
- snapshot.subxcnt = 0;
- snapshot.subxip = NULL;
- }
snapshot.takenDuringRecovery = parseIntFromText("rec:", &filebuf, path);
diff --git a/src/test/recovery/meson.build b/src/test/recovery/meson.build
index 39ec8c494..72113c5ac 100644
--- a/src/test/recovery/meson.build
+++ b/src/test/recovery/meson.build
@@ -64,6 +64,7 @@ tests += {
't/053_standby_login_event_trigger.pl',
't/054_unlogged_sequence_promotion.pl',
't/055_cascade_reconnect.pl',
+ 't/056_standby_snapshot_export.pl',
],
},
}
diff --git a/src/test/recovery/t/056_standby_snapshot_export.pl b/src/test/recovery/t/056_standby_snapshot_export.pl
new file mode 100644
index 000000000..d5300336a
--- /dev/null
+++ b/src/test/recovery/t/056_standby_snapshot_export.pl
@@ -0,0 +1,222 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+#
+# Test snapshot export and import on a standby.
+#
+# A snapshot taken during recovery keeps its whole in-progress set, including
+# every running top-level XID, in subxip, and leaves xip empty. Exporting it
+# must write subxip out even once the snapshot is marked suboverflowed, which
+# happens as soon as a transaction on the primary reports 64 subtransactions.
+# A buggy importer that loses subxip believes nothing at all is running between
+# xmin and xmax, treats live transactions as aborted, and stamps
+# HEAP_XMIN_INVALID/HEAP_XMAX_INVALID on their tuples. Every other session
+# on the standby then sees those hint bits.
+
+use strict;
+use warnings FATAL => 'all';
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+# Enough to arm lastOverflowedXid on the standby (which happens at 64) and to
+# overflow the primary's own subxid cache (at 65), so that later
+# xl_running_xacts records keep it armed.
+my $nsubxacts = 80;
+
+my $primary = PostgreSQL::Test::Cluster->new('primary');
+$primary->init(allows_streaming => 1);
+$primary->append_conf('postgresql.conf', 'autovacuum = off');
+$primary->start;
+
+$primary->backup('backup');
+my $standby = PostgreSQL::Test::Cluster->new('standby');
+$standby->init_from_backup($primary, 'backup', has_streaming => 1);
+$standby->append_conf('postgresql.conf', 'max_standby_streaming_delay = -1');
+$standby->start;
+
+# visibility_test holds the rows whose hint bits the test checks. xid_burner
+# exists only to consume XIDs and subxids.
+$primary->safe_psql(
+ 'postgres', q[
+CREATE TABLE visibility_test(k int, pad text);
+INSERT INTO visibility_test
+ SELECT g, repeat('x', 1200) FROM generate_series(1, 40) g;
+CREATE TABLE xid_burner(i int);
+]);
+
+# Hold the primary's removal horizon below the deleting transaction, so that
+# the row version it deletes stays RECENTLY_DEAD. Otherwise the re-INSERT
+# prunes that version on the primary, and replay of the prune record removes
+# the evidence from the standby.
+my $horizon_holder = $primary->background_psql('postgres');
+$horizon_holder->query_safe(
+ 'BEGIN ISOLATION LEVEL REPEATABLE READ;
+ SELECT count(*) FROM visibility_test');
+
+# This transaction deletes a row and stays open, so every snapshot taken from
+# here on must report its XID as running.
+my $toplevel_deleter = $primary->background_psql('postgres');
+$toplevel_deleter->query_safe('BEGIN');
+$toplevel_deleter->query_safe('DELETE FROM visibility_test WHERE k = 7');
+my $deleter_xid =
+ $toplevel_deleter->query_safe('SELECT pg_current_xact_id()');
+
+# This transaction deletes another row in an early subtransaction, then
+# overflows its subxid cache and stays open. Recovery removes the deleting
+# subtransaction's XID from KnownAssignedXids, so checking that tuple's xmax
+# later has to map the child XID back to its parent through pg_subtrans.
+# Keeping the transaction open also keeps lastOverflowedXid armed on the
+# standby.
+my $subxact_deleter = $primary->background_psql('postgres');
+$subxact_deleter->query_safe('BEGIN');
+$subxact_deleter->query_safe('SAVEPOINT early');
+$subxact_deleter->query_safe('DELETE FROM visibility_test WHERE k = 8');
+$subxact_deleter->query_safe('RELEASE early');
+$subxact_deleter->query_safe(
+ qq[DO \$\$ BEGIN
+ FOR i IN 1..$nsubxacts LOOP
+ BEGIN INSERT INTO xid_burner VALUES (i);
+ EXCEPTION WHEN OTHERS THEN NULL; END;
+ END LOOP; END \$\$]);
+
+# Commit something so that the standby's xmax ends up past the deleting XID,
+# and to flush the xid-assignment WAL that those subtransactions wrote. Then
+# wait until the standby has removed the subxids and marks snapshots
+# overflowed. Without the extra commit, XidInMVCCSnapshot() answers through
+# its xid >= xmax range test, and the test would pass without examining the
+# in-progress array.
+$primary->safe_psql('postgres', 'INSERT INTO xid_burner VALUES (0)');
+$primary->wait_for_replay_catchup($standby);
+
+my $exporter = $standby->background_psql('postgres');
+$exporter->query_safe('BEGIN ISOLATION LEVEL REPEATABLE READ');
+my $recovery_snap = $exporter->query_safe('SELECT pg_export_snapshot()');
+
+my $recovery_file =
+ slurp_file($standby->data_dir . "/pg_snapshots/$recovery_snap");
+note("exported snapshot $recovery_snap:\n$recovery_file");
+
+like($recovery_file, qr/^rec:1$/m, 'snapshot was taken during recovery');
+like($recovery_file, qr/^sof:1$/m, 'snapshot is suboverflowed');
+
+# Without these the test could pass while exercising nothing: an xmax below
+# the deleting XID answers "not running" through the plain range test instead.
+my ($exported_xmin) = $recovery_file =~ /^xmin:(\d+)$/m;
+my ($exported_xmax) = $recovery_file =~ /^xmax:(\d+)$/m;
+cmp_ok($exported_xmin, '<=', $deleter_xid,
+ 'exported xmin does not follow the running XID');
+cmp_ok($deleter_xid, '<', $exported_xmax,
+ 'running XID precedes exported xmax');
+
+like($recovery_file, qr/^sxcnt:[1-9]/m,
+ 'suboverflowed recovery snapshot exports its subxip array');
+like($recovery_file, qr/^sxp:$deleter_xid$/m, 'running XID is exported');
+
+# Import the snapshot and read the table. This answers correctly (a
+# still-running transaction and an aborted one are indistinguishable here),
+# but it is what writes the hint bits.
+my $importer = $standby->background_psql('postgres');
+$importer->query_safe(
+ qq[BEGIN ISOLATION LEVEL REPEATABLE READ;
+ SET TRANSACTION SNAPSHOT '$recovery_snap']);
+is($importer->query_safe('SELECT count(*) FROM visibility_test'),
+ 40, 'importing backend sees its own snapshot');
+$importer->query_safe('COMMIT');
+$exporter->query_safe('COMMIT');
+
+# The deleting transaction commits, and the freed key is used again.
+$toplevel_deleter->query_safe('COMMIT');
+$primary->safe_psql('postgres',
+ q[INSERT INTO visibility_test VALUES (7, repeat('y', 1200))]);
+$primary->wait_for_replay_catchup($standby);
+
+# Sessions that never touched the exported snapshot must agree with the
+# primary. A stale HEAP_XMAX_INVALID on the deleted row version brings it
+# back to life, so the key is returned twice.
+is( $standby->safe_psql(
+ 'postgres', 'SELECT count(*) FROM visibility_test WHERE k = 7'),
+ 1,
+ 'standby sees the re-inserted row once');
+
+my $digest =
+ q[SELECT md5(string_agg(k::text, ',' ORDER BY k)) FROM visibility_test];
+is( $standby->safe_psql('postgres', $digest),
+ $primary->safe_psql('postgres', $digest),
+ 'primary and standby agree on table contents');
+
+# pg_snapshots is only emptied at startup, and ImportSnapshot() takes rec:
+# from the file rather than from RecoveryInProgress(). A snapshot exported
+# by a standby therefore outlives promotion, and keeps taking
+# XidInMVCCSnapshot()'s recovery branch on a server no longer in recovery.
+my $promotion_exporter = $standby->background_psql('postgres');
+$promotion_exporter->query_safe('BEGIN ISOLATION LEVEL REPEATABLE READ');
+my $promotion_snap =
+ $promotion_exporter->query_safe('SELECT pg_export_snapshot()');
+my $promotion_count =
+ $promotion_exporter->query_safe('SELECT count(*) FROM visibility_test');
+
+like(slurp_file($standby->data_dir . "/pg_snapshots/$promotion_snap"),
+ qr/^sof:1$/m,
+ 'snapshot saved for post-promotion re-export is suboverflowed');
+
+# The deleting child XID is no longer in KnownAssignedXids, so this needs
+# pg_subtrans to reach its parent.
+$subxact_deleter->query_safe('COMMIT');
+$primary->wait_for_replay_catchup($standby);
+is( $standby->safe_psql(
+ 'postgres', 'SELECT count(*) FROM visibility_test WHERE k = 8'),
+ 0,
+ 'standby sees the committed subtransaction delete');
+
+$standby->promote;
+$standby->poll_query_until('postgres', 'SELECT NOT pg_is_in_recovery()')
+ or die "standby never finished promotion";
+
+my $reexporter = $standby->background_psql('postgres');
+$reexporter->query_safe(
+ qq[BEGIN ISOLATION LEVEL REPEATABLE READ;
+ SET TRANSACTION SNAPSHOT '$promotion_snap']);
+is($reexporter->query_safe('SELECT count(*) FROM visibility_test'),
+ $promotion_count,
+ 'recovery-taken snapshot still imports after promotion');
+
+# That importing transaction is read-write, since the read-only requirement
+# binds only a SERIALIZABLE source. So it can acquire subcommitted children
+# and export again, which hands ExportSnapshot() a snapshot taken during
+# recovery together with a non-empty children array. The subxip array must
+# still be written out.
+$reexporter->query_safe('SAVEPOINT sp');
+$reexporter->query_safe(
+ q[INSERT INTO visibility_test VALUES (5000, repeat('z', 10))]);
+$reexporter->query_safe('RELEASE sp');
+my $reexported_snap = $reexporter->query_safe('SELECT pg_export_snapshot()');
+
+my $reexported_file =
+ slurp_file($standby->data_dir . "/pg_snapshots/$reexported_snap");
+like($reexported_file, qr/^sof:1$/m,
+ 're-exported recovery snapshot remains suboverflowed');
+like($reexported_file, qr/^sxcnt:[1-9]/m,
+ 'a recovery-taken snapshot re-exports its subxip array after a write');
+
+my $reimporter = $standby->background_psql('postgres');
+$reimporter->query_safe(
+ qq[BEGIN ISOLATION LEVEL REPEATABLE READ;
+ SET TRANSACTION SNAPSHOT '$reexported_snap']);
+is($reimporter->query_safe('SELECT count(*) FROM visibility_test'),
+ $promotion_count, 're-exported recovery snapshot can be imported');
+
+$reimporter->query_safe('COMMIT');
+$reexporter->query_safe('COMMIT');
+$promotion_exporter->query_safe('COMMIT');
+
+$horizon_holder->quit;
+$toplevel_deleter->quit;
+$subxact_deleter->quit;
+$exporter->quit;
+$importer->quit;
+$promotion_exporter->quit;
+$reexporter->quit;
+$reimporter->quit;
+$standby->stop;
+$primary->stop;
+
+done_testing();
--
2.53.0
^ permalink raw reply [nested|flat] 6+ messages in thread
* Re: Snapshot export on a standby corrupts hint bits on subxact overflow
2026-07-27 02:17 Snapshot export on a standby corrupts hint bits on subxact overflow Peter Geoghegan <pg@bowt.ie>
2026-07-28 06:07 ` Re: Snapshot export on a standby corrupts hint bits on subxact overflow Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-07-29 09:36 ` Re: Snapshot export on a standby corrupts hint bits on subxact overflow Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-08-24 23:07 ` Re: Snapshot export on a standby corrupts hint bits on subxact overflow Peter Geoghegan <pg@bowt.ie>
@ 2026-08-25 09:32 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-08-26 20:50 ` Re: Snapshot export on a standby corrupts hint bits on subxact overflow Peter Geoghegan <pg@bowt.ie>
0 siblings, 1 reply; 6+ messages in thread
From: Bertrand Drouvot @ 2026-08-25 09:32 UTC (permalink / raw)
To: Peter Geoghegan <pg@bowt.ie>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; Andres Freund <andres@anarazel.de>; scott@scottray.io
Hi,
On Mon, Aug 24, 2026 at 07:07:07PM -0400, Peter Geoghegan wrote:
> On Wed, Jul 29, 2026 at 5:36 AM Bertrand Drouvot
> <bertranddrouvot.pg@gmail.com> wrote:
> > > 1/ In ExportSnapshot(), do not include recovery subxip entries and committed
> > > child XIDs at or above xmax when counting and serializing them, so unnecessary
> > > entries do not consume the limited recovery subxip capacity.
>
> That is a valid issue, but I wonder if it's worth including in a
> back-patchable fix. Is the special case worth the added risk?
Yeah, probably not. What about adding an XXX here:
+ /*
+ * Ignore the subxid array if it has overflowed, unless the snapshot was
+ * taken during recovery - in that case, top-level XIDs are in subxip as
+ * well, and we mustn't lose them. CopySnapshot() and SerializeSnapshot()
+ * make the same exception.
+ */
Like:
"
* XXX: After promotion, an imported recovery snapshot can have subxip
* entries and committed children at or above xmax. These entries cannot
* affect visibility, but can make sxcnt exceed
* GetMaxSnapshotSubxidCount(), causing ImportSnapshot() to reject a
* snapshot we exported. Filtering entries outside [xmin, xmax) would avoid that.
"
so that we don't forget about it?
> Attached v3 simplifies 0001, partly by leaving that part out entirely.
Thanks for the new version! Yeah, it looks simpler, let's keep it that way.
> It also simplifies the logic by always writing "sof:%u" and "sxcnt:%d"
> to the temp file -- the idea is to make ImportSnapshot import any
> subxacts it finds in the file (while still sanitizing the inputs).
Good idea! That makes sense to me. The format change is safe to backpatch too,
since exported snapshot files are removed at startup.
> Maybe we could improve the error message, but I want the committed
> solution to be as simple as possible.
That makes sense. I'm not sure we should modify the error message in this
commit, let's keep the patch focus on fixing the bug?
> > > 2/ In pg_current_snapshot(), do not include source XIDs outside [xmin, xmax),
> > > so that it enforces the rule regardless of how the source snapshot was produced.
>
> I'm not treating this one as a priority, so I haven't worked on it.
>
> I'm focused on committing 0001 in the next few days, since it's a bug
> that has caused users real harm.
Sounds good!
I looked at v3 and LGTM.
Regards,
--
Bertrand Drouvot
PostgreSQL Contributors Team
RDS Open Source Databases
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 6+ messages in thread
* Re: Snapshot export on a standby corrupts hint bits on subxact overflow
2026-07-27 02:17 Snapshot export on a standby corrupts hint bits on subxact overflow Peter Geoghegan <pg@bowt.ie>
2026-07-28 06:07 ` Re: Snapshot export on a standby corrupts hint bits on subxact overflow Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-07-29 09:36 ` Re: Snapshot export on a standby corrupts hint bits on subxact overflow Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-08-24 23:07 ` Re: Snapshot export on a standby corrupts hint bits on subxact overflow Peter Geoghegan <pg@bowt.ie>
2026-08-25 09:32 ` Re: Snapshot export on a standby corrupts hint bits on subxact overflow Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
@ 2026-08-26 20:50 ` Peter Geoghegan <pg@bowt.ie>
0 siblings, 0 replies; 6+ messages in thread
From: Peter Geoghegan @ 2026-08-26 20:50 UTC (permalink / raw)
To: Bertrand Drouvot <bertranddrouvot.pg@gmail.com>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; Andres Freund <andres@anarazel.de>; scott@scottray.io
Pushed.
On Tue, Aug 25, 2026 at 5:32 AM Bertrand Drouvot
<bertranddrouvot.pg@gmail.com> wrote:
> On Mon, Aug 24, 2026 at 07:07:07PM -0400, Peter Geoghegan wrote:
> > On Wed, Jul 29, 2026 at 5:36 AM Bertrand Drouvot
> > <bertranddrouvot.pg@gmail.com> wrote:
> > > > 1/ In ExportSnapshot(), do not include recovery subxip entries and committed
> > > > child XIDs at or above xmax when counting and serializing them, so unnecessary
> > > > entries do not consume the limited recovery subxip capacity.
> >
> > That is a valid issue, but I wonder if it's worth including in a
> > back-patchable fix. Is the special case worth the added risk?
>
> Yeah, probably not. What about adding an XXX here:
In the committed version, we raise a specific ERROR when this happens.
I think it's unlikely that any user will ever see this ERROR, but it's
better to have it and not need it.
Thanks
--
Peter Geoghegan
^ permalink raw reply [nested|flat] 6+ messages in thread
end of thread, other threads:[~2026-08-26 20:50 UTC | newest]
Thread overview: 6+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2026-07-27 02:17 Snapshot export on a standby corrupts hint bits on subxact overflow Peter Geoghegan <pg@bowt.ie>
2026-07-28 06:07 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-07-29 09:36 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-08-24 23:07 ` Peter Geoghegan <pg@bowt.ie>
2026-08-25 09:32 ` Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-08-26 20:50 ` Peter Geoghegan <pg@bowt.ie>
This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox