pg.ddx.io  pgsql-hackers@postgresql.org mailing list archive  
help / color / mirror / Atom feed
From: Michael Paquier <michael@paquier.xyz>
To: Hayato Kuroda (Fujitsu) <kuroda.hayato@fujitsu.com>
Cc: Alexander Lakhin <exclusion@gmail.com>
Cc: pgsql-hackers <pgsql-hackers@postgresql.org>
Cc: Aleksander Alekseev <aleksander@timescale.com>
Subject: Re: BUG: Former primary node might stuck when started as a standby
Date: Wed, 4 Mar 2026 14:31:29 +0900
Message-ID: <aafDsb5snkfkNfdS@paquier.xyz> (raw)
In-Reply-To: <OS9PR01MB121491FE3BD5D9D7341B83537F57FA@OS9PR01MB12149.jpnprd01.prod.outlook.com>
References: <OS9PR01MB121498EFA4CBF3003B83C9BCCF56CA@OS9PR01MB12149.jpnprd01.prod.outlook.com>
	<63e55743-669e-4300-a561-7b7ff63723b6@gmail.com>
	<OS9PR01MB12149D4F1A2BC23637688CE4DF56BA@OS9PR01MB12149.jpnprd01.prod.outlook.com>
	<045cab6f-4738-417e-b551-01adba44d6c3@gmail.com>
	<c7284102-01fb-4704-8a80-8e65f59fa933@gmail.com>
	<aaU5EiNHfyMb9Bvu@paquier.xyz>
	<TYRPR01MB12156CC5A9AC774B07FA7E40BF57FA@TYRPR01MB12156.jpnprd01.prod.outlook.com>
	<aaZ77VvZ4Oabp30A@paquier.xyz>
	<493401a8-063f-436a-8287-a235d9e065fc@gmail.com>
	<OS9PR01MB121491FE3BD5D9D7341B83537F57FA@OS9PR01MB12149.jpnprd01.prod.outlook.com>

On Tue, Mar 03, 2026 at 09:17:16AM +0000, Hayato Kuroda (Fujitsu) wrote:
> Thanks for the info. So I can provide the patch after the issue for 009_twophase.pl
> is fixed. For better understanding we may be able to fork new
> thread.

Regarding your posted v4, I am actually not convinced that there is a
need for injection points and disabling standby snapshots, for the
three sequences of tests proposed.

While the first wait_for_replay_catchup() can be useful before the
teardown_node() of the primary in the "Check that prepared
transactions can be committed on promoted standby" sequence, it still
has a limited impact.  It looks like we could have other parasite
records as well, depending on how slowly the primary is stopped?  I
think that we should switch to a plain stop() of the primary, the test
wants to check that prepared transactions can be committed on a
standby.  Stopping the primary abruptly does not matter for this
sequence.

For the second wait_for_replay_catchup(), after the PREPARE of
xact_009_11.  I may be missing something but in how does it change
things?  A plain stop() of the primary means that it would have
received all the WAL records from the primary on disk in its pg_wal,
no?  Upon restart, it should replay everything it finds in pg_wal/.  I
don't see a change required here.

For the third wait_for_replay_catchup(), after the PREPARE of
xact_009_12, same dance.  The primary is cleanly stopped first.  All
the WAL records of the primary should have been flushed to the
standby.

As a whole, it looks like we should just switch the teardown() call to
a stop() call in the first test with xact_009_10, backpatch it, and
call it a day.  No need for injection points and no need for GUC
tweaks.  I have not looked at 004_timeline_switch yet.

> I guess so. cluster::stop does the `pg_ctl stop -m fast` command. In this case
> the walsender waits till there are nothing to be sent, see WalSndLoop().
> Do let me know if you have observed the similar failure here.

Exactly.  Doing a clean stop of the primary offers a strong guarantee
here.  We are sure that the standby will have received all the records
from the primary.  Timeline forking is an impossible thing in
012_subtransactions.pl based on how the switchover from the primary to
the standby happens.  I don't see a need for tweaking this test at
all.  Or perhaps you did see a failure of some kind in this test,
Alexander?
--
Michael

Attachments:

  [application/pgp-signature] signature.asc (832B, ../aafDsb5snkfkNfdS@paquier.xyz/2-signature.asc)
  download

view thread (30+ messages)  latest in thread

Message-ID: <aafDsb5snkfkNfdS@paquier.xyz>
Permalink:  ../aafDsb5snkfkNfdS@paquier.xyz/
Also on:    postgresql.org/message-id/aafDsb5snkfkNfdS@paquier.xyz

 · 

reply

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Reply to all the recipients using the --to and --cc options:
  reply via email

  To: pgsql-hackers@postgresql.org
  Cc: michael@paquier.xyz, kuroda.hayato@fujitsu.com, exclusion@gmail.com, aleksander@timescale.com
  Subject: Re: BUG: Former primary node might stuck when started as a standby
  In-Reply-To: <aafDsb5snkfkNfdS@paquier.xyz>

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox