pg.ddx.io pgsql-hackers@postgresql.org mailing list archive
help / color / mirror / Atom feed From: Alexander Lakhin <exclusion@gmail.com>
To: Michael Paquier <michael@paquier.xyz>
To: Hayato Kuroda (Fujitsu) <kuroda.hayato@fujitsu.com>
Cc: pgsql-hackers <pgsql-hackers@postgresql.org>
Cc: Aleksander Alekseev <aleksander@timescale.com>
Subject: Re: BUG: Former primary node might stuck when started as a standby
Date: Wed, 4 Mar 2026 10:00:00 +0200
Message-ID: <9d4abbe2-95aa-47e0-9ce2-842196a662a7@gmail.com> (raw )
In-Reply-To: <aafDsb5snkfkNfdS@paquier.xyz >
References: <OS9PR01MB121498EFA4CBF3003B83C9BCCF56CA@OS9PR01MB12149.jpnprd01.prod.outlook.com >
<63e55743-669e-4300-a561-7b7ff63723b6@gmail.com >
<OS9PR01MB12149D4F1A2BC23637688CE4DF56BA@OS9PR01MB12149.jpnprd01.prod.outlook.com >
<045cab6f-4738-417e-b551-01adba44d6c3@gmail.com >
<c7284102-01fb-4704-8a80-8e65f59fa933@gmail.com >
<aaU5EiNHfyMb9Bvu@paquier.xyz >
<TYRPR01MB12156CC5A9AC774B07FA7E40BF57FA@TYRPR01MB12156.jpnprd01.prod.outlook.com >
<aaZ77VvZ4Oabp30A@paquier.xyz >
<493401a8-063f-436a-8287-a235d9e065fc@gmail.com >
<OS9PR01MB121491FE3BD5D9D7341B83537F57FA@OS9PR01MB12149.jpnprd01.prod.outlook.com >
<aafDsb5snkfkNfdS@paquier.xyz >
Hello Michael,
04.03.2026 07:31, Michael Paquier wrote
>> I guess so. cluster::stop does the `pg_ctl stop -m fast` command. In this case
>> the walsender waits till there are nothing to be sent, see WalSndLoop().
>> Do let me know if you have observed the similar failure here.
> Exactly. Doing a clean stop of the primary offers a strong guarantee
> here. We are sure that the standby will have received all the records
> from the primary. Timeline forking is an impossible thing in
> 012_subtransactions.pl based on how the switchover from the primary to
> the standby happens. I don't see a need for tweaking this test at
> all. Or perhaps you did see a failure of some kind in this test,
> Alexander?
Yes, 012_subtransactions doesn't fail with aggressive bgwriter, as I noted
before. I mentioned it exactly to show that stop does matter here. But if
we recognize teardown_node in this context as risky, maybe it would make
sense to review also other tests in recovery/. I already wrote about
004_timeline_switch, but probably there are more. E.g., 028_pitr_timelines
(I haven't tested it intensively yet) does:
$node_primary->stop('immediate');
# Promote the standby, and switch WAL so that it archives a WAL segment
# that contains all the INSERTs, on a new timeline.
$node_standby->promote;
Best regards,
Alexander
view thread (30+ messages) latest in thread
Message-ID: <9d4abbe2-95aa-47e0-9ce2-842196a662a7@gmail.com>
Permalink: ../9d4abbe2-95aa-47e0-9ce2-842196a662a7@gmail.com/
Also on: postgresql.org/message-id/9d4abbe2-95aa-47e0-9ce2-842196a662a7@gmail.com
copy link · copy postgr.es
reply Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Reply to all the recipients using the --to and --cc options:
reply via email
To: pgsql-hackers@postgresql.org
Cc: exclusion@gmail.com, michael@paquier.xyz, kuroda.hayato@fujitsu.com, aleksander@timescale.com
Subject: Re: BUG: Former primary node might stuck when started as a standby
In-Reply-To: <9d4abbe2-95aa-47e0-9ce2-842196a662a7@gmail.com>
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox