agora inbox for pgsql-hackers@postgresql.org  
help / color / mirror / Atom feed
From: Andres Freund <andres@anarazel.de>
To: Xuneng Zhou <xunengzhou@gmail.com>
Cc: Alexander Korotkov <aekorotkov@gmail.com>
Cc: Tom Lane <tgl@sss.pgh.pa.us>
Cc: Heikki Linnakangas <hlinnaka@iki.fi>
Cc: Peter Eisentraut <peter@eisentraut.org>
Cc: Thomas Munro <thomas.munro@gmail.com>
Cc: Álvaro Herrera <alvherre@kurilemu.de>
Cc: Chao Li <li.evan.chao@gmail.com>
Cc: pgsql-hackers <pgsql-hackers@lists.postgresql.org>
Cc: Michael Paquier <michael@paquier.xyz>
Cc: jian he <jian.universality@gmail.com>
Cc: Tomas Vondra <tomas@vondra.me>
Cc: Yura Sokolov <y.sokolov@postgrespro.ru>
Subject: Re: Implement waiting for wal lsn replay: reloaded
Date: Tue, 7 Apr 2026 09:18:37 -0400
Message-ID: <ytjxycbqtqcdw6o57ljfmerzcudhxcymlmqspoqubjgxep5t4g@nzilrsos3e3v> (raw)
In-Reply-To: <CABPTF7VRP97GPwPTiu89xYQMA5pfWknsLSSxnpq11mu+-FiDRA@mail.gmail.com>
References: <CABPTF7V-E_e3kQ2vtwUz6Jy7u-8_YeUT0SDoAbu7EKPgNp=ndA@mail.gmail.com>
	<CAPpHfdtNiSqQCu+YtTYcc+K4q9FwtZuAtQ5Qs+KoaZZM4QyYTA@mail.gmail.com>
	<1957514.1775526774@sss.pgh.pa.us>
	<1959506.1775527693@sss.pgh.pa.us>
	<CABPTF7W0GGEpTeRS8YiMX=77DJbe9jqoUEaWpWHcxxs1xvWkkA@mail.gmail.com>
	<zqbppucpmkeqecfy4s5kscnru4tbk6khp3ozqz6ad2zijz354k@w4bdf4z3wqoz>
	<jzq5shdewncpxc35r3s2mcfsmo4bjovkza5mnqf5bdfumhfi3g@bglckf7dxmw5>
	<CABPTF7WPRVJGdDeuWo-=3csnDs5FGfUsg97xppPyzoj9fRAMeQ@mail.gmail.com>
	<CAPpHfds=B=ZNcfxPqFJ8ZVJB6surey+cP+8iDLfj5qM3Xvx=bg@mail.gmail.com>
	<CABPTF7VRP97GPwPTiu89xYQMA5pfWknsLSSxnpq11mu+-FiDRA@mail.gmail.com>

Hi,

On 2026-04-07 21:05:40 +0800, Xuneng Zhou wrote:
> I’ve posted two patches. The first fixes the duplication issue
> reported by Andres and is fairly straightforward. The second turned
> out to be more complex than expected, and I’m still working through
> possible solutions. Feedback or alternative approaches would be very
> helpful.
> I also spent some time drafting a patch to address the memory ordering
> issue and will post it later.

I propose quickly applying a minimal patch like the attached, to get the test
performance back to normal.

Will do so unless somebody protests within in one CI cycle and one coffee.

Greetings,

Andres Freund

Attachments:

  [text/x-diff] v1-0001-Minimal-fix-for-WAIT-FOR-.-MODE-standby_flush.patch (2.0K, ../ytjxycbqtqcdw6o57ljfmerzcudhxcymlmqspoqubjgxep5t4g@nzilrsos3e3v/2-v1-0001-Minimal-fix-for-WAIT-FOR-.-MODE-standby_flush.patch)
  download | inline diff:
From 0a9c10fe36c6b2d08d1f4fbd0825b76bdd389c10 Mon Sep 17 00:00:00 2001
From: Andres Freund <andres@anarazel.de>
Date: Tue, 7 Apr 2026 09:11:07 -0400
Subject: [PATCH v1] Minimal fix for WAIT FOR ... MODE 'standby_flush'

The investigation into the negative test performance impact of 7e8aeb9e483
lead to discovering that there are a few issues with WAIT FOR.

This commit is just a minimal fix to prevent hangs in standby_flush mode, due
to WAIT FOR ... 'standby_flush' seeing a 0 LSN if a newly started walreceiver
does not receive any writes, because the stanby is already caught up.

There are several other issues and this is isn't necessarily the best fix. But
this way we get the hangs out of the way.

Reported-by: Tom Lane <tgl@sss.pgh.pa.us>
Discussion: https://postgr.es/m/zqbppucpmkeqecfy4s5kscnru4tbk6khp3ozqz6ad2zijz354k@w4bdf4z3wqoz
---
 src/backend/replication/walreceiver.c      | 2 --
 src/backend/replication/walreceiverfuncs.c | 1 +
 2 files changed, 1 insertion(+), 2 deletions(-)

diff --git a/src/backend/replication/walreceiver.c b/src/backend/replication/walreceiver.c
index a437273cf9a..09fde92bfd7 100644
--- a/src/backend/replication/walreceiver.c
+++ b/src/backend/replication/walreceiver.c
@@ -242,8 +242,6 @@ WalReceiverMain(const void *startup_data, size_t startup_data_len)
 
 	SpinLockRelease(&walrcv->mutex);
 
-	pg_atomic_write_u64(&WalRcv->writtenUpto, 0);
-
 	/* Arrange to clean up at walreceiver exit */
 	on_shmem_exit(WalRcvDie, PointerGetDatum(&startpointTLI));
 
diff --git a/src/backend/replication/walreceiverfuncs.c b/src/backend/replication/walreceiverfuncs.c
index 4e03e721872..bd5d47be964 100644
--- a/src/backend/replication/walreceiverfuncs.c
+++ b/src/backend/replication/walreceiverfuncs.c
@@ -321,6 +321,7 @@ RequestXLogStreaming(TimeLineID tli, XLogRecPtr recptr, const char *conninfo,
 		walrcv->flushedUpto = recptr;
 		walrcv->receivedTLI = tli;
 		walrcv->latestChunkStart = recptr;
+		pg_atomic_write_u64(&walrcv->writtenUpto, recptr);
 	}
 	walrcv->receiveStart = recptr;
 	walrcv->receiveStartTLI = tli;
-- 
2.53.0.1.gb2826b52eb

view thread (181+ messages)  latest in thread

Message-ID: <ytjxycbqtqcdw6o57ljfmerzcudhxcymlmqspoqubjgxep5t4g@nzilrsos3e3v>
Permalink:  ../ytjxycbqtqcdw6o57ljfmerzcudhxcymlmqspoqubjgxep5t4g@nzilrsos3e3v/
Also on:    postgresql.org/message-id/ytjxycbqtqcdw6o57ljfmerzcudhxcymlmqspoqubjgxep5t4g@nzilrsos3e3v

reply

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Reply to all the recipients using the --to and --cc options:
  reply via email

  To: pgsql-hackers@postgresql.org
  Cc: andres@anarazel.de, xunengzhou@gmail.com, aekorotkov@gmail.com, tgl@sss.pgh.pa.us, hlinnaka@iki.fi, peter@eisentraut.org, thomas.munro@gmail.com, alvherre@kurilemu.de, li.evan.chao@gmail.com, pgsql-hackers@lists.postgresql.org, michael@paquier.xyz, jian.universality@gmail.com, tomas@vondra.me, y.sokolov@postgrespro.ru
  Subject: Re: Implement waiting for wal lsn replay: reloaded
  In-Reply-To: <ytjxycbqtqcdw6o57ljfmerzcudhxcymlmqspoqubjgxep5t4g@nzilrsos3e3v>

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox