pg.ddx.io  pgsql-hackers@postgresql.org mailing list archive  
help / color / mirror / Atom feed
injection_points: canceled or terminated waiters leak their wait slots
21+ messages / 5 participants
[nested] [flat]

* injection_points: canceled or terminated waiters leak their wait slots
@ 2026-07-21 08:28  Zsolt Parragi <zsolt.parragi@percona.com>
  0 siblings, 1 reply; 21+ messages in thread

From: Zsolt Parragi @ 2026-07-21 08:28 UTC (permalink / raw)
  To: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>

Hello hackers,

While stress-testing REPACK CONCURRENTLY on 19beta2 I saw a logical
decoding activation race, and I extended 051_effective_wal_level.pl
with a test that waits on an injection point and wakes it up later.
051 cancels two injection point waiters earlier in the script, and
injection_wait() never cleans up after a canceled waiter, the wakeup
never arrived and the test deadlocked.

I think we are missing an ENSURE_ERROR_CLEANUP block there. See
attached patch with a testcase reproducing the issue.

A wakeup racing against a canceled waiter with no other live waiter
now errors with "could not find injection point ... to wake up"
instead of silently bumping the leaked slot.

I also attached a separate version for pg19, as master has a
refactored version of injection_wait. All previous branches have the
19 version, it should be easy to backport to other branches.

Attachments:

  [application/octet-stream] nocfbot-pg19-0001-injection_points-clear-waiter-slot-on-error-and-exit.patch (9.2K, ../../CAN4CZFO+KF=cc0-iEg28RhqRBp_fTs6D4b8b7D7DB-pGYP3Ccg@mail.gmail.com/2-nocfbot-pg19-0001-injection_points-clear-waiter-slot-on-error-and-exit.patch)
  download | inline diff:
From 08e1d5b36086f4646df8a203dfc21ceff2a9afa7 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Mon, 20 Jul 2026 21:22:09 +0000
Subject: [PATCH] injection_points: clear waiter slot on error and exit

injection_wait() only clears its slot in the waiter array after the
wait loop finishes. When the waiting query is canceled or the backend
is terminated, the slot leaks. Later wakeups of the same point then
bump the counter of the leaked slot instead of the real waiter, which
sleeps forever. Repeated leaks also exhaust the 8 slots.

Wrap the wait loop in PG_ENSURE_ERROR_CLEANUP, its cleanup callback
covers both ERROR and FATAL.

Add an isolation test: cancel one waiter, terminate another, then
check that a later waiter still receives the wakeup.
---
 src/test/modules/injection_points/Makefile    |  1 +
 .../expected/wait_cleanup.out                 | 87 +++++++++++++++++++
 .../injection_points/injection_points.c       | 45 +++++++---
 src/test/modules/injection_points/meson.build |  1 +
 .../injection_points/specs/wait_cleanup.spec  | 50 +++++++++++
 5 files changed, 172 insertions(+), 12 deletions(-)
 create mode 100644 src/test/modules/injection_points/expected/wait_cleanup.out
 create mode 100644 src/test/modules/injection_points/specs/wait_cleanup.spec

diff --git a/src/test/modules/injection_points/Makefile b/src/test/modules/injection_points/Makefile
index c01d2fb095c..2f77c974e0a 100644
--- a/src/test/modules/injection_points/Makefile
+++ b/src/test/modules/injection_points/Makefile
@@ -13,6 +13,7 @@ REGRESS = injection_points hashagg reindex_conc vacuum
 REGRESS_OPTS = --dlpath=$(top_builddir)/src/test/regress
 
 ISOLATION = basic \
+	    wait_cleanup \
 	    inplace \
 	    repack \
 	    repack_temporal \
diff --git a/src/test/modules/injection_points/expected/wait_cleanup.out b/src/test/modules/injection_points/expected/wait_cleanup.out
new file mode 100644
index 00000000000..c5be17428fc
--- /dev/null
+++ b/src/test/modules/injection_points/expected/wait_cleanup.out
@@ -0,0 +1,87 @@
+Parsed test spec with 3 sessions
+
+starting permutation: wait1 cancel3 noop3 wait2 wakeup3 noop2 detach3
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step wait1: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step cancel3: 
+	SELECT pg_cancel_backend(pid) FROM pg_stat_activity
+	  WHERE wait_event = 'injection-points-wait';
+ <waiting ...>
+step wait1: <... completed>
+ERROR:  canceling statement due to user request
+step cancel3: <... completed>
+pg_cancel_backend
+-----------------
+t                
+(1 row)
+
+step noop3: 
+step wait2: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step wakeup3: SELECT injection_points_wakeup('injection-points-wait');
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait2: <... completed>
+injection_points_run
+--------------------
+                    
+(1 row)
+
+step noop2: 
+step detach3: SELECT injection_points_detach('injection-points-wait');
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
+
+starting permutation: wait1 terminate3 noop3 wait2 wakeup3 noop2 detach3
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step wait1: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step terminate3: 
+	SELECT pg_terminate_backend(pid) FROM pg_stat_activity
+	  WHERE wait_event = 'injection-points-wait';
+ <waiting ...>
+step wait1: <... completed>
+FATAL:  terminating connection due to administrator command
+server closed the connection unexpectedly
+	This probably means the server terminated abnormally
+	before or while processing the request.
+
+step terminate3: <... completed>
+pg_terminate_backend
+--------------------
+t                   
+(1 row)
+
+step noop3: 
+step wait2: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step wakeup3: SELECT injection_points_wakeup('injection-points-wait');
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait2: <... completed>
+injection_points_run
+--------------------
+                    
+(1 row)
+
+step noop2: 
+step detach3: SELECT injection_points_detach('injection-points-wait');
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/injection_points.c b/src/test/modules/injection_points/injection_points.c
index ba282e3dcab..418c5bc756a 100644
--- a/src/test/modules/injection_points/injection_points.c
+++ b/src/test/modules/injection_points/injection_points.c
@@ -222,6 +222,25 @@ injection_notice(const char *name, const void *private_data, void *arg)
 		elog(NOTICE, "notice triggered for injection point %s", name);
 }
 
+/*
+ * Clear the shared waiter slot of this backend, on error or process exit.
+ *
+ * Without this a canceled or terminated waiter would leak its slot, and
+ * later wakeups of the same injection point would bump the leaked slot's
+ * counter instead of the real waiter's. Registered through
+ * PG_ENSURE_ERROR_CLEANUP so that it also runs on FATAL, which exits
+ * without unwinding the stack.
+ */
+static void
+injection_wait_cleanup(int code, Datum arg)
+{
+	int			index = DatumGetInt32(arg);
+
+	SpinLockAcquire(&inj_state->lock);
+	inj_state->name[index][0] = '\0';
+	SpinLockRelease(&inj_state->lock);
+}
+
 /* Wait on a condition variable, awaken by injection_points_wakeup() */
 void
 injection_wait(const char *name, const void *private_data, void *arg)
@@ -266,24 +285,26 @@ injection_wait(const char *name, const void *private_data, void *arg)
 
 	/* And sleep.. */
 	ConditionVariablePrepareToSleep(&inj_state->wait_point);
-	for (;;)
+	PG_ENSURE_ERROR_CLEANUP(injection_wait_cleanup, Int32GetDatum(index));
 	{
-		uint32		new_wait_counts;
+		for (;;)
+		{
+			uint32		new_wait_counts;
 
-		SpinLockAcquire(&inj_state->lock);
-		new_wait_counts = inj_state->wait_counts[index];
-		SpinLockRelease(&inj_state->lock);
+			SpinLockAcquire(&inj_state->lock);
+			new_wait_counts = inj_state->wait_counts[index];
+			SpinLockRelease(&inj_state->lock);
 
-		if (old_wait_counts != new_wait_counts)
-			break;
-		ConditionVariableSleep(&inj_state->wait_point, injection_wait_event);
+			if (old_wait_counts != new_wait_counts)
+				break;
+			ConditionVariableSleep(&inj_state->wait_point, injection_wait_event);
+		}
+		ConditionVariableCancelSleep();
 	}
-	ConditionVariableCancelSleep();
+	PG_END_ENSURE_ERROR_CLEANUP(injection_wait_cleanup, Int32GetDatum(index));
 
 	/* Remove this injection point from the waiters. */
-	SpinLockAcquire(&inj_state->lock);
-	inj_state->name[index][0] = '\0';
-	SpinLockRelease(&inj_state->lock);
+	injection_wait_cleanup(0, Int32GetDatum(index));
 }
 
 /*
diff --git a/src/test/modules/injection_points/meson.build b/src/test/modules/injection_points/meson.build
index 59dba1cb023..a903bc94a88 100644
--- a/src/test/modules/injection_points/meson.build
+++ b/src/test/modules/injection_points/meson.build
@@ -44,6 +44,7 @@ tests += {
   'isolation': {
     'specs': [
       'basic',
+      'wait_cleanup',
       'inplace',
       'repack',
       'repack_temporal',
diff --git a/src/test/modules/injection_points/specs/wait_cleanup.spec b/src/test/modules/injection_points/specs/wait_cleanup.spec
new file mode 100644
index 00000000000..4ef8c293e83
--- /dev/null
+++ b/src/test/modules/injection_points/specs/wait_cleanup.spec
@@ -0,0 +1,50 @@
+# Check that a canceled or terminated waiter does not leave a stale slot
+# behind in the waiter array. A leaked slot would make later wakeups of
+# the same injection point bump the leaked slot's counter instead of the
+# real waiter's, leaving the real waiter stuck.
+
+setup
+{
+	CREATE EXTENSION injection_points;
+}
+teardown
+{
+	DROP EXTENSION injection_points;
+}
+
+# The first waiter, which gets canceled or terminated. No set_local:
+# the injection point must survive s1's termination so that s3 can
+# still detach it.
+session s1
+setup	{
+	SELECT injection_points_attach('injection-points-wait', 'wait');
+}
+step wait1	{ SELECT injection_points_run('injection-points-wait'); }
+
+# The second waiter, which must still receive the wakeup.
+session s2
+step wait2	{ SELECT injection_points_run('injection-points-wait'); }
+step noop2	{ }
+
+# The control session. The blocker annotations on cancel3/terminate3
+# plus noop3 make the tester wait until wait1 actually finished before
+# starting wait2, otherwise wait2 could register its own waiter slot
+# while s1 still holds the old one.
+session s3
+step cancel3	{
+	SELECT pg_cancel_backend(pid) FROM pg_stat_activity
+	  WHERE wait_event = 'injection-points-wait';
+}
+step terminate3	{
+	SELECT pg_terminate_backend(pid) FROM pg_stat_activity
+	  WHERE wait_event = 'injection-points-wait';
+}
+step wakeup3	{ SELECT injection_points_wakeup('injection-points-wait'); }
+step detach3	{ SELECT injection_points_detach('injection-points-wait'); }
+step noop3	{ }
+
+permutation wait1 cancel3(wait1) noop3 wait2 wakeup3 noop2 detach3
+
+# The terminate permutation has to stay last: s1's connection is dead
+# afterwards, and the tester never reconnects a session.
+permutation wait1 terminate3(wait1) noop3 wait2 wakeup3 noop2 detach3
-- 
2.54.0



  [application/octet-stream] 0001-injection_points-clear-waiter-slot-on-error-and-exit.patch (8.9K, ../../CAN4CZFO+KF=cc0-iEg28RhqRBp_fTs6D4b8b7D7DB-pGYP3Ccg@mail.gmail.com/3-0001-injection_points-clear-waiter-slot-on-error-and-exit.patch)
  download | inline diff:
From 42df9177b36a236b0a10cc1c1352b756228ad412 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Mon, 20 Jul 2026 22:35:06 +0000
Subject: [PATCH] injection_points: clear waiter slot on error and exit

injection_wait() only clears its slot in the waiter array after the
wait loop finishes. When the waiting query is canceled or the backend
is terminated, the slot leaks. Later wakeups of the same point then
bump the counter of the leaked slot instead of the real waiter, which
sleeps forever. Repeated leaks also exhaust the 8 slots.

Wrap the wait loop in PG_ENSURE_ERROR_CLEANUP, its cleanup callback
covers both ERROR and FATAL.

Add an isolation test: cancel one waiter, terminate another, then
check that a later waiter still receives the wakeup.
---
 src/test/modules/injection_points/Makefile    |  1 +
 .../expected/wait_cleanup.out                 | 87 +++++++++++++++++++
 .../injection_points/injection_points.c       | 37 ++++++--
 src/test/modules/injection_points/meson.build |  1 +
 .../injection_points/specs/wait_cleanup.spec  | 50 +++++++++++
 5 files changed, 168 insertions(+), 8 deletions(-)
 create mode 100644 src/test/modules/injection_points/expected/wait_cleanup.out
 create mode 100644 src/test/modules/injection_points/specs/wait_cleanup.spec

diff --git a/src/test/modules/injection_points/Makefile b/src/test/modules/injection_points/Makefile
index c01d2fb095c..2f77c974e0a 100644
--- a/src/test/modules/injection_points/Makefile
+++ b/src/test/modules/injection_points/Makefile
@@ -13,6 +13,7 @@ REGRESS = injection_points hashagg reindex_conc vacuum
 REGRESS_OPTS = --dlpath=$(top_builddir)/src/test/regress
 
 ISOLATION = basic \
+	    wait_cleanup \
 	    inplace \
 	    repack \
 	    repack_temporal \
diff --git a/src/test/modules/injection_points/expected/wait_cleanup.out b/src/test/modules/injection_points/expected/wait_cleanup.out
new file mode 100644
index 00000000000..c5be17428fc
--- /dev/null
+++ b/src/test/modules/injection_points/expected/wait_cleanup.out
@@ -0,0 +1,87 @@
+Parsed test spec with 3 sessions
+
+starting permutation: wait1 cancel3 noop3 wait2 wakeup3 noop2 detach3
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step wait1: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step cancel3: 
+	SELECT pg_cancel_backend(pid) FROM pg_stat_activity
+	  WHERE wait_event = 'injection-points-wait';
+ <waiting ...>
+step wait1: <... completed>
+ERROR:  canceling statement due to user request
+step cancel3: <... completed>
+pg_cancel_backend
+-----------------
+t                
+(1 row)
+
+step noop3: 
+step wait2: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step wakeup3: SELECT injection_points_wakeup('injection-points-wait');
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait2: <... completed>
+injection_points_run
+--------------------
+                    
+(1 row)
+
+step noop2: 
+step detach3: SELECT injection_points_detach('injection-points-wait');
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
+
+starting permutation: wait1 terminate3 noop3 wait2 wakeup3 noop2 detach3
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step wait1: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step terminate3: 
+	SELECT pg_terminate_backend(pid) FROM pg_stat_activity
+	  WHERE wait_event = 'injection-points-wait';
+ <waiting ...>
+step wait1: <... completed>
+FATAL:  terminating connection due to administrator command
+server closed the connection unexpectedly
+	This probably means the server terminated abnormally
+	before or while processing the request.
+
+step terminate3: <... completed>
+pg_terminate_backend
+--------------------
+t                   
+(1 row)
+
+step noop3: 
+step wait2: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step wakeup3: SELECT injection_points_wakeup('injection-points-wait');
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait2: <... completed>
+injection_points_run
+--------------------
+                    
+(1 row)
+
+step noop2: 
+step detach3: SELECT injection_points_detach('injection-points-wait');
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/injection_points.c b/src/test/modules/injection_points/injection_points.c
index 2d26ecedd5d..a549720ed75 100644
--- a/src/test/modules/injection_points/injection_points.c
+++ b/src/test/modules/injection_points/injection_points.c
@@ -223,6 +223,25 @@ injection_notice(const char *name, const void *private_data, void *arg)
 		elog(NOTICE, "notice triggered for injection point %s", name);
 }
 
+/*
+ * Clear the shared waiter slot of this backend, on error or process exit.
+ *
+ * Without this a canceled or terminated waiter would leak its slot, and
+ * later wakeups of the same injection point would bump the leaked slot's
+ * counter instead of the real waiter's. Registered through
+ * PG_ENSURE_ERROR_CLEANUP so that it also runs on FATAL, which exits
+ * without unwinding the stack.
+ */
+static void
+injection_wait_cleanup(int code, Datum arg)
+{
+	int			index = DatumGetInt32(arg);
+
+	SpinLockAcquire(&inj_state->lock);
+	inj_state->name[index][0] = '\0';
+	SpinLockRelease(&inj_state->lock);
+}
+
 /* Wait until injection_points_wakeup() is called */
 void
 injection_wait(const char *name, const void *private_data, void *arg)
@@ -275,19 +294,21 @@ injection_wait(const char *name, const void *private_data, void *arg)
 	delay_us = INJ_WAIT_INITIAL_US;
 
 	pgstat_report_wait_start(injection_wait_event);
-	while (pg_atomic_read_u32(&inj_state->wait_counts[index]) == old_wait_counts)
+	PG_ENSURE_ERROR_CLEANUP(injection_wait_cleanup, Int32GetDatum(index));
 	{
-		CHECK_FOR_INTERRUPTS();
-		pg_usleep(delay_us);
-		if (delay_us < INJ_WAIT_MAX_US)
-			delay_us = Min(delay_us * 2, INJ_WAIT_MAX_US);
+		while (pg_atomic_read_u32(&inj_state->wait_counts[index]) == old_wait_counts)
+		{
+			CHECK_FOR_INTERRUPTS();
+			pg_usleep(delay_us);
+			if (delay_us < INJ_WAIT_MAX_US)
+				delay_us = Min(delay_us * 2, INJ_WAIT_MAX_US);
+		}
 	}
+	PG_END_ENSURE_ERROR_CLEANUP(injection_wait_cleanup, Int32GetDatum(index));
 	pgstat_report_wait_end();
 
 	/* Remove this injection point from the waiters. */
-	SpinLockAcquire(&inj_state->lock);
-	inj_state->name[index][0] = '\0';
-	SpinLockRelease(&inj_state->lock);
+	injection_wait_cleanup(0, Int32GetDatum(index));
 }
 
 /*
diff --git a/src/test/modules/injection_points/meson.build b/src/test/modules/injection_points/meson.build
index 59dba1cb023..a903bc94a88 100644
--- a/src/test/modules/injection_points/meson.build
+++ b/src/test/modules/injection_points/meson.build
@@ -44,6 +44,7 @@ tests += {
   'isolation': {
     'specs': [
       'basic',
+      'wait_cleanup',
       'inplace',
       'repack',
       'repack_temporal',
diff --git a/src/test/modules/injection_points/specs/wait_cleanup.spec b/src/test/modules/injection_points/specs/wait_cleanup.spec
new file mode 100644
index 00000000000..4ef8c293e83
--- /dev/null
+++ b/src/test/modules/injection_points/specs/wait_cleanup.spec
@@ -0,0 +1,50 @@
+# Check that a canceled or terminated waiter does not leave a stale slot
+# behind in the waiter array. A leaked slot would make later wakeups of
+# the same injection point bump the leaked slot's counter instead of the
+# real waiter's, leaving the real waiter stuck.
+
+setup
+{
+	CREATE EXTENSION injection_points;
+}
+teardown
+{
+	DROP EXTENSION injection_points;
+}
+
+# The first waiter, which gets canceled or terminated. No set_local:
+# the injection point must survive s1's termination so that s3 can
+# still detach it.
+session s1
+setup	{
+	SELECT injection_points_attach('injection-points-wait', 'wait');
+}
+step wait1	{ SELECT injection_points_run('injection-points-wait'); }
+
+# The second waiter, which must still receive the wakeup.
+session s2
+step wait2	{ SELECT injection_points_run('injection-points-wait'); }
+step noop2	{ }
+
+# The control session. The blocker annotations on cancel3/terminate3
+# plus noop3 make the tester wait until wait1 actually finished before
+# starting wait2, otherwise wait2 could register its own waiter slot
+# while s1 still holds the old one.
+session s3
+step cancel3	{
+	SELECT pg_cancel_backend(pid) FROM pg_stat_activity
+	  WHERE wait_event = 'injection-points-wait';
+}
+step terminate3	{
+	SELECT pg_terminate_backend(pid) FROM pg_stat_activity
+	  WHERE wait_event = 'injection-points-wait';
+}
+step wakeup3	{ SELECT injection_points_wakeup('injection-points-wait'); }
+step detach3	{ SELECT injection_points_detach('injection-points-wait'); }
+step noop3	{ }
+
+permutation wait1 cancel3(wait1) noop3 wait2 wakeup3 noop2 detach3
+
+# The terminate permutation has to stay last: s1's connection is dead
+# afterwards, and the tester never reconnects a session.
+permutation wait1 terminate3(wait1) noop3 wait2 wakeup3 noop2 detach3
-- 
2.54.0



^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-07-22 01:29  Michael Paquier <michael@paquier.xyz>
  parent: Zsolt Parragi <zsolt.parragi@percona.com>
  0 siblings, 1 reply; 21+ messages in thread

From: Michael Paquier @ 2026-07-22 01:29 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>

On Tue, Jul 21, 2026 at 09:28:22AM +0100, Zsolt Parragi wrote:
> While stress-testing REPACK CONCURRENTLY on 19beta2 I saw a logical
> decoding activation race, and I extended 051_effective_wal_level.pl
> with a test that waits on an injection point and wakes it up later.
> 051 cancels two injection point waiters earlier in the script, and
> injection_wait() never cleans up after a canceled waiter, the wakeup
> never arrived and the test deadlocked.
> 
> I think we are missing an ENSURE_ERROR_CLEANUP block there. See
> attached patch with a testcase reproducing the issue.

Hmm.  I think that I'd rather use a PG_TRY/PG_FINALLY and avoid the
refactoring with the extra routine required, keeping the cleanup
action local to injection_wait().  That's also because the cleanup
action is the same for both the "normal" exit path and the interrupt
path.

> I also attached a separate version for pg19, as master has a
> refactored version of injection_wait. All previous branches have the
> 19 version, it should be easy to backport to other branches.

Thanks for that.  I'm always OK to deal with a backpatch as required.
Posting versions saves some time, of course, just don't feel obliged
if you feel that this is extra work on your side.

On an unpatched code, the test would hang due to the fact that we are
doing a wait but we should not because the slot was not cleaned up.
It means that a failure mode equals to a timeout.  Why not, we have
other tests of this class.  Another thought: the addition of a SQL
function that provides the list of waiters that we reuse here.  I
don't really see why this is worth the cost compared to your test, but
opinions of others are welcome, of course.
--
Michael

Attachments:

  [application/pgp-signature] signature.asc (832B, ../../amAc853flInQAEI5@paquier.xyz/2-signature.asc)
  download

^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-07-22 06:07  Zsolt Parragi <zsolt.parragi@percona.com>
  parent: Michael Paquier <michael@paquier.xyz>
  0 siblings, 1 reply; 21+ messages in thread

From: Zsolt Parragi @ 2026-07-22 06:07 UTC (permalink / raw)
  To: Michael Paquier <michael@paquier.xyz>; +Cc: pgsql-hackers@lists.postgresql.org

> I think that I'd rather use a PG_TRY/PG_FINALLY and avoid the
> refactoring with the extra routine required, keeping the cleanup
> action local to injection_wait()

Wouldn't that miss FATAL? (pg_terminate_backend in the testcase)

> Thanks for that. I'm always OK to deal with a backpatch as required.
> Posting versions saves some time, of course, just don't feel obliged
> if you feel that this is extra work on your side.

I don't always consistently do this (sometimes it's interesting to
check that its different in back branches, sometimes I completely
forgot about it)

In this case, I started the testing / debugging on pg19, and only
checked that the code is different on master after I already had the
fix. Verifying that the code is the same on earlier branches wasn't
much extra work after that.

> Another thought: the addition of a SQL
> function that provides the list of waiters that we reuse here.

I can implement that if you think it is better, my initial logic was
to keep the fix smaller and strictly a bugfix.






^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-07-22 06:28  Michael Paquier <michael@paquier.xyz>
  parent: Zsolt Parragi <zsolt.parragi@percona.com>
  0 siblings, 1 reply; 21+ messages in thread

From: Michael Paquier @ 2026-07-22 06:28 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: pgsql-hackers@lists.postgresql.org

On Tue, Jul 21, 2026 at 11:07:01PM -0700, Zsolt Parragi wrote:
> Wouldn't that miss FATAL? (pg_terminate_backend in the testcase)

Arf.  I've forgotten that elog.h documents that.  Thanks.

> I can implement that if you think it is better, my initial logic was
> to keep the fix smaller and strictly a bugfix.

Minimal sounds good here.  I'll make something happen after
double-checking what you have sent.
--
Michael

Attachments:

  [application/pgp-signature] signature.asc (832B, ../../amBjKAJLyPR5hdlE@paquier.xyz/2-signature.asc)
  download

^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-07-23 05:40  Michael Paquier <michael@paquier.xyz>
  parent: Michael Paquier <michael@paquier.xyz>
  0 siblings, 1 reply; 21+ messages in thread

From: Michael Paquier @ 2026-07-23 05:40 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: pgsql-hackers@lists.postgresql.org

On Wed, Jul 22, 2026 at 03:28:56PM +0900, Michael Paquier wrote:
> Minimal sounds good here.  I'll make something happen after
> double-checking what you have sent.

And now done down to v17, as of a49b6a610946.
--
Michael

Attachments:

  [application/pgp-signature] signature.asc (832B, ../../amGpM1v-hhvVav3n@paquier.xyz/2-signature.asc)
  download

^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-08-23 09:13  Andrey Borodin <x4mmm@yandex-team.ru>
  parent: Michael Paquier <michael@paquier.xyz>
  0 siblings, 1 reply; 21+ messages in thread

From: Andrey Borodin @ 2026-08-23 09:13 UTC (permalink / raw)
  To: Michael Paquier <michael@paquier.xyz>; +Cc: Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-hackers mailing list <pgsql-hackers@lists.postgresql.org>



> On 23 Jul 2026, at 08:40, Michael Paquier <michael@paquier.xyz> wrote:
> 
> done

Hi,

While running CI for an unrelated patch, I saw wait_cleanup fail in the
Windows Visual Studio job[0].

The server log contains the expected FATAL, but isolationtester only saw:

  PQconsumeInput failed: server closed the connection unexpectedly

It then exited without running the rest of the permutation or teardown, so
heap_lock_update failed afterwards because the injection_points extension
still existed.

This seems to be another instance of the known Windows behavior where the
last server message can be lost when a connection is closed [1].  The test
added in a49b6a61094 intentionally terminates an isolationtester connection,
so it is exposed to that behavior.

The attached patch makes isolationtester treat PQconsumeInput() failure with
CONNECTION_BAD as completion of the step.  It reports any complete server
error already buffered by libpq, followed by the saved connection error, and
the rest of the test and teardown can run.  Other PQconsumeInput() failures
remain fatal.  An alternative expected file covers the case where Windows
loses the server's FATAL and only the libpq-generated connection error
remains.

The alternative output is synthetic.  I tested it by temporarily suppressing
the final ErrorResponse while leaving backend termination unchanged.  The
output then matched wait_cleanup_1.out.  With normal error delivery it matched
wait_cleanup.out.  If anyone knows a way to reproduce the actual Windows
message loss on demand, that would be useful.  Otherwise, the next occurrence
in CI with this patch applied will give us an output to compare with the
alternative file.

The injection_points isolation tests pass through Windows CI.

If I have misdiagnosed the cause of this CI failure, apologies for the noise.


Best regards, Andrey Borodin.

[0] https://github.com/x4m/postgres_g/actions/runs/32582921356/job/97055012805
[1] https://postgr.es/m/CA%2BhUKGLR10ZqRCvdoRrkQusq75wF5%3DvEetRSs2_u1s%2BFAUosFQ%40mail.gmail.com

Attachments:

  [application/octet-stream] v1-0001-Let-isolationtester-report-connection-loss-as-a-s.patch (5.5K, ../../088AB35D-0860-494A-A9CF-8B301543AA51@yandex-team.ru/2-v1-0001-Let-isolationtester-report-connection-loss-as-a-s.patch)
  download | inline diff:
From 2de6e4b3548ba979fe413a5796132fd0e5761f49 Mon Sep 17 00:00:00 2001
From: Andrey Borodin <amborodin@acm.org>
Date: Sat, 22 Aug 2026 20:38:51 +0300
Subject: [PATCH v1] Let isolationtester report connection loss as a step
 result

---
 src/test/isolation/isolationtester.c          | 45 ++++++++--
 .../expected/wait_cleanup_1.out               | 86 +++++++++++++++++++
 2 files changed, 124 insertions(+), 7 deletions(-)
 create mode 100644 src/test/modules/injection_points/expected/wait_cleanup_1.out

diff --git a/src/test/isolation/isolationtester.c b/src/test/isolation/isolationtester.c
index e64dc020b18..31e7341a67a 100644
--- a/src/test/isolation/isolationtester.c
+++ b/src/test/isolation/isolationtester.c
@@ -828,6 +828,7 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 	PGresult   *res;
 	PGnotify   *notify;
 	bool		canceled = false;
+	char	   *connection_error = NULL;
 
 	/*
 	 * If the step is annotated with (*), then on the first call, force it to
@@ -910,9 +911,19 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 					 */
 					if (!PQconsumeInput(conn))
 					{
-						fprintf(stderr, "PQconsumeInput failed: %s\n",
-								PQerrorMessage(conn));
-						exit(1);
+						if (PQstatus(conn) != CONNECTION_BAD)
+						{
+							fprintf(stderr, "PQconsumeInput failed: %s\n",
+									PQerrorMessage(conn));
+							exit(1);
+						}
+
+						/*
+						 * Save the error before PQgetResult() adds another complaint
+						 * about attempting to read from the dead socket.
+						 */
+						connection_error = pg_strdup(PQerrorMessage(conn));
+						break;
 					}
 					if (!PQisBusy(conn))
 						break;
@@ -979,9 +990,19 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 		}
 		else if (!PQconsumeInput(conn)) /* select(): data available */
 		{
-			fprintf(stderr, "PQconsumeInput failed: %s\n",
-					PQerrorMessage(conn));
-			exit(1);
+			if (PQstatus(conn) != CONNECTION_BAD)
+			{
+				fprintf(stderr, "PQconsumeInput failed: %s\n",
+						PQerrorMessage(conn));
+				exit(1);
+			}
+
+			/*
+			 * Save the error before PQgetResult() adds another complaint about
+			 * attempting to read from the dead socket.
+			 */
+			connection_error = pg_strdup(PQerrorMessage(conn));
+			break;
 		}
 	}
 
@@ -1028,7 +1049,7 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 
 					if (sev && msg)
 						printf("%s:  %s\n", sev, msg);
-					else
+					else if (!connection_error)
 						printf("%s\n", PQresultErrorMessage(res));
 				}
 				break;
@@ -1037,6 +1058,16 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 					   PQresStatus(PQresultStatus(res)));
 		}
 		PQclear(res);
+
+		/* The connection is dead, so don't ask libpq for another result. */
+		if (connection_error)
+			break;
+	}
+
+	if (connection_error)
+	{
+		printf("%s\n", connection_error);
+		pg_free(connection_error);
 	}
 
 	/* Report any available NOTIFY messages, too */
diff --git a/src/test/modules/injection_points/expected/wait_cleanup_1.out b/src/test/modules/injection_points/expected/wait_cleanup_1.out
new file mode 100644
index 00000000000..516a428b364
--- /dev/null
+++ b/src/test/modules/injection_points/expected/wait_cleanup_1.out
@@ -0,0 +1,86 @@
+Parsed test spec with 3 sessions
+
+starting permutation: wait1 cancel3 noop3 wait2 wakeup3 noop2 detach3
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step wait1: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step cancel3: 
+	SELECT pg_cancel_backend(pid) FROM pg_stat_activity
+	  WHERE wait_event = 'injection-points-wait';
+ <waiting ...>
+step wait1: <... completed>
+ERROR:  canceling statement due to user request
+step cancel3: <... completed>
+pg_cancel_backend
+-----------------
+t                
+(1 row)
+
+step noop3: 
+step wait2: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step wakeup3: SELECT injection_points_wakeup('injection-points-wait');
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait2: <... completed>
+injection_points_run
+--------------------
+                    
+(1 row)
+
+step noop2: 
+step detach3: SELECT injection_points_detach('injection-points-wait');
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
+
+starting permutation: wait1 terminate3 noop3 wait2 wakeup3 noop2 detach3
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step wait1: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step terminate3: 
+	SELECT pg_terminate_backend(pid) FROM pg_stat_activity
+	  WHERE wait_event = 'injection-points-wait';
+ <waiting ...>
+step wait1: <... completed>
+server closed the connection unexpectedly
+	This probably means the server terminated abnormally
+	before or while processing the request.
+
+step terminate3: <... completed>
+pg_terminate_backend
+--------------------
+t                   
+(1 row)
+
+step noop3: 
+step wait2: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step wakeup3: SELECT injection_points_wakeup('injection-points-wait');
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait2: <... completed>
+injection_points_run
+--------------------
+                    
+(1 row)
+
+step noop2: 
+step detach3: SELECT injection_points_detach('injection-points-wait');
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
-- 
That's all, folks. May the source be with you.

=

^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-09-23 16:42  Nikolay Samokhvalov <nik@postgres.ai>
  parent: Andrey Borodin <x4mmm@yandex-team.ru>
  0 siblings, 1 reply; 21+ messages in thread

From: Nikolay Samokhvalov @ 2026-09-23 16:42 UTC (permalink / raw)
  To: Andrey Borodin <x4mmm@yandex-team.ru>; +Cc: Michael Paquier <michael@paquier.xyz>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-hackers mailing list <pgsql-hackers@lists.postgresql.org>

On Sun, Aug 23, 2026, Andrey Borodin <x4mmm@yandex-team.ru> wrote:
> The attached patch makes isolationtester treat PQconsumeInput() failure with
> CONNECTION_BAD as completion of the step.

Thanks Andrey. CF 6992 hit the same symptom in slot_creation_error on
Windows. Attached are three follow-ups on top of your v1.

There is another case with a notice blocker: try_complete_step() can return
while the step is still blocked, losing the saved connection error. The next
attempt then fails with "invalid socket". The fix keeps the error in
IsoConnInfo until the step completes. The other patches add the alternate
slot_creation_error output and extend wait_cleanup to cover this case.

With the extended test and the final ErrorResponse suppressed, v1 fails
before release3/detach3; the fixed version completes both. The suppression
is test-only and leaves the server's FATAL in the log. It does not force
Windows transport loss.

The Windows run below passed core regression/isolation, test_decoding,
injection_points isolation and subscription TAP (one ICU-dependent skip),
with your v1 plus these patches. The workflow includes setup, the missing
ErrorResponse test and old-driver controls:

https://github.com/NikolayS/postgres/actions/runs/35770093308
https://github.com/NikolayS/postgres/blob/8039ee47313f7ac860d26bd6fd9892c0ac296fee/.github/workflows...

AI found the blocker case and prepared these follow-up patches.
I have not manually reviewed these patches.

Nik

Attachments:

  [application/x-patch] 0002-test-decoding-windows-expected-output.patch (3.9K, ../../CAM527d8gNR9MW-BSC=DFm7aufx6P+_hALSVr-ML9Er0KpE9jjQ@mail.gmail.com/2-0002-test-decoding-windows-expected-output.patch)
  download | inline diff:
From 30c1a14b115299ee10a43a3497f4de763ddf8e4e Mon Sep 17 00:00:00 2001
From: Nik Samokhvalov <nik@postgres.ai>
Date: Tue, 22 Sep 2026 08:21:30 -0700
Subject: [PATCH] test_decoding: accept lost FATAL message on backend
 termination

With isolationtester able to finish a step after connection loss, accept the Windows case where the final ErrorResponse is not received. Keep all subsequent checks, including slot cleanup, in the expected output.
---
 .../expected/slot_creation_error_1.out        | 113 ++++++++++++++++++
 1 file changed, 113 insertions(+)
 create mode 100644 contrib/test_decoding/expected/slot_creation_error_1.out

diff --git a/contrib/test_decoding/expected/slot_creation_error_1.out b/contrib/test_decoding/expected/slot_creation_error_1.out
new file mode 100644
index 00000000000..2922022ce32
--- /dev/null
+++ b/contrib/test_decoding/expected/slot_creation_error_1.out
@@ -0,0 +1,113 @@
+Parsed test spec with 2 sessions
+
+starting permutation: s1_b s1_xid s2_init s1_view_slot s1_cancel_s2 s1_view_slot s1_c
+step s1_b: BEGIN;
+step s1_xid: SELECT 'xid' FROM txid_current();
+?column?
+--------
+xid     
+(1 row)
+
+step s2_init: 
+    SELECT 'init' FROM pg_create_logical_replication_slot('slot_creation_error', 'test_decoding');
+ <waiting ...>
+step s1_view_slot: 
+    SELECT slot_name, slot_type, active FROM pg_replication_slots WHERE slot_name = 'slot_creation_error'
+
+slot_name          |slot_type|active
+-------------------+---------+------
+slot_creation_error|logical  |t     
+(1 row)
+
+step s1_cancel_s2: 
+    SELECT pg_cancel_backend(pid)
+    FROM pg_stat_activity
+    WHERE application_name = 'isolation/slot_creation_error/s2';
+ <waiting ...>
+step s2_init: <... completed>
+ERROR:  canceling statement due to user request
+step s1_cancel_s2: <... completed>
+pg_cancel_backend
+-----------------
+t                
+(1 row)
+
+step s1_view_slot: 
+    SELECT slot_name, slot_type, active FROM pg_replication_slots WHERE slot_name = 'slot_creation_error'
+
+slot_name|slot_type|active
+---------+---------+------
+(0 rows)
+
+step s1_c: COMMIT;
+
+starting permutation: s1_b s1_xid s2_init s1_c s1_view_slot s1_drop_slot
+step s1_b: BEGIN;
+step s1_xid: SELECT 'xid' FROM txid_current();
+?column?
+--------
+xid     
+(1 row)
+
+step s2_init: 
+    SELECT 'init' FROM pg_create_logical_replication_slot('slot_creation_error', 'test_decoding');
+ <waiting ...>
+step s1_c: COMMIT;
+step s2_init: <... completed>
+?column?
+--------
+init    
+(1 row)
+
+step s1_view_slot: 
+    SELECT slot_name, slot_type, active FROM pg_replication_slots WHERE slot_name = 'slot_creation_error'
+
+slot_name          |slot_type|active
+-------------------+---------+------
+slot_creation_error|logical  |f     
+(1 row)
+
+step s1_drop_slot: 
+    SELECT pg_drop_replication_slot('slot_creation_error');
+
+pg_drop_replication_slot
+------------------------
+                        
+(1 row)
+
+
+starting permutation: s1_b s1_xid s2_init s1_terminate_s2 s1_c s1_view_slot
+step s1_b: BEGIN;
+step s1_xid: SELECT 'xid' FROM txid_current();
+?column?
+--------
+xid     
+(1 row)
+
+step s2_init: 
+    SELECT 'init' FROM pg_create_logical_replication_slot('slot_creation_error', 'test_decoding');
+ <waiting ...>
+step s1_terminate_s2: 
+    SELECT pg_terminate_backend(pid)
+    FROM pg_stat_activity
+    WHERE application_name = 'isolation/slot_creation_error/s2';
+ <waiting ...>
+step s2_init: <... completed>
+server closed the connection unexpectedly
+	This probably means the server terminated abnormally
+	before or while processing the request.
+
+step s1_terminate_s2: <... completed>
+pg_terminate_backend
+--------------------
+t                   
+(1 row)
+
+step s1_c: COMMIT;
+step s1_view_slot: 
+    SELECT slot_name, slot_type, active FROM pg_replication_slots WHERE slot_name = 'slot_creation_error'
+
+slot_name|slot_type|active
+---------+---------+------
+(0 rows)
+
-- 
2.50.1 (Apple Git-155)



  [application/x-patch] 0004-test-connection-loss-with-active-blocker.patch (6.1K, ../../CAM527d8gNR9MW-BSC=DFm7aufx6P+_hALSVr-ML9Er0KpE9jjQ@mail.gmail.com/3-0004-test-connection-loss-with-active-blocker.patch)
  download | inline diff:
From 0dda85366844bb48157980992f47d19a68152537 Mon Sep 17 00:00:00 2001
From: Nik Samokhvalov <nik@postgres.ai>
Date: Tue, 22 Sep 2026 11:32:09 -0700
Subject: [PATCH] isolationtester: test connection loss with an active blocker

Wait for the terminated backend to exit before releasing the notice blocker. This makes the missing-error case exercise retaining a connection error across scheduler retries.
---
 .../expected/wait_cleanup.out                 | 19 ++++++++++---------
 .../expected/wait_cleanup_1.out               | 17 +++++++++--------
 .../injection_points/specs/wait_cleanup.spec  | 16 ++++++++++------
 3 files changed, 29 insertions(+), 23 deletions(-)

diff --git a/src/test/modules/injection_points/expected/wait_cleanup.out b/src/test/modules/injection_points/expected/wait_cleanup.out
index c5be17428fc..9c60ecfb2d4 100644
--- a/src/test/modules/injection_points/expected/wait_cleanup.out
+++ b/src/test/modules/injection_points/expected/wait_cleanup.out
@@ -41,7 +41,7 @@ injection_points_detach
 (1 row)
 
 
-starting permutation: wait1 terminate3 noop3 wait2 wakeup3 noop2 detach3
+starting permutation: wait1 terminate3 noop3 release3 wait2 wakeup3 noop2 detach3
 injection_points_attach
 -----------------------
                        
@@ -49,22 +49,23 @@ injection_points_attach
 
 step wait1: SELECT injection_points_run('injection-points-wait'); <waiting ...>
 step terminate3: 
-	SELECT pg_terminate_backend(pid) FROM pg_stat_activity
+	SELECT pg_terminate_backend(pid, 180000) FROM pg_stat_activity
 	  WHERE wait_event = 'injection-points-wait';
- <waiting ...>
-step wait1: <... completed>
-FATAL:  terminating connection due to administrator command
-server closed the connection unexpectedly
-	This probably means the server terminated abnormally
-	before or while processing the request.
 
-step terminate3: <... completed>
 pg_terminate_backend
 --------------------
 t                   
 (1 row)
 
 step noop3: 
+s3: NOTICE:  release wait1
+step release3: DO $$BEGIN RAISE NOTICE 'release wait1'; END$$;
+step wait1: <... completed>
+FATAL:  terminating connection due to administrator command
+server closed the connection unexpectedly
+	This probably means the server terminated abnormally
+	before or while processing the request.
+
 step wait2: SELECT injection_points_run('injection-points-wait'); <waiting ...>
 step wakeup3: SELECT injection_points_wakeup('injection-points-wait');
 injection_points_wakeup
diff --git a/src/test/modules/injection_points/expected/wait_cleanup_1.out b/src/test/modules/injection_points/expected/wait_cleanup_1.out
index 516a428b364..71ae751cbee 100644
--- a/src/test/modules/injection_points/expected/wait_cleanup_1.out
+++ b/src/test/modules/injection_points/expected/wait_cleanup_1.out
@@ -41,7 +41,7 @@ injection_points_detach
 (1 row)
 
 
-starting permutation: wait1 terminate3 noop3 wait2 wakeup3 noop2 detach3
+starting permutation: wait1 terminate3 noop3 release3 wait2 wakeup3 noop2 detach3
 injection_points_attach
 -----------------------
                        
@@ -49,21 +49,22 @@ injection_points_attach
 
 step wait1: SELECT injection_points_run('injection-points-wait'); <waiting ...>
 step terminate3: 
-	SELECT pg_terminate_backend(pid) FROM pg_stat_activity
+	SELECT pg_terminate_backend(pid, 180000) FROM pg_stat_activity
 	  WHERE wait_event = 'injection-points-wait';
- <waiting ...>
-step wait1: <... completed>
-server closed the connection unexpectedly
-	This probably means the server terminated abnormally
-	before or while processing the request.
 
-step terminate3: <... completed>
 pg_terminate_backend
 --------------------
 t                   
 (1 row)
 
 step noop3: 
+s3: NOTICE:  release wait1
+step release3: DO $$BEGIN RAISE NOTICE 'release wait1'; END$$;
+step wait1: <... completed>
+server closed the connection unexpectedly
+	This probably means the server terminated abnormally
+	before or while processing the request.
+
 step wait2: SELECT injection_points_run('injection-points-wait'); <waiting ...>
 step wakeup3: SELECT injection_points_wakeup('injection-points-wait');
 injection_points_wakeup
diff --git a/src/test/modules/injection_points/specs/wait_cleanup.spec b/src/test/modules/injection_points/specs/wait_cleanup.spec
index ed7d21c4de4..fa482f2422c 100644
--- a/src/test/modules/injection_points/specs/wait_cleanup.spec
+++ b/src/test/modules/injection_points/specs/wait_cleanup.spec
@@ -26,9 +26,9 @@ session s2
 step wait2	{ SELECT injection_points_run('injection-points-wait'); }
 step noop2	{ }
 
-# Control session.  The blocker annotations on cancel3/terminate3,
-# together with noop3, make the tester wait until wait1 has fully
-# completed before starting wait2.  Otherwise, wait2 could register a
+# Control session.  The blocker on cancel3 and the notice from release3
+# make the tester wait until wait1 has fully completed before starting
+# wait2.  Otherwise, wait2 could register a
 # new waiter slot while s1 still owns the previous one.
 session s3
 step cancel3	{
@@ -36,15 +36,19 @@ step cancel3	{
 	  WHERE wait_event = 'injection-points-wait';
 }
 step terminate3	{
-	SELECT pg_terminate_backend(pid) FROM pg_stat_activity
+	SELECT pg_terminate_backend(pid, 180000) FROM pg_stat_activity
 	  WHERE wait_event = 'injection-points-wait';
 }
 step wakeup3	{ SELECT injection_points_wakeup('injection-points-wait'); }
 step detach3	{ SELECT injection_points_detach('injection-points-wait'); }
+step release3	{ DO $$BEGIN RAISE NOTICE 'release wait1'; END$$; }
 step noop3	{ }
 
 permutation wait1 cancel3(wait1) noop3 wait2 wakeup3 noop2 detach3
 
 # The terminate permutation has to stay last: s1's connection is dead
-# afterwards, and the tester never reconnects a session.
-permutation wait1 terminate3(wait1) noop3 wait2 wakeup3 noop2 detach3
+# afterwards, and the tester never reconnects a session.  Delay reporting
+# wait1 until release3 sends its notice, even if the connection is already
+# closed, to exercise retaining the connection error across step retries.
+# terminate3 waits for backend exit before the notice blocker is released.
+permutation wait1(release3 notices 1) terminate3 noop3 release3 wait2 wakeup3 noop2 detach3
-- 
2.50.1 (Apple Git-155)



  [application/x-patch] 0003-preserve-connection-error-across-blockers.patch (3.5K, ../../CAM527d8gNR9MW-BSC=DFm7aufx6P+_hALSVr-ML9Er0KpE9jjQ@mail.gmail.com/4-0003-preserve-connection-error-across-blockers.patch)
  download | inline diff:
From ce461649028aba48133acd7a8270c4fb1145fdd0 Mon Sep 17 00:00:00 2001
From: Nik Samokhvalov <nik@postgres.ai>
Date: Tue, 22 Sep 2026 08:25:16 -0700
Subject: [PATCH] isolationtester: retain connection errors across blocked
 steps

Keep the saved error with the active connection until the step can be reported. A blocked step may be retried after its socket has closed. Drain all complete buffered results before reporting the saved connection error.
---
 src/test/isolation/isolationtester.c | 24 +++++++++++-------------
 1 file changed, 11 insertions(+), 13 deletions(-)

diff --git a/src/test/isolation/isolationtester.c b/src/test/isolation/isolationtester.c
index 31e7341a67a..b7e39c48672 100644
--- a/src/test/isolation/isolationtester.c
+++ b/src/test/isolation/isolationtester.c
@@ -33,6 +33,8 @@ typedef struct IsoConnInfo
 	const char *sessionname;
 	/* Active step on this connection, or NULL if idle. */
 	PermutationStep *active_step;
+	/* Connection error retained while an active step has blockers. */
+	char	   *connection_error;
 	/* Number of NOTICE messages received from connection. */
 	int			total_notices;
 } IsoConnInfo;
@@ -828,7 +830,6 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 	PGresult   *res;
 	PGnotify   *notify;
 	bool		canceled = false;
-	char	   *connection_error = NULL;
 
 	/*
 	 * If the step is annotated with (*), then on the first call, force it to
@@ -852,7 +853,7 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 		}
 	}
 
-	if (sock < 0)
+	if (sock < 0 && !iconn->connection_error)
 	{
 		fprintf(stderr, "invalid socket: %s", PQerrorMessage(conn));
 		exit(1);
@@ -861,7 +862,7 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 	gettimeofday(&start_time, NULL);
 	FD_ZERO(&read_set);
 
-	while (PQisBusy(conn))
+	while (!iconn->connection_error && PQisBusy(conn))
 	{
 		FD_SET(sock, &read_set);
 		timeout.tv_sec = 0;
@@ -922,7 +923,7 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 						 * Save the error before PQgetResult() adds another complaint
 						 * about attempting to read from the dead socket.
 						 */
-						connection_error = pg_strdup(PQerrorMessage(conn));
+						iconn->connection_error = pg_strdup(PQerrorMessage(conn));
 						break;
 					}
 					if (!PQisBusy(conn))
@@ -1001,7 +1002,7 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 			 * Save the error before PQgetResult() adds another complaint about
 			 * attempting to read from the dead socket.
 			 */
-			connection_error = pg_strdup(PQerrorMessage(conn));
+			iconn->connection_error = pg_strdup(PQerrorMessage(conn));
 			break;
 		}
 	}
@@ -1049,7 +1050,7 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 
 					if (sev && msg)
 						printf("%s:  %s\n", sev, msg);
-					else if (!connection_error)
+					else if (!iconn->connection_error)
 						printf("%s\n", PQresultErrorMessage(res));
 				}
 				break;
@@ -1058,16 +1059,13 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 					   PQresStatus(PQresultStatus(res)));
 		}
 		PQclear(res);
-
-		/* The connection is dead, so don't ask libpq for another result. */
-		if (connection_error)
-			break;
 	}
 
-	if (connection_error)
+	if (iconn->connection_error)
 	{
-		printf("%s\n", connection_error);
-		pg_free(connection_error);
+		printf("%s\n", iconn->connection_error);
+		pg_free(iconn->connection_error);
+		iconn->connection_error = NULL;
 	}
 
 	/* Report any available NOTIFY messages, too */
-- 
2.50.1 (Apple Git-155)



^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-09-28 22:12  Zsolt Parragi <zsolt.parragi@percona.com>
  parent: Nikolay Samokhvalov <nik@postgres.ai>
  0 siblings, 1 reply; 21+ messages in thread

From: Zsolt Parragi @ 2026-09-28 22:12 UTC (permalink / raw)
  To: Nikolay Samokhvalov <nik@postgres.ai>; +Cc: pgsql-hackers@lists.postgresql.org, Andrey Borodin <x4mmm@yandex-team.ru>; Michael Paquier <michael@paquier.xyz>

Thanks for the patches!

I was able to reproduce the issue both on windows, and on linux by
"patching" the server to reproduce the described behavior.

The overall changes look good to me, it just needs some
squashing/organizing to be a proper patchset.





^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-09-28 22:57  Michael Paquier <michael@paquier.xyz>
  parent: Zsolt Parragi <zsolt.parragi@percona.com>
  0 siblings, 1 reply; 21+ messages in thread

From: Michael Paquier @ 2026-09-28 22:57 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Nikolay Samokhvalov <nik@postgres.ai>; pgsql-hackers@lists.postgresql.org, Andrey Borodin <x4mmm@yandex-team.ru>

On Mon, Sep 28, 2026 at 10:12:37PM +0000, Zsolt Parragi wrote:
> I was able to reproduce the issue both on windows, and on linux by
> "patching" the server to reproduce the described behavior.

How exactly?

> The overall changes look good to me, it just needs some
> squashing/organizing to be a proper patchset.

TBH, I am not completely sure what you are proposing here.  v1-0001
from Andrey and AI-generated-not-reviewed 0003 from Nikolay step on
each other with the latter patch requiring the former patch, in terms
of the way the last buffered error messages can be consumed, which
should be either something local to try_complete_step() or tracked by
IsoConnInfo.  I'm OK with the basic idea of having more predictible
output here, with some alternate outputs to make the CI happier on
Windows, let's just organize a bit the whole..

-       SELECT pg_terminate_backend(pid) FROM pg_stat_activity
+       SELECT pg_terminate_backend(pid, 180000) FROM pg_stat_activity

Avoiding hardcoded timeouts would be nice.  They are not liked on slow
machines.
--
Michael

Attachments:

  [application/pgp-signature] signature.asc (832B, ../../arrw8isOSbfkK5Fd@paquier.xyz/2-signature.asc)
  download

^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-09-29 00:38  Zsolt Parragi <zsolt.parragi@percona.com>
  parent: Michael Paquier <michael@paquier.xyz>
  0 siblings, 1 reply; 21+ messages in thread

From: Zsolt Parragi @ 2026-09-29 00:38 UTC (permalink / raw)
  To: Michael Paquier <michael@paquier.xyz>; +Cc: Nikolay Samokhvalov <nik@postgres.ai>; pgsql-hackers@lists.postgresql.org, Andrey Borodin <x4mmm@yandex-team.ru>

> How exactly?

If we assume the linked thread about windows skipping the last message
is correct, we can simulate that by adding an option not to send the
fatal message to the client:

diff --git a/src/backend/utils/error/elog.c b/src/backend/utils/error/elog.c
index b9d2c96b97a..e043717b943 100644
--- a/src/backend/utils/error/elog.c
+++ b/src/backend/utils/error/elog.c
@@ -1924,7 +1924,8 @@ EmitErrorReport(void)
                send_message_to_server_log(edata);

        /* Send to client, if enabled */
-       if (edata->output_to_client)
+       if (edata->output_to_client &&
+               !(edata->elevel == FATAL && getenv("PG_TEST_DROP_FATAL")))
                send_message_to_frontend(edata);

        MemoryContextSwitchTo(oldcontext);

On windows, it seems it can randomly happen, but it's very unlikely

> TBH, I am not completely sure what you are proposing here. v1-0001
> from Andrey and AI-generated-not-reviewed 0003 from Nikolay step on
> each other

0001+0002 is enough to fix the existing reports. I think
temp-schema-cleanup could also use an alternative output, as it uses
the same construction and fails with the above modification.

0003+0004 seems to fix a blocking issue Nikolay mentioned, but I don't
think that's currently required by any of the existing test cases, it
seems more like an AI-discovered scenario.

One thing I realized after sending that message is that if we accept
the limitation in this comment:

+ /*
+ * Save the error before PQgetResult() adds another complaint about
+ * attempting to read from the dead socket.
+ */
+ connection_error = pg_strdup(PQerrorMessage(conn));

Which means an additional "invalid socket" message in the output,
0001+0003 together becomes 3 simple condition changes. See the
attached patch, the only difference is that there's one more extra
line in the alternative test outputs, but now the code change is
simpler, so this might be actually better. (I also validated Nikolay's
0004 test case with this, but I didn't include it in the patch)

This still need actual windows testing, I didn't do that part yet.

Attachments:

  [application/octet-stream] v2-0001-isolationtester-Report-lost-connection-as-step-re.patch (12.2K, ../../CAN4CZFM4iAESOc08pNiA87nboNP35Nj_7z7eXvfS1-g=08KZ7A@mail.gmail.com/2-v2-0001-isolationtester-Report-lost-connection-as-step-re.patch)
  download | inline diff:
From d29eccc7c73d6e1c804c5673eb02d4c0ce841b03 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Tue, 29 Sep 2026 00:11:58 +0000
Subject: [PATCH v2] isolationtester: Report lost connection as step result

isolationtester exits when PQconsumeInput() fails.  On Windows the
FATAL of a terminated backend can get lost, and the connection error
then aborts the permutation, skipping teardown.

libpq already treats a dead connection as not busy, so let the step
complete with the connection error instead.  Add alternative outputs
for the tests that terminate a backend.
---
 .../expected/slot_creation_error_1.out        | 114 +++++++++++++++++
 .../expected/temp-schema-cleanup_1.out        | 119 ++++++++++++++++++
 src/test/isolation/isolationtester.c          |   7 +-
 .../expected/wait_cleanup_1.out               |  87 +++++++++++++
 4 files changed, 324 insertions(+), 3 deletions(-)
 create mode 100644 contrib/test_decoding/expected/slot_creation_error_1.out
 create mode 100644 src/test/isolation/expected/temp-schema-cleanup_1.out
 create mode 100644 src/test/modules/injection_points/expected/wait_cleanup_1.out

diff --git a/contrib/test_decoding/expected/slot_creation_error_1.out b/contrib/test_decoding/expected/slot_creation_error_1.out
new file mode 100644
index 00000000000..041cc547fb3
--- /dev/null
+++ b/contrib/test_decoding/expected/slot_creation_error_1.out
@@ -0,0 +1,114 @@
+Parsed test spec with 2 sessions
+
+starting permutation: s1_b s1_xid s2_init s1_view_slot s1_cancel_s2 s1_view_slot s1_c
+step s1_b: BEGIN;
+step s1_xid: SELECT 'xid' FROM txid_current();
+?column?
+--------
+xid     
+(1 row)
+
+step s2_init: 
+    SELECT 'init' FROM pg_create_logical_replication_slot('slot_creation_error', 'test_decoding');
+ <waiting ...>
+step s1_view_slot: 
+    SELECT slot_name, slot_type, active FROM pg_replication_slots WHERE slot_name = 'slot_creation_error'
+
+slot_name          |slot_type|active
+-------------------+---------+------
+slot_creation_error|logical  |t     
+(1 row)
+
+step s1_cancel_s2: 
+    SELECT pg_cancel_backend(pid)
+    FROM pg_stat_activity
+    WHERE application_name = 'isolation/slot_creation_error/s2';
+ <waiting ...>
+step s2_init: <... completed>
+ERROR:  canceling statement due to user request
+step s1_cancel_s2: <... completed>
+pg_cancel_backend
+-----------------
+t                
+(1 row)
+
+step s1_view_slot: 
+    SELECT slot_name, slot_type, active FROM pg_replication_slots WHERE slot_name = 'slot_creation_error'
+
+slot_name|slot_type|active
+---------+---------+------
+(0 rows)
+
+step s1_c: COMMIT;
+
+starting permutation: s1_b s1_xid s2_init s1_c s1_view_slot s1_drop_slot
+step s1_b: BEGIN;
+step s1_xid: SELECT 'xid' FROM txid_current();
+?column?
+--------
+xid     
+(1 row)
+
+step s2_init: 
+    SELECT 'init' FROM pg_create_logical_replication_slot('slot_creation_error', 'test_decoding');
+ <waiting ...>
+step s1_c: COMMIT;
+step s2_init: <... completed>
+?column?
+--------
+init    
+(1 row)
+
+step s1_view_slot: 
+    SELECT slot_name, slot_type, active FROM pg_replication_slots WHERE slot_name = 'slot_creation_error'
+
+slot_name          |slot_type|active
+-------------------+---------+------
+slot_creation_error|logical  |f     
+(1 row)
+
+step s1_drop_slot: 
+    SELECT pg_drop_replication_slot('slot_creation_error');
+
+pg_drop_replication_slot
+------------------------
+                        
+(1 row)
+
+
+starting permutation: s1_b s1_xid s2_init s1_terminate_s2 s1_c s1_view_slot
+step s1_b: BEGIN;
+step s1_xid: SELECT 'xid' FROM txid_current();
+?column?
+--------
+xid     
+(1 row)
+
+step s2_init: 
+    SELECT 'init' FROM pg_create_logical_replication_slot('slot_creation_error', 'test_decoding');
+ <waiting ...>
+step s1_terminate_s2: 
+    SELECT pg_terminate_backend(pid)
+    FROM pg_stat_activity
+    WHERE application_name = 'isolation/slot_creation_error/s2';
+ <waiting ...>
+step s2_init: <... completed>
+server closed the connection unexpectedly
+	This probably means the server terminated abnormally
+	before or while processing the request.
+invalid socket
+
+step s1_terminate_s2: <... completed>
+pg_terminate_backend
+--------------------
+t                   
+(1 row)
+
+step s1_c: COMMIT;
+step s1_view_slot: 
+    SELECT slot_name, slot_type, active FROM pg_replication_slots WHERE slot_name = 'slot_creation_error'
+
+slot_name|slot_type|active
+---------+---------+------
+(0 rows)
+
diff --git a/src/test/isolation/expected/temp-schema-cleanup_1.out b/src/test/isolation/expected/temp-schema-cleanup_1.out
new file mode 100644
index 00000000000..5ac3054e6eb
--- /dev/null
+++ b/src/test/isolation/expected/temp-schema-cleanup_1.out
@@ -0,0 +1,119 @@
+Parsed test spec with 2 sessions
+
+starting permutation: s1_create_temp_objects s1_discard_temp s2_check_schema
+step s1_create_temp_objects: 
+
+    -- create function large enough to be toasted, to ensure we correctly clean those up, a prior bug
+    -- https://postgr.es/m/CAOFAq3BU5Mf2TTvu8D9n_ZOoFAeQswuzk7yziAb7xuw_qyw5gw%40mail.gmail.com
+    SELECT exec(format($outer$
+        CREATE OR REPLACE FUNCTION pg_temp.long() RETURNS text LANGUAGE sql AS $body$ SELECT %L; $body$$outer$,
+	(SELECT string_agg(g.i::text||':'||random()::text, '|') FROM generate_series(1, 100) g(i))));
+
+    -- The above bug requires function removal to happen after a catalog
+    -- invalidation. dependency.c sorts objects in descending oid order so
+    -- that newer objects are deleted before older objects, so create a
+    -- table after.
+    CREATE TEMPORARY TABLE invalidate_catalog_cache();
+
+    -- test non-temp function is dropped when depending on temp table
+    CREATE TEMPORARY TABLE just_give_me_a_type(id serial primary key);
+
+    CREATE FUNCTION uses_a_temp_type(just_give_me_a_type) RETURNS int LANGUAGE sql AS $$SELECT 1;$$;
+
+exec
+----
+    
+(1 row)
+
+s1: NOTICE:  function "uses_a_temp_type" will be effectively temporary
+DETAIL:  It depends on temporary type just_give_me_a_type.
+step s1_discard_temp: 
+    DISCARD TEMP;
+
+step s2_check_schema: 
+    SELECT oid::regclass FROM pg_class WHERE relnamespace = (SELECT oid FROM s1_temp_schema);
+    SELECT oid::regproc FROM pg_proc WHERE pronamespace = (SELECT oid FROM s1_temp_schema);
+    SELECT oid::regproc FROM pg_type WHERE typnamespace = (SELECT oid FROM s1_temp_schema);
+
+oid
+---
+(0 rows)
+
+oid
+---
+(0 rows)
+
+oid
+---
+(0 rows)
+
+
+starting permutation: s1_advisory s2_advisory s1_create_temp_objects s1_exit s2_check_schema
+step s1_advisory: 
+    SELECT pg_advisory_lock('pg_namespace'::regclass::int8);
+
+pg_advisory_lock
+----------------
+                
+(1 row)
+
+step s2_advisory: 
+    SELECT pg_advisory_lock('pg_namespace'::regclass::int8);
+ <waiting ...>
+step s1_create_temp_objects: 
+
+    -- create function large enough to be toasted, to ensure we correctly clean those up, a prior bug
+    -- https://postgr.es/m/CAOFAq3BU5Mf2TTvu8D9n_ZOoFAeQswuzk7yziAb7xuw_qyw5gw%40mail.gmail.com
+    SELECT exec(format($outer$
+        CREATE OR REPLACE FUNCTION pg_temp.long() RETURNS text LANGUAGE sql AS $body$ SELECT %L; $body$$outer$,
+	(SELECT string_agg(g.i::text||':'||random()::text, '|') FROM generate_series(1, 100) g(i))));
+
+    -- The above bug requires function removal to happen after a catalog
+    -- invalidation. dependency.c sorts objects in descending oid order so
+    -- that newer objects are deleted before older objects, so create a
+    -- table after.
+    CREATE TEMPORARY TABLE invalidate_catalog_cache();
+
+    -- test non-temp function is dropped when depending on temp table
+    CREATE TEMPORARY TABLE just_give_me_a_type(id serial primary key);
+
+    CREATE FUNCTION uses_a_temp_type(just_give_me_a_type) RETURNS int LANGUAGE sql AS $$SELECT 1;$$;
+
+exec
+----
+    
+(1 row)
+
+s1: NOTICE:  function "uses_a_temp_type" will be effectively temporary
+DETAIL:  It depends on temporary type just_give_me_a_type.
+step s1_exit: 
+    SELECT pg_terminate_backend(pg_backend_pid());
+
+server closed the connection unexpectedly
+	This probably means the server terminated abnormally
+	before or while processing the request.
+invalid socket
+
+step s2_advisory: <... completed>
+pg_advisory_lock
+----------------
+                
+(1 row)
+
+step s2_check_schema: 
+    SELECT oid::regclass FROM pg_class WHERE relnamespace = (SELECT oid FROM s1_temp_schema);
+    SELECT oid::regproc FROM pg_proc WHERE pronamespace = (SELECT oid FROM s1_temp_schema);
+    SELECT oid::regproc FROM pg_type WHERE typnamespace = (SELECT oid FROM s1_temp_schema);
+
+oid
+---
+(0 rows)
+
+oid
+---
+(0 rows)
+
+oid
+---
+(0 rows)
+
diff --git a/src/test/isolation/isolationtester.c b/src/test/isolation/isolationtester.c
index e64dc020b18..66241a633c5 100644
--- a/src/test/isolation/isolationtester.c
+++ b/src/test/isolation/isolationtester.c
@@ -851,7 +851,7 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 		}
 	}
 
-	if (sock < 0)
+	if (sock < 0 && PQstatus(conn) != CONNECTION_BAD)
 	{
 		fprintf(stderr, "invalid socket: %s", PQerrorMessage(conn));
 		exit(1);
@@ -908,7 +908,7 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 					 * returns false, we might as well go examine the
 					 * available result.
 					 */
-					if (!PQconsumeInput(conn))
+					if (!PQconsumeInput(conn) && PQstatus(conn) != CONNECTION_BAD)
 					{
 						fprintf(stderr, "PQconsumeInput failed: %s\n",
 								PQerrorMessage(conn));
@@ -977,7 +977,8 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 				exit(1);
 			}
 		}
-		else if (!PQconsumeInput(conn)) /* select(): data available */
+		else if (!PQconsumeInput(conn) &&
+				 PQstatus(conn) != CONNECTION_BAD)	/* select(): data available */
 		{
 			fprintf(stderr, "PQconsumeInput failed: %s\n",
 					PQerrorMessage(conn));
diff --git a/src/test/modules/injection_points/expected/wait_cleanup_1.out b/src/test/modules/injection_points/expected/wait_cleanup_1.out
new file mode 100644
index 00000000000..10da6bd4fcd
--- /dev/null
+++ b/src/test/modules/injection_points/expected/wait_cleanup_1.out
@@ -0,0 +1,87 @@
+Parsed test spec with 3 sessions
+
+starting permutation: wait1 cancel3 noop3 wait2 wakeup3 noop2 detach3
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step wait1: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step cancel3: 
+	SELECT pg_cancel_backend(pid) FROM pg_stat_activity
+	  WHERE wait_event = 'injection-points-wait';
+ <waiting ...>
+step wait1: <... completed>
+ERROR:  canceling statement due to user request
+step cancel3: <... completed>
+pg_cancel_backend
+-----------------
+t                
+(1 row)
+
+step noop3: 
+step wait2: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step wakeup3: SELECT injection_points_wakeup('injection-points-wait');
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait2: <... completed>
+injection_points_run
+--------------------
+                    
+(1 row)
+
+step noop2: 
+step detach3: SELECT injection_points_detach('injection-points-wait');
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
+
+starting permutation: wait1 terminate3 noop3 wait2 wakeup3 noop2 detach3
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step wait1: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step terminate3: 
+	SELECT pg_terminate_backend(pid) FROM pg_stat_activity
+	  WHERE wait_event = 'injection-points-wait';
+ <waiting ...>
+step wait1: <... completed>
+server closed the connection unexpectedly
+	This probably means the server terminated abnormally
+	before or while processing the request.
+invalid socket
+
+step terminate3: <... completed>
+pg_terminate_backend
+--------------------
+t                   
+(1 row)
+
+step noop3: 
+step wait2: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step wakeup3: SELECT injection_points_wakeup('injection-points-wait');
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait2: <... completed>
+injection_points_run
+--------------------
+                    
+(1 row)
+
+step noop2: 
+step detach3: SELECT injection_points_detach('injection-points-wait');
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
-- 
2.55.0



^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-09-29 11:24  Zsolt Parragi <zsolt.parragi@percona.com>
  parent: Zsolt Parragi <zsolt.parragi@percona.com>
  0 siblings, 1 reply; 21+ messages in thread

From: Zsolt Parragi @ 2026-09-29 11:24 UTC (permalink / raw)
  To: Michael Paquier <michael@paquier.xyz>; +Cc: Nikolay Samokhvalov <nik@postgres.ai>; pgsql-hackers@lists.postgresql.org, Andrey Borodin <x4mmm@yandex-team.ru>

> On windows, it seems it can randomly happen, but it's very unlikely

I ended up debugging for some more on windows, and then looking into
the history of this issue.

The following modification works for testing Andrey's patch or also my
v2, and doesn't use a hard coded timeout, it reproduces the original
reported issue 100% in my tests on windows. It could be still timing
dependent on CI under load, but that just means that sometimes the
test would pass with the original output, not _1.

 step terminate3        {
        SELECT pg_terminate_backend(pid) FROM pg_stat_activity
          WHERE wait_event = 'injection-points-wait';
+       SELECT pg_sleep(0.1);
 }

It doesn't test the remainder issue fixed by Nikolay's 0003 or also by
v2-0001, but it is a simple change. Alternatively, if we do this in a
wait loop we can also reproduce the additional failure reported/fixed
by Nikolay without a hardcoded timeout.

I attached v3, which contains this additional change (the looped version).

I also looked into the history of this, and I think we might solve the
problem instead, at least on master:

6051857fc on 2021-12-02 added closesocket() in socket_close() under
#ifdef WIN32.
ed52c3707 on 2021-12-07 also shutdown(sock, SD_SEND)

these would fix the issue properly, making the alternative outputs and
the isolation tester fix unnecessary, however

75674c7ec on 2022-01-25 reverted in back branches because of issues
with walreciever
29992a6a5 on 2022-03-22 reverted also in master for the same reason

but also

a8458f508a7 on 2024-07-13 fixed the issue causing that instability, so
now it should be safe to revert the revert, at least I can't reproduce
the problem mentioned in the reverts with this fix in place, and I can
without it.
In theory a8458f508a7 was backported everywhere, but that doesn't mean
this change would be completely risk-free.

Attachments:

  [application/octet-stream] v3-0001-isolationtester-Report-lost-connection-as-step-re.patch (17.3K, ../../CAN4CZFMtxYXiJV+eAA7XmYx4RDU1SFnzFq1UB=aoWa+GX6hFug@mail.gmail.com/2-v3-0001-isolationtester-Report-lost-connection-as-step-re.patch)
  download | inline diff:
From a04813acfaf06be576ae0c83a063755ddd04bfaa Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Tue, 29 Sep 2026 12:11:46 +0100
Subject: [PATCH v3] isolationtester: Report lost connection as step result

isolationtester exited when PQconsumeInput() failed.  On Windows, the
FATAL of a terminated backend can get lost, as the backend's exit
resets the connection and the client discards data it has not read
yet.  The resulting connection error aborted the permutation, skipping
teardown, which also broke later tests in the same run.

libpq treats a dead connection as not busy, so let the step complete
with the connection error instead.  Add alternative outputs for the
tests that terminate a backend.  wait_cleanup now also holds the
terminated step until the backend has exited, so that on Windows it
covers reporting a lost connection on a later retry of a blocked step.

Author: Zsolt Parragi <zsolt.parragi@percona.com>
Author: Andrey Borodin <x4mmm@yandex-team.ru>
Author: Nikolay Samokhvalov <nik@postgres.ai>
Discussion: https://postgr.es/m/088AB35D-0860-494A-A9CF-8B301543AA51@yandex-team.ru
---
 .../expected/slot_creation_error_1.out        | 114 +++++++++++++++++
 .../expected/temp-schema-cleanup_1.out        | 119 ++++++++++++++++++
 src/test/isolation/isolationtester.c          |   7 +-
 .../expected/wait_cleanup.out                 |  27 ++--
 .../expected/wait_cleanup_1.out               |  96 ++++++++++++++
 .../injection_points/specs/wait_cleanup.spec  |  27 +++-
 6 files changed, 373 insertions(+), 17 deletions(-)
 create mode 100644 contrib/test_decoding/expected/slot_creation_error_1.out
 create mode 100644 src/test/isolation/expected/temp-schema-cleanup_1.out
 create mode 100644 src/test/modules/injection_points/expected/wait_cleanup_1.out

diff --git a/contrib/test_decoding/expected/slot_creation_error_1.out b/contrib/test_decoding/expected/slot_creation_error_1.out
new file mode 100644
index 00000000000..041cc547fb3
--- /dev/null
+++ b/contrib/test_decoding/expected/slot_creation_error_1.out
@@ -0,0 +1,114 @@
+Parsed test spec with 2 sessions
+
+starting permutation: s1_b s1_xid s2_init s1_view_slot s1_cancel_s2 s1_view_slot s1_c
+step s1_b: BEGIN;
+step s1_xid: SELECT 'xid' FROM txid_current();
+?column?
+--------
+xid     
+(1 row)
+
+step s2_init: 
+    SELECT 'init' FROM pg_create_logical_replication_slot('slot_creation_error', 'test_decoding');
+ <waiting ...>
+step s1_view_slot: 
+    SELECT slot_name, slot_type, active FROM pg_replication_slots WHERE slot_name = 'slot_creation_error'
+
+slot_name          |slot_type|active
+-------------------+---------+------
+slot_creation_error|logical  |t     
+(1 row)
+
+step s1_cancel_s2: 
+    SELECT pg_cancel_backend(pid)
+    FROM pg_stat_activity
+    WHERE application_name = 'isolation/slot_creation_error/s2';
+ <waiting ...>
+step s2_init: <... completed>
+ERROR:  canceling statement due to user request
+step s1_cancel_s2: <... completed>
+pg_cancel_backend
+-----------------
+t                
+(1 row)
+
+step s1_view_slot: 
+    SELECT slot_name, slot_type, active FROM pg_replication_slots WHERE slot_name = 'slot_creation_error'
+
+slot_name|slot_type|active
+---------+---------+------
+(0 rows)
+
+step s1_c: COMMIT;
+
+starting permutation: s1_b s1_xid s2_init s1_c s1_view_slot s1_drop_slot
+step s1_b: BEGIN;
+step s1_xid: SELECT 'xid' FROM txid_current();
+?column?
+--------
+xid     
+(1 row)
+
+step s2_init: 
+    SELECT 'init' FROM pg_create_logical_replication_slot('slot_creation_error', 'test_decoding');
+ <waiting ...>
+step s1_c: COMMIT;
+step s2_init: <... completed>
+?column?
+--------
+init    
+(1 row)
+
+step s1_view_slot: 
+    SELECT slot_name, slot_type, active FROM pg_replication_slots WHERE slot_name = 'slot_creation_error'
+
+slot_name          |slot_type|active
+-------------------+---------+------
+slot_creation_error|logical  |f     
+(1 row)
+
+step s1_drop_slot: 
+    SELECT pg_drop_replication_slot('slot_creation_error');
+
+pg_drop_replication_slot
+------------------------
+                        
+(1 row)
+
+
+starting permutation: s1_b s1_xid s2_init s1_terminate_s2 s1_c s1_view_slot
+step s1_b: BEGIN;
+step s1_xid: SELECT 'xid' FROM txid_current();
+?column?
+--------
+xid     
+(1 row)
+
+step s2_init: 
+    SELECT 'init' FROM pg_create_logical_replication_slot('slot_creation_error', 'test_decoding');
+ <waiting ...>
+step s1_terminate_s2: 
+    SELECT pg_terminate_backend(pid)
+    FROM pg_stat_activity
+    WHERE application_name = 'isolation/slot_creation_error/s2';
+ <waiting ...>
+step s2_init: <... completed>
+server closed the connection unexpectedly
+	This probably means the server terminated abnormally
+	before or while processing the request.
+invalid socket
+
+step s1_terminate_s2: <... completed>
+pg_terminate_backend
+--------------------
+t                   
+(1 row)
+
+step s1_c: COMMIT;
+step s1_view_slot: 
+    SELECT slot_name, slot_type, active FROM pg_replication_slots WHERE slot_name = 'slot_creation_error'
+
+slot_name|slot_type|active
+---------+---------+------
+(0 rows)
+
diff --git a/src/test/isolation/expected/temp-schema-cleanup_1.out b/src/test/isolation/expected/temp-schema-cleanup_1.out
new file mode 100644
index 00000000000..5ac3054e6eb
--- /dev/null
+++ b/src/test/isolation/expected/temp-schema-cleanup_1.out
@@ -0,0 +1,119 @@
+Parsed test spec with 2 sessions
+
+starting permutation: s1_create_temp_objects s1_discard_temp s2_check_schema
+step s1_create_temp_objects: 
+
+    -- create function large enough to be toasted, to ensure we correctly clean those up, a prior bug
+    -- https://postgr.es/m/CAOFAq3BU5Mf2TTvu8D9n_ZOoFAeQswuzk7yziAb7xuw_qyw5gw%40mail.gmail.com
+    SELECT exec(format($outer$
+        CREATE OR REPLACE FUNCTION pg_temp.long() RETURNS text LANGUAGE sql AS $body$ SELECT %L; $body$$outer$,
+	(SELECT string_agg(g.i::text||':'||random()::text, '|') FROM generate_series(1, 100) g(i))));
+
+    -- The above bug requires function removal to happen after a catalog
+    -- invalidation. dependency.c sorts objects in descending oid order so
+    -- that newer objects are deleted before older objects, so create a
+    -- table after.
+    CREATE TEMPORARY TABLE invalidate_catalog_cache();
+
+    -- test non-temp function is dropped when depending on temp table
+    CREATE TEMPORARY TABLE just_give_me_a_type(id serial primary key);
+
+    CREATE FUNCTION uses_a_temp_type(just_give_me_a_type) RETURNS int LANGUAGE sql AS $$SELECT 1;$$;
+
+exec
+----
+    
+(1 row)
+
+s1: NOTICE:  function "uses_a_temp_type" will be effectively temporary
+DETAIL:  It depends on temporary type just_give_me_a_type.
+step s1_discard_temp: 
+    DISCARD TEMP;
+
+step s2_check_schema: 
+    SELECT oid::regclass FROM pg_class WHERE relnamespace = (SELECT oid FROM s1_temp_schema);
+    SELECT oid::regproc FROM pg_proc WHERE pronamespace = (SELECT oid FROM s1_temp_schema);
+    SELECT oid::regproc FROM pg_type WHERE typnamespace = (SELECT oid FROM s1_temp_schema);
+
+oid
+---
+(0 rows)
+
+oid
+---
+(0 rows)
+
+oid
+---
+(0 rows)
+
+
+starting permutation: s1_advisory s2_advisory s1_create_temp_objects s1_exit s2_check_schema
+step s1_advisory: 
+    SELECT pg_advisory_lock('pg_namespace'::regclass::int8);
+
+pg_advisory_lock
+----------------
+                
+(1 row)
+
+step s2_advisory: 
+    SELECT pg_advisory_lock('pg_namespace'::regclass::int8);
+ <waiting ...>
+step s1_create_temp_objects: 
+
+    -- create function large enough to be toasted, to ensure we correctly clean those up, a prior bug
+    -- https://postgr.es/m/CAOFAq3BU5Mf2TTvu8D9n_ZOoFAeQswuzk7yziAb7xuw_qyw5gw%40mail.gmail.com
+    SELECT exec(format($outer$
+        CREATE OR REPLACE FUNCTION pg_temp.long() RETURNS text LANGUAGE sql AS $body$ SELECT %L; $body$$outer$,
+	(SELECT string_agg(g.i::text||':'||random()::text, '|') FROM generate_series(1, 100) g(i))));
+
+    -- The above bug requires function removal to happen after a catalog
+    -- invalidation. dependency.c sorts objects in descending oid order so
+    -- that newer objects are deleted before older objects, so create a
+    -- table after.
+    CREATE TEMPORARY TABLE invalidate_catalog_cache();
+
+    -- test non-temp function is dropped when depending on temp table
+    CREATE TEMPORARY TABLE just_give_me_a_type(id serial primary key);
+
+    CREATE FUNCTION uses_a_temp_type(just_give_me_a_type) RETURNS int LANGUAGE sql AS $$SELECT 1;$$;
+
+exec
+----
+    
+(1 row)
+
+s1: NOTICE:  function "uses_a_temp_type" will be effectively temporary
+DETAIL:  It depends on temporary type just_give_me_a_type.
+step s1_exit: 
+    SELECT pg_terminate_backend(pg_backend_pid());
+
+server closed the connection unexpectedly
+	This probably means the server terminated abnormally
+	before or while processing the request.
+invalid socket
+
+step s2_advisory: <... completed>
+pg_advisory_lock
+----------------
+                
+(1 row)
+
+step s2_check_schema: 
+    SELECT oid::regclass FROM pg_class WHERE relnamespace = (SELECT oid FROM s1_temp_schema);
+    SELECT oid::regproc FROM pg_proc WHERE pronamespace = (SELECT oid FROM s1_temp_schema);
+    SELECT oid::regproc FROM pg_type WHERE typnamespace = (SELECT oid FROM s1_temp_schema);
+
+oid
+---
+(0 rows)
+
+oid
+---
+(0 rows)
+
+oid
+---
+(0 rows)
+
diff --git a/src/test/isolation/isolationtester.c b/src/test/isolation/isolationtester.c
index e64dc020b18..66241a633c5 100644
--- a/src/test/isolation/isolationtester.c
+++ b/src/test/isolation/isolationtester.c
@@ -851,7 +851,7 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 		}
 	}
 
-	if (sock < 0)
+	if (sock < 0 && PQstatus(conn) != CONNECTION_BAD)
 	{
 		fprintf(stderr, "invalid socket: %s", PQerrorMessage(conn));
 		exit(1);
@@ -908,7 +908,7 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 					 * returns false, we might as well go examine the
 					 * available result.
 					 */
-					if (!PQconsumeInput(conn))
+					if (!PQconsumeInput(conn) && PQstatus(conn) != CONNECTION_BAD)
 					{
 						fprintf(stderr, "PQconsumeInput failed: %s\n",
 								PQerrorMessage(conn));
@@ -977,7 +977,8 @@ try_complete_step(TestSpec *testspec, PermutationStep *pstep, int flags)
 				exit(1);
 			}
 		}
-		else if (!PQconsumeInput(conn)) /* select(): data available */
+		else if (!PQconsumeInput(conn) &&
+				 PQstatus(conn) != CONNECTION_BAD)	/* select(): data available */
 		{
 			fprintf(stderr, "PQconsumeInput failed: %s\n",
 					PQerrorMessage(conn));
diff --git a/src/test/modules/injection_points/expected/wait_cleanup.out b/src/test/modules/injection_points/expected/wait_cleanup.out
index c5be17428fc..e7300508999 100644
--- a/src/test/modules/injection_points/expected/wait_cleanup.out
+++ b/src/test/modules/injection_points/expected/wait_cleanup.out
@@ -41,7 +41,7 @@ injection_points_detach
 (1 row)
 
 
-starting permutation: wait1 terminate3 noop3 wait2 wakeup3 noop2 detach3
+starting permutation: wait1 terminate3 release3 wait2 wakeup3 noop2 detach3
 injection_points_attach
 -----------------------
                        
@@ -51,20 +51,29 @@ step wait1: SELECT injection_points_run('injection-points-wait'); <waiting ...>
 step terminate3: 
 	SELECT pg_terminate_backend(pid) FROM pg_stat_activity
 	  WHERE wait_event = 'injection-points-wait';
- <waiting ...>
-step wait1: <... completed>
-FATAL:  terminating connection due to administrator command
-server closed the connection unexpectedly
-	This probably means the server terminated abnormally
-	before or while processing the request.
+	DO $$
+	BEGIN
+	  WHILE EXISTS (SELECT FROM pg_stat_activity
+					WHERE application_name = 'isolation/wait_cleanup/s1')
+	  LOOP
+		PERFORM pg_sleep(0.01);
+		PERFORM pg_stat_clear_snapshot();
+	  END LOOP;
+	END$$;
 
-step terminate3: <... completed>
 pg_terminate_backend
 --------------------
 t                   
 (1 row)
 
-step noop3: 
+s3: NOTICE:  release wait1
+step release3: DO $$BEGIN RAISE NOTICE 'release wait1'; END$$;
+step wait1: <... completed>
+FATAL:  terminating connection due to administrator command
+server closed the connection unexpectedly
+	This probably means the server terminated abnormally
+	before or while processing the request.
+
 step wait2: SELECT injection_points_run('injection-points-wait'); <waiting ...>
 step wakeup3: SELECT injection_points_wakeup('injection-points-wait');
 injection_points_wakeup
diff --git a/src/test/modules/injection_points/expected/wait_cleanup_1.out b/src/test/modules/injection_points/expected/wait_cleanup_1.out
new file mode 100644
index 00000000000..19f7ff22f20
--- /dev/null
+++ b/src/test/modules/injection_points/expected/wait_cleanup_1.out
@@ -0,0 +1,96 @@
+Parsed test spec with 3 sessions
+
+starting permutation: wait1 cancel3 noop3 wait2 wakeup3 noop2 detach3
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step wait1: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step cancel3: 
+	SELECT pg_cancel_backend(pid) FROM pg_stat_activity
+	  WHERE wait_event = 'injection-points-wait';
+ <waiting ...>
+step wait1: <... completed>
+ERROR:  canceling statement due to user request
+step cancel3: <... completed>
+pg_cancel_backend
+-----------------
+t                
+(1 row)
+
+step noop3: 
+step wait2: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step wakeup3: SELECT injection_points_wakeup('injection-points-wait');
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait2: <... completed>
+injection_points_run
+--------------------
+                    
+(1 row)
+
+step noop2: 
+step detach3: SELECT injection_points_detach('injection-points-wait');
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
+
+starting permutation: wait1 terminate3 release3 wait2 wakeup3 noop2 detach3
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step wait1: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step terminate3: 
+	SELECT pg_terminate_backend(pid) FROM pg_stat_activity
+	  WHERE wait_event = 'injection-points-wait';
+	DO $$
+	BEGIN
+	  WHILE EXISTS (SELECT FROM pg_stat_activity
+					WHERE application_name = 'isolation/wait_cleanup/s1')
+	  LOOP
+		PERFORM pg_sleep(0.01);
+		PERFORM pg_stat_clear_snapshot();
+	  END LOOP;
+	END$$;
+
+pg_terminate_backend
+--------------------
+t                   
+(1 row)
+
+s3: NOTICE:  release wait1
+step release3: DO $$BEGIN RAISE NOTICE 'release wait1'; END$$;
+step wait1: <... completed>
+server closed the connection unexpectedly
+	This probably means the server terminated abnormally
+	before or while processing the request.
+invalid socket
+
+step wait2: SELECT injection_points_run('injection-points-wait'); <waiting ...>
+step wakeup3: SELECT injection_points_wakeup('injection-points-wait');
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait2: <... completed>
+injection_points_run
+--------------------
+                    
+(1 row)
+
+step noop2: 
+step detach3: SELECT injection_points_detach('injection-points-wait');
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/wait_cleanup.spec b/src/test/modules/injection_points/specs/wait_cleanup.spec
index ed7d21c4de4..54ec5fe0ef8 100644
--- a/src/test/modules/injection_points/specs/wait_cleanup.spec
+++ b/src/test/modules/injection_points/specs/wait_cleanup.spec
@@ -26,10 +26,11 @@ session s2
 step wait2	{ SELECT injection_points_run('injection-points-wait'); }
 step noop2	{ }
 
-# Control session.  The blocker annotations on cancel3/terminate3,
-# together with noop3, make the tester wait until wait1 has fully
-# completed before starting wait2.  Otherwise, wait2 could register a
-# new waiter slot while s1 still owns the previous one.
+# Control session.  The blocker annotation on cancel3, together with
+# noop3, makes the tester wait until wait1 has fully completed before
+# starting wait2.  Otherwise, wait2 could register a new waiter slot
+# while s1 still owns the previous one.  In the terminate permutation,
+# terminate3 waits for s1's backend to exit instead.
 session s3
 step cancel3	{
 	SELECT pg_cancel_backend(pid) FROM pg_stat_activity
@@ -38,13 +39,29 @@ step cancel3	{
 step terminate3	{
 	SELECT pg_terminate_backend(pid) FROM pg_stat_activity
 	  WHERE wait_event = 'injection-points-wait';
+	DO $$
+	BEGIN
+	  WHILE EXISTS (SELECT FROM pg_stat_activity
+					WHERE application_name = 'isolation/wait_cleanup/s1')
+	  LOOP
+		PERFORM pg_sleep(0.01);
+		PERFORM pg_stat_clear_snapshot();
+	  END LOOP;
+	END$$;
 }
 step wakeup3	{ SELECT injection_points_wakeup('injection-points-wait'); }
 step detach3	{ SELECT injection_points_detach('injection-points-wait'); }
+step release3	{ DO $$BEGIN RAISE NOTICE 'release wait1'; END$$; }
 step noop3	{ }
 
 permutation wait1 cancel3(wait1) noop3 wait2 wakeup3 noop2 detach3
 
 # The terminate permutation has to stay last: s1's connection is dead
 # afterwards, and the tester never reconnects a session.
-permutation wait1 terminate3(wait1) noop3 wait2 wakeup3 noop2 detach3
+#
+# On Windows, the FATAL of a terminated backend can get lost, as the
+# backend's exit resets the connection and unread data is discarded.
+# Holding wait1 until release3 makes the tester notice the lost
+# connection while the step is still blocked, and report it on a later
+# retry, see wait_cleanup_1.out.
+permutation wait1(release3 notices 1) terminate3 release3 wait2 wakeup3 noop2 detach3
-- 
2.43.0



^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-09-29 23:46  Michael Paquier <michael@paquier.xyz>
  parent: Zsolt Parragi <zsolt.parragi@percona.com>
  0 siblings, 2 replies; 21+ messages in thread

From: Michael Paquier @ 2026-09-29 23:46 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Nikolay Samokhvalov <nik@postgres.ai>; pgsql-hackers@lists.postgresql.org, Andrey Borodin <x4mmm@yandex-team.ru>

On Tue, Sep 29, 2026 at 12:24:46PM +0100, Zsolt Parragi wrote:
> a8458f508a7 on 2024-07-13 fixed the issue causing that instability, so
> now it should be safe to revert the revert, at least I can't reproduce
> the problem mentioned in the reverts with this fix in place, and I can
> without it.
> In theory a8458f508a7 was backported everywhere, but that doesn't mean
> this change would be completely risk-free.

Thanks for looking at some history and putting some pieces together.
I was not aware of this stuff. 

In this context, the revert of the revert can be translated as `git
revert 29992a6a509b`, right?  If we do that, being able to get rid of
the alternate outputs would be super nice, and we would not even need
to have alternate outputs like that:
--- temp-schema-cleanup.out	2026-07-27 08:05:11.192854428 +0900
+++ temp-schema-cleanup_1.out	2026-09-30 08:03:16.982528616 +0900
@@ -89,10 +89,10 @@ DETAIL:  It depends on temporary type ju
 step s1_exit: 
     SELECT pg_terminate_backend(pg_backend_pid());
 
-FATAL:  terminating connection due to administrator command
 server closed the connection unexpectedly
 	This probably means the server terminated abnormally
 	before or while processing the request.
+invalid socket
 
 step s2_advisory: <... completed>
 pg_advisory_lock

This points to the fact that losing the messages is wrong, because the
tests do not check what we want them to do in some environments as the
data we want is not received in the backends.  And we would not even
need the isolationtester.c tweaks, assuming that things work out and
that we can rely on the fact that the backend has sent all its
messages, would we?

+    DO $$
+    BEGIN
+      WHILE EXISTS (SELECT FROM pg_stat_activity
+                    WHERE application_name = 'isolation/wait_cleanup/s1')
+      LOOP
+        PERFORM pg_sleep(0.01);
+        PERFORM pg_stat_clear_snapshot();
+      END LOOP;
+    END$$;

Even that feels like the wrong thing to do, spreading a tweak that
ought to be simpler for folks implement tests.

In short, it means that the proposal is papering over the actual
Windows problem: pqcomm.c should really close that socket, so my issue
here is that everybody has lost track of the inter-dependency between
all these issues; pieces that you are just putting together.  I would
not mind experimenting with a revert of 29992a6a509b on HEAD, at
least, and give it some time to brew before deciding what to do with
the stable branches.

Another option that would be on the table for me would be to make
these tests conditional, not running on WIN32 where we know they're
unstable.  We cannot do that within schedule files so
temp-schema-cleanup would be an issue if not moved out somewhere else
(just test_misc with a conditional ISOLATION list?), and I don't
remember a way to control that cleanly at SQL level..  We could do a
solution based on a test list filtering in test_decoding and
injection_points, at least, taking take of two instabilities out of
three.

The perfect scenario for me would be to prove that undoing
29992a6a509b is now really-absolutely-stable safe, as it's still a
server bug to me to not send back this information back to the client
on WIN32.

What do you think?
--
Michael

Attachments:

  [application/pgp-signature] signature.asc (832B, ../../arxN1TpPi75XHGRb@paquier.xyz/2-signature.asc)
  download

^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-09-30 12:31  Zsolt Parragi <zsolt.parragi@percona.com>
  parent: Michael Paquier <michael@paquier.xyz>
  1 sibling, 2 replies; 21+ messages in thread

From: Zsolt Parragi @ 2026-09-30 12:31 UTC (permalink / raw)
  To: Michael Paquier <michael@paquier.xyz>; +Cc: pgsql-hackers@lists.postgresql.org, Nikolay Samokhvalov <nik@postgres.ai>

> In this context, the revert of the revert can be translated as `git
> revert 29992a6a509b`, right?  If we do that, being able to get rid of
> the alternate outputs would be super nice, and we would not even need
> to have alternate outputs like that:

Yes. If we revert 29992a6a509b, I can't reproduce any of the issues
anymore. We can still make a test modification if we want to
explicitly test for this scenario, to make sure it doesn't reappear,
but we won't need the alternative outputs or the isolationtester
modifications.

> +    DO $$
> +    BEGIN
> +      WHILE EXISTS (SELECT FROM pg_stat_activity
> +                    WHERE application_name = 'isolation/wait_cleanup/s1')
> +      LOOP
> +        PERFORM pg_sleep(0.01);
> +        PERFORM pg_stat_clear_snapshot();
> +      END LOOP;
> +    END$$;
>
> Even that feels like the wrong thing to do, spreading a tweak that
> ought to be simpler for folks implement tests.

We don't have to spread this around, I added this to one scenario to
explicitly test the missing last message issue. This, or the simpler
single pg_sleep call makes it deterministic. This is the part either
in this form or just as the one line pg_sleep addition that might be
worth keeping even in the 29992a6a509b direction, so we notice if the
issue comes back / still happens sometimes.

> What do you think?

I agree with the let's try the revert on master approach. That won't
help with random failures on the stable branches, but at least we are
aware why it is happening now, and later we can either apply the
revert on them, or disable these tests on them, or apply v3.

> The perfect scenario for me would be to prove that undoing
> 29992a6a509b is now really-absolutely-stable safe, as it's still a
> server bug to me to not send back this information back to the client
> on WIN32.

The question is, what would be good enough proof? I can reproduce the
walreceiver issue with around 1% failure rate on my laptop, and it
didn't reproduce even once with only reverting 29992a6a509b in more
than 5000 runs.
I'll try to do the same thing on back branches, and I'll also set it
up on github actions to test it there.






^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-09-30 16:14  Kacper Kuras <kacperkuras@hotmail.com>
  parent: Zsolt Parragi <zsolt.parragi@percona.com>
  1 sibling, 0 replies; 21+ messages in thread

From: Kacper Kuras @ 2026-09-30 16:14 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: pgsql-hackers@lists.postgresql.org <pgsql-hackers@lists.postgresql.org>; Nikolay Samokhvalov <nik@postgres.ai>; Thomas Munro <thomas.munro@gmail.com>; Michael Paquier <michael@paquier.xyz>; Andrey Borodin <x4mmm@yandex-team.ru>

On Wed, Sep 30, 2026 at 12:31:53PM +0100, Zsolt Parragi wrote:
> The question is, what would be good enough proof?

I got to the same conclusion independently, from the client side:
clients on Windows get "connection reset" instead of the server's
FATAL, and there's nothing a client can do about it, because by the
time it reads, the data is already gone.  Here are my results, in case
another machine helps.

Setup: Windows 11 Pro 25H2 (build 26200), MSVC 19.50, meson, OpenSSL
3.6.2, master at 9510a826e4a, all over localhost.  "Patched" means
master plus the revert of 29992a6a509b.

1. wait_cleanup with your pg_sleep(0.1) after pg_terminate_backend(),
   and the rest of the injection_points isolation suite, 20 runs each:

     unpatched: 20/20 lose the FATAL ("PQconsumeInput failed: server
                closed the connection unexpectedly"), as in Andrey's CI
     patched:   20/20 get the FATAL, and the suite passes every time

2. Startup FATAL.  The attached script sends a startup packet for a
   role that doesn't exist and sleeps before reading; 50 attempts per
   delay:

     delay before read   unpatched   patched
     0 ms                        50/50           50/50
     50 ms                      4/50             50/50
     300 ms                    0/50             50/50

3. The tests that hung in 2022 (commit_ts/002_standby,
   commit_ts/003_standby_2, recovery/001_stream_rep), patched, 20 runs
   each: all 60 passed, no hangs.

4. SSL, patched: ssl/001-004, 20 runs each, all passed.  Alexander
   reported in [1] that the revoked-client-cert case in 001_ssltests.pl
   sometimes got "Software caused connection abort" with the earlier
   patch set, so I also looped just that case 2000 times on both builds.
   It reported "certificate revoked" every time on each.

5. A full meson test run, patched, with PG_TEST_EXTRA=ssl but without
   injection points (those are covered by 1).  Everything passed except
   pg_test_timing/001_basic and psql/001_basic, which fail here because
   of the Polish locale's decimal comma, not because of the patch.

Not tested: a connection that isn't over localhost, and the back
branches.

Two more pieces of history that may help: Thomas already proposed
re-committing 6051857fc on master in March 2025 [2], and nobody
objected, but it didn't happen.  And there's a remaining walreceiver
hang on WSAECONNRESET that a8458f508 doesn't cover [3], but it happens
without the revert too, so I don't think it's related.

So +1 for trying the revert on master.

[1] https://postgr.es/m/32d112ee-0b6f-d4ab-441b-e2bba66a1d83@gmail.com
[2] https://postgr.es/m/CA+hUKGKJSOAdAukP4QTkR3-jFws39+8C197XC-a97dgYr=cdBA@mail.gmail.com
[3] https://postgr.es/m/93515d62-edce-9041-ec6e-7122f6e92bea@gmail.com

--
Kacper Kuras=

Attachments:

  [application/octet-stream] repro_startup_fatal.pl (950B, ../../VI0P193MB311184FC0AC2815008C46C4BBF8B2@VI0P193MB3111.EURP193.PROD.OUTLOOK.COM/2-repro_startup_fatal.pl)
  download

^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-09-30 18:00  Andrey Borodin <x4mmm@yandex-team.ru>
  parent: Michael Paquier <michael@paquier.xyz>
  1 sibling, 1 reply; 21+ messages in thread

From: Andrey Borodin @ 2026-09-30 18:00 UTC (permalink / raw)
  To: Michael Paquier <michael@paquier.xyz>; +Cc: Zsolt Parragi <zsolt.parragi@percona.com>; Nikolay Samokhvalov <nik@postgres.ai>; pgsql-hackers@lists.postgresql.org



> On 30 Sep 2026, at 04:46, Michael Paquier <michael@paquier.xyz> wrote:
> 
> 29992a6a509b

Wow, what a deep story for just a couple of lines!

+1 for trying to revert in master.

I think there's more benefit than possible harm. There are ~4 similar
failures a month on master. If we resurrect walreceiver problems, we will
notice soon. There were like ~20 failures a month. And a8458f508a7 commit
message looks promising (I do not understand commit itself).


Best regards, Andrey Borodin.






^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-09-30 22:55  Michael Paquier <michael@paquier.xyz>
  parent: Zsolt Parragi <zsolt.parragi@percona.com>
  1 sibling, 0 replies; 21+ messages in thread

From: Michael Paquier @ 2026-09-30 22:55 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: pgsql-hackers@lists.postgresql.org, Nikolay Samokhvalov <nik@postgres.ai>

On Wed, Sep 30, 2026 at 05:31:53AM -0700, Zsolt Parragi wrote:
> I agree with the let's try the revert on master approach. That won't
> help with random failures on the stable branches, but at least we are
> aware why it is happening now, and later we can either apply the
> revert on them, or disable these tests on them, or apply v3.

It would mean that the spurious failures would still be annoying on
stable branches while we evaluate the HEAD case.  See below.

> The question is, what would be good enough proof? I can reproduce the
> walreceiver issue with around 1% failure rate on my laptop, and it
> didn't reproduce even once with only reverting 29992a6a509b in more
> than 5000 runs.
> I'll try to do the same thing on back branches, and I'll also set it
> up on github actions to test it there.

How about the following plan, instead?  I would suggest moving forward
with the following steps:
- First disable these tests on WIN32, backpatch this change to
entirely silence the buildfarm.  I am not really sure that we gain
much coverage by keeping them while the non-WIN32 paths would still
stress the cases in a stable manner.  In order to do that, some
Makefile and meson.build can be manipulated to make the tests
conditional depending on the platform, something that we already do
that.  One small-ish issue here is temp-schema-cleanup in
src/test/isolation/.  Let's just move that to test_misc, I guess
(cannot think of a better location), create an ISOLATION target list
in a conditional manner.
- Revert the revert on HEAD, allow the tests to work on WIN32.
- Monitor the result for a few days or weeks.  Investigate if more
actions are required and what to do in stable branches, and if we
actually want to do something in stable branches.  My gut feeling is
that we'd do that only on HEAD, this has been reverted a lot already.

The meson and Makefile tricks already exist in the tree:
* For meson "if host_system == 'windows'".
* For Makefile, that should be "ifeq ($(PORTNAME), win32)".

This plan is risk-free: even if step 1 is reverted, we still have the
benefit of not running these unstable tests on WIN32, keeping the
buildfarm stable in the long run anyway.

So, thoughts?
--
Michael

Attachments:

  [application/pgp-signature] signature.asc (832B, ../../ar2TeDSQIuQJknfn@paquier.xyz/2-signature.asc)
  download

^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-09-30 23:05  Michael Paquier <michael@paquier.xyz>
  parent: Andrey Borodin <x4mmm@yandex-team.ru>
  0 siblings, 1 reply; 21+ messages in thread

From: Michael Paquier @ 2026-09-30 23:05 UTC (permalink / raw)
  To: Andrey Borodin <x4mmm@yandex-team.ru>; +Cc: Zsolt Parragi <zsolt.parragi@percona.com>; Nikolay Samokhvalov <nik@postgres.ai>; pgsql-hackers@lists.postgresql.org

On Wed, Sep 30, 2026 at 11:00:26PM +0500, Andrey Borodin wrote:
> I think there's more benefit than possible harm. There are ~4 similar
> failures a month on master. If we resurrect walreceiver problems, we will
> notice soon. There were like ~20 failures a month. And a8458f508a7 commit
> message looks promising (I do not understand commit itself).

Well, as mentioned in this commit, peeking at a socket before sleeping
is qualified as a kludge, probably not the best long-term approach.  I
doubt that a lot of people are motivated to redesign the event vs
socket handling just to make for a better socket implementation of
Postgres on WIN32, and I suspect that relying on Thomas's mage skills
should be good enough.  FWIW, I trust his skills, and also I have no
plans to dive into this leel of details for the WIN32 issues; there's
more than enough going on. 

Now I don't mind running the revert/revert experiment and see how it
goes, at the condition that we make all the branches more stable as a
starting point.  Bypassing these tests on WIN32 is part of the plan
I'd be OK with.
--
Michael

Attachments:

  [application/pgp-signature] signature.asc (832B, ../../ar2VsK9d-UJiC7x7@paquier.xyz/2-signature.asc)
  download

^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-10-02 21:44  Zsolt Parragi <zsolt.parragi@percona.com>
  parent: Michael Paquier <michael@paquier.xyz>
  0 siblings, 2 replies; 21+ messages in thread

From: Zsolt Parragi @ 2026-10-02 21:44 UTC (permalink / raw)
  To: Michael Paquier <michael@paquier.xyz>; +Cc: Andrey Borodin <x4mmm@yandex-team.ru>; Nikolay Samokhvalov <nik@postgres.ai>; pgsql-hackers@lists.postgresql.org

I attached v4.

I only disabled the two tests that already had failures on CI. After
looking into it again I realized that temp-schema-cleanup could only
fail with my hacked together linux repro, I can't reproduce it on
windows. In wait_cleanup and slot_creation_error one session kills
another session. In temp-schema-cleanup s1 kills itself. When the
fatal happens the isolation tester is reading s1's socket, and
receives the message properly, I can't make that fail without
modifying the testcase.

Attachments:

  [application/octet-stream] v4-0001-Skip-on-Windows-isolation-tests-that-terminate-ot.patch (4.5K, ../../CAN4CZFOoiKa9E4Gnh_OzLk3S5Z6RFyg_oY1e_WddCVYhkpTVow@mail.gmail.com/2-v4-0001-Skip-on-Windows-isolation-tests-that-terminate-ot.patch)
  download | inline diff:
From 9bbb65bcd3506beee54324387801a9a8e62e77f1 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <dutow@mentalstatic.info>
Date: Fri, 2 Oct 2026 22:10:18 +0100
Subject: [PATCH v4] Skip on Windows isolation tests that terminate other
 backends

On Windows, a backend exiting resets its connection, and the client
discards any data not yet read, including the FATAL message of a
terminated backend.  isolationtester polls one connection at a time, so
when it waits on a different session the message can get lost, failing
the permutation and the tests that run after it.

wait_cleanup in injection_points and slot_creation_error in
test_decoding have this pattern, and have been failing randomly in the
buildfarm.  Let's skip them on Windows.

Discussion: https://postgr.es/m/088AB35D-0860-494A-A9CF-8B301543AA51@yandex-team.ru
---
 contrib/test_decoding/Makefile                |  7 +++++++
 contrib/test_decoding/meson.build             | 11 +++++++++--
 src/test/modules/injection_points/Makefile    |  7 +++++++
 src/test/modules/injection_points/meson.build | 11 +++++++++--
 4 files changed, 32 insertions(+), 4 deletions(-)

diff --git a/contrib/test_decoding/Makefile b/contrib/test_decoding/Makefile
index 0111124399a..322e7dc185f 100644
--- a/contrib/test_decoding/Makefile
+++ b/contrib/test_decoding/Makefile
@@ -31,6 +31,13 @@ include $(top_builddir)/src/Makefile.global
 include $(top_srcdir)/contrib/contrib-global.mk
 endif
 
+# slot_creation_error is unstable on Windows, where the FATAL message of a
+# terminated backend can get lost.  See
+# https://postgr.es/m/088AB35D-0860-494A-A9CF-8B301543AA51@yandex-team.ru
+ifeq ($(PORTNAME), win32)
+ISOLATION := $(filter-out slot_creation_error,$(ISOLATION))
+endif
+
 # But it can nonetheless be very helpful to run tests on preexisting
 # installation, allow to do so, but only if requested explicitly.
 installcheck-force:
diff --git a/contrib/test_decoding/meson.build b/contrib/test_decoding/meson.build
index ac655853d26..08efa7b5d99 100644
--- a/contrib/test_decoding/meson.build
+++ b/contrib/test_decoding/meson.build
@@ -16,6 +16,14 @@ test_decoding = shared_module('test_decoding',
 )
 contrib_targets += test_decoding
 
+# slot_creation_error is unstable on Windows, where the FATAL message of a
+# terminated backend can get lost.  See
+# https://postgr.es/m/088AB35D-0860-494A-A9CF-8B301543AA51@yandex-team.ru
+test_decoding_extra_specs = []
+if host_system != 'windows'
+  test_decoding_extra_specs += 'slot_creation_error'
+endif
+
 tests += {
   'name': 'test_decoding',
   'sd': meson.current_source_dir(),
@@ -62,11 +70,10 @@ tests += {
       'subxact_without_top',
       'concurrent_stream',
       'twophase_snapshot',
-      'slot_creation_error',
       'skip_snapshot_restore',
       'invalidation_distribution',
       'parallel_session_origin',
-    ],
+    ] + test_decoding_extra_specs,
     'regress_args': [
       '--temp-config', files('logical.conf'),
     ],
diff --git a/src/test/modules/injection_points/Makefile b/src/test/modules/injection_points/Makefile
index 136f0f77951..87528590992 100644
--- a/src/test/modules/injection_points/Makefile
+++ b/src/test/modules/injection_points/Makefile
@@ -61,3 +61,10 @@ check:
 endif
 
 endif
+
+# wait_cleanup is unstable on Windows, where the FATAL message of a
+# terminated backend can get lost.  See
+# https://postgr.es/m/088AB35D-0860-494A-A9CF-8B301543AA51@yandex-team.ru
+ifeq ($(PORTNAME), win32)
+ISOLATION := $(filter-out wait_cleanup,$(ISOLATION))
+endif
diff --git a/src/test/modules/injection_points/meson.build b/src/test/modules/injection_points/meson.build
index db6b93a8115..dd4e64c2102 100644
--- a/src/test/modules/injection_points/meson.build
+++ b/src/test/modules/injection_points/meson.build
@@ -26,6 +26,14 @@ test_install_data += files(
   'injection_points--1.0.sql',
 )
 
+# wait_cleanup is unstable on Windows, where the FATAL message of a
+# terminated backend can get lost.  See
+# https://postgr.es/m/088AB35D-0860-494A-A9CF-8B301543AA51@yandex-team.ru
+injection_points_extra_specs = []
+if host_system != 'windows'
+  injection_points_extra_specs += 'wait_cleanup'
+endif
+
 tests += {
   'name': 'injection_points',
   'sd': meson.current_source_dir(),
@@ -56,10 +64,9 @@ tests += {
       'ri_fastpath_reindex',
       'ri_fastpath_snapshot',
       'syscache-update-pruned',
-      'wait_cleanup',
       'heap_lock_update',
       'on_conflict_probe_window',
-    ],
+    ] + injection_points_extra_specs,
     'runningcheck': false, # see syscache-update-pruned
     # Some tests wait for all snapshots, so avoid parallel execution
     'runningcheck-parallel': false,
-- 
2.43.0



^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-10-02 23:48  Michael Paquier <michael@paquier.xyz>
  parent: Zsolt Parragi <zsolt.parragi@percona.com>
  1 sibling, 0 replies; 21+ messages in thread

From: Michael Paquier @ 2026-10-02 23:48 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Andrey Borodin <x4mmm@yandex-team.ru>; Nikolay Samokhvalov <nik@postgres.ai>; pgsql-hackers@lists.postgresql.org

On Fri, Oct 02, 2026 at 10:44:20PM +0100, Zsolt Parragi wrote:
> I only disabled the two tests that already had failures on CI. After
> looking into it again I realized that temp-schema-cleanup could only
> fail with my hacked together linux repro, I can't reproduce it on
> windows. In wait_cleanup and slot_creation_error one session kills
> another session.

Thanks for the patch.  That should be able to do the job.  Good idea
to mention the thread, for future references.

> In temp-schema-cleanup s1 kills itself. When the
> fatal happens the isolation tester is reading s1's socket, and
> receives the message properly, I can't make that fail without
> modifying the testcase.

Okay, that's good to know.  That leads to an even simpler patch.
--
Michael

Attachments:

  [application/pgp-signature] signature.asc (832B, ../../asBCsCfFefC4lMQE@paquier.xyz/2-signature.asc)
  download

^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-10-05 02:14  Michael Paquier <michael@paquier.xyz>
  parent: Zsolt Parragi <zsolt.parragi@percona.com>
  1 sibling, 1 reply; 21+ messages in thread

From: Michael Paquier @ 2026-10-05 02:14 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Andrey Borodin <x4mmm@yandex-team.ru>; Nikolay Samokhvalov <nik@postgres.ai>; pgsql-hackers@lists.postgresql.org

On Fri, Oct 02, 2026 at 10:44:20PM +0100, Zsolt Parragi wrote:
> I only disabled the two tests that already had failures on CI. After
> looking into it again I realized that temp-schema-cleanup could only
> fail with my hacked together linux repro, I can't reproduce it on
> windows. In wait_cleanup and slot_creation_error one session kills
> another session. In temp-schema-cleanup s1 kills itself. When the
> fatal happens the isolation tester is reading s1's socket, and
> receives the message properly, I can't make that fail without
> modifying the testcase.

Okay, let's only disable these two as a first step, then.  For the
Makefiles, I have reused the same thing as pgcrypto, so as the
conditional test is stored in a variable before saving the ISOLATION
list.  For meson, the lists have been moved outside the "tests".  All
that felt slightly cleaner.

Now back to the main issue..
--
Michael

Attachments:

  [application/pgp-signature] signature.asc (832B, ../../asMIDM-g3ZMzuKEE@paquier.xyz/2-signature.asc)
  download

^ permalink  raw  reply  [nested|flat] 21+ messages in thread

* Re: injection_points: canceled or terminated waiters leak their wait slots
@ 2026-10-05 02:25  Michael Paquier <michael@paquier.xyz>
  parent: Michael Paquier <michael@paquier.xyz>
  0 siblings, 0 replies; 21+ messages in thread

From: Michael Paquier @ 2026-10-05 02:25 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Andrey Borodin <x4mmm@yandex-team.ru>; Nikolay Samokhvalov <nik@postgres.ai>; pgsql-hackers@lists.postgresql.org

On Mon, Oct 05, 2026 at 11:14:36AM +0900, Michael Paquier wrote:
> Now back to the main issue..

As the subject is much broader than what this thread is about, perhaps
the next steps had better be discussed on a new thread.  Could it be
possible to start a new thread with a proposal of patch there?
--
Michael

Attachments:

  [application/pgp-signature] signature.asc (832B, ../../asMKm68JpU5bPY5w@paquier.xyz/2-signature.asc)
  download

^ permalink  raw  reply  [nested|flat] 21+ messages in thread


end of thread, other threads:[~2026-10-05 02:25 UTC | newest]

Thread overview: 21+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2026-07-21 08:28 injection_points: canceled or terminated waiters leak their wait slots Zsolt Parragi <zsolt.parragi@percona.com>
2026-07-22 01:29 ` Michael Paquier <michael@paquier.xyz>
2026-07-22 06:07   ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-07-22 06:28     ` Michael Paquier <michael@paquier.xyz>
2026-07-23 05:40       ` Michael Paquier <michael@paquier.xyz>
2026-08-23 09:13         ` Andrey Borodin <x4mmm@yandex-team.ru>
2026-09-23 16:42           ` Nikolay Samokhvalov <nik@postgres.ai>
2026-09-28 22:12             ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-09-28 22:57               ` Michael Paquier <michael@paquier.xyz>
2026-09-29 00:38                 ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-09-29 11:24                   ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-09-29 23:46                     ` Michael Paquier <michael@paquier.xyz>
2026-09-30 12:31                       ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-09-30 16:14                         ` Kacper Kuras <kacperkuras@hotmail.com>
2026-09-30 22:55                         ` Michael Paquier <michael@paquier.xyz>
2026-09-30 18:00                       ` Andrey Borodin <x4mmm@yandex-team.ru>
2026-09-30 23:05                         ` Michael Paquier <michael@paquier.xyz>
2026-10-02 21:44                           ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-10-02 23:48                             ` Michael Paquier <michael@paquier.xyz>
2026-10-05 02:14                             ` Michael Paquier <michael@paquier.xyz>
2026-10-05 02:25                               ` Michael Paquier <michael@paquier.xyz>

This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox