pg.ddx.io  pgsql-bugs@postgresql.org mailing list archive  
help / color / mirror / Atom feed
autovacuum: automatically propagate updated parameters
43+ messages / 6 participants
[nested] [flat]

* autovacuum: automatically propagate updated parameters
@ 2026-07-24 08:33  Zsolt Parragi <zsolt.parragi@percona.com>
  0 siblings, 1 reply; 43+ messages in thread

From: Zsolt Parragi @ 2026-07-24 08:33 UTC (permalink / raw)
  To: pgsql-bugs@lists.postgresql.org

From the documentation:

Parallel workers launched for Parallel Vacuum are using the same cost
delay parameters as the leader worker. If any of these parameters are
changed in the leader worker, it will propagate the new parameter
values to all of its parallel workers.

But in practice, parallel_vacuum_propagate_shared_delay_params was
only called during config reload. Otherwise when the leader adjusted
the parameters, it didn't share them with the other workers.

See the attached patch which adds a test case about this and adds an
additional parallel_vacuum_propagate_shared_delay_params call to the
update logic.

Attachments:

  [application/octet-stream] 0001-Propagate-rebalanced-cost-limit-to-parallel-vacuum-w.patch (6.4K, ../../CAN4CZFOZtEPwGQ6oa9LvHvN522zEp8h_dW9hHExSh7pVXofoKQ@mail.gmail.com/2-0001-Propagate-rebalanced-cost-limit-to-parallel-vacuum-w.patch)
  download | inline diff:
From 3b2f1bcea74c8523756b66c49a0b92e387e3e02d Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 24 Jul 2026 08:18:54 +0000
Subject: [PATCH] Propagate rebalanced cost limit to parallel vacuum workers

AutoVacuumUpdateCostLimit() runs after each nap in vacuum_delay_point()
and follows av_nworkersForBalance, but the new limit never reached the
shared cost params in the vacuum DSM: propagation only happened on
config reload. Parallel workers computed their delays from the stale
limit, so a parallel autovacuum could run at up to twice the configured
budget (or half of it) until the next SIGHUP, contradicting the
propagation promise in maintenance.sgml.

Call parallel_vacuum_propagate_shared_delay_params() after rebalancing.
Gated on the leader: parallel workers take the same nap path and must
not overwrite the shared parameters.

Add a test: pause the leader at the existing injection point, start a
second autovacuum worker, check the first parameter load of the
parallel workers reports the balanced limit.
---
 src/backend/commands/vacuum.c                 |   3 +
 src/test/modules/test_autovacuum/meson.build  |   1 +
 .../test_autovacuum/t/002_cost_rebalance.pl   | 113 ++++++++++++++++++
 3 files changed, 117 insertions(+)
 create mode 100644 src/test/modules/test_autovacuum/t/002_cost_rebalance.pl

diff --git a/src/backend/commands/vacuum.c b/src/backend/commands/vacuum.c
index 38539a6fd3d..c17aecce7ad 100644
--- a/src/backend/commands/vacuum.c
+++ b/src/backend/commands/vacuum.c
@@ -2577,6 +2577,9 @@ vacuum_delay_point(bool is_analyze)
 		 */
 		AutoVacuumUpdateCostLimit();
 
+		if (AmAutoVacuumWorkerProcess())
+			parallel_vacuum_propagate_shared_delay_params();
+
 		/* Might have gotten an interrupt while sleeping */
 		CHECK_FOR_INTERRUPTS();
 	}
diff --git a/src/test/modules/test_autovacuum/meson.build b/src/test/modules/test_autovacuum/meson.build
index 86e392bc0de..7fbe50daa28 100644
--- a/src/test/modules/test_autovacuum/meson.build
+++ b/src/test/modules/test_autovacuum/meson.build
@@ -10,6 +10,7 @@ tests += {
     },
     'tests': [
       't/001_parallel_autovacuum.pl',
+      't/002_cost_rebalance.pl',
     ],
   },
 }
diff --git a/src/test/modules/test_autovacuum/t/002_cost_rebalance.pl b/src/test/modules/test_autovacuum/t/002_cost_rebalance.pl
new file mode 100644
index 00000000000..0089d9af780
--- /dev/null
+++ b/src/test/modules/test_autovacuum/t/002_cost_rebalance.pl
@@ -0,0 +1,113 @@
+
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+# Test that cost limit rebalancing reaches parallel autovacuum workers.
+#
+# Leader pauses at the injection point after the shared cost param snapshot
+# (balance = 1, limit 200). A second autovacuum worker starts (balance = 2,
+# limit 100). After resume, the parallel workers' first param load must show
+# the rebalanced 100, not the snapshotted 200.
+
+use strict;
+use warnings FATAL => 'all';
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+	plan skip_all => 'Injection points not supported by this build';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+
+$node->append_conf(
+	'postgresql.conf', qq{
+autovacuum_naptime = '1s'
+autovacuum_max_workers = 3
+autovacuum_worker_slots = 4
+autovacuum_max_parallel_workers = 2
+autovacuum_vacuum_cost_delay = '20ms'
+vacuum_cost_limit = 200
+max_worker_processes = 16
+max_parallel_workers = 8
+log_min_messages = debug2
+min_parallel_index_scan_size = 0
+log_autovacuum_min_duration = -1
+});
+$node->start;
+
+if (!$node->check_extension('injection_points'))
+{
+	plan skip_all => 'Extension injection_points not installed';
+}
+
+$node->safe_psql('postgres', 'CREATE EXTENSION injection_points');
+$node->safe_psql('postgres', 'CREATE DATABASE regress_db2');
+
+# leader's table
+$node->safe_psql(
+	'postgres', qq{
+	CREATE TABLE test_av (id serial primary key, col_1 int, col_2 int)
+	  WITH (autovacuum_parallel_workers = 2, autovacuum_enabled = false);
+	INSERT INTO test_av SELECT g, g + 1, g + 2 FROM generate_series(1, 50000) g;
+	CREATE INDEX test_av_col_1 ON test_av (col_1);
+	CREATE INDEX test_av_col_2 ON test_av (col_2);
+});
+
+# second worker's table: no indexes, no cost reloptions (participates in
+# balancing), sized to outlast the test
+$node->safe_psql(
+	'regress_db2', qq{
+	CREATE TABLE filler (id int, pad text) WITH (autovacuum_enabled = false);
+	INSERT INTO filler SELECT g, repeat('x', 100) FROM generate_series(1, 200000) g;
+});
+
+# quiesce catalogs so no extra worker skews the balance
+$node->safe_psql($_, 'VACUUM ANALYZE')
+  for ('postgres', 'regress_db2', 'template1');
+
+my $db2oid = $node->safe_psql('postgres',
+	"SELECT oid FROM pg_database WHERE datname = 'regress_db2'");
+my $filleroid = $node->safe_psql('regress_db2', "SELECT 'filler'::regclass::oid");
+
+$node->safe_psql('postgres', 'UPDATE test_av SET col_1 = col_1 + 1');
+$node->safe_psql('regress_db2', 'UPDATE filler SET id = id + 1');
+
+my $log_offset = -s $node->logfile;
+
+# pause leader after the shared cost param snapshot
+$node->safe_psql('postgres',
+	"SELECT injection_points_attach('autovacuum-start-parallel-vacuum', 'wait')");
+$node->safe_psql('postgres', 'ALTER TABLE test_av SET (autovacuum_enabled = true)');
+$node->wait_for_event('autovacuum worker', 'autovacuum-start-parallel-vacuum');
+
+# second worker -> balance = 2
+$node->safe_psql('regress_db2', 'ALTER TABLE filler SET (autovacuum_enabled = true)');
+$node->wait_for_log(
+	qr/VacuumUpdateCosts\(db=$db2oid, rel=$filleroid, dobalance=yes, cost_limit=100,/,
+	$log_offset);
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('autovacuum-start-parallel-vacuum')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('autovacuum-start-parallel-vacuum')");
+
+# first param load must show the rebalanced limit
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=\d+,/,
+	$log_offset);
+my $log = slurp_file($node->logfile, $log_offset);
+my @limits =
+  $log =~ /parallel autovacuum worker updated cost params: cost_limit=(\d+),/g;
+note("parallel worker cost_limit sequence: @limits");
+is($limits[0], '100',
+	'parallel workers see the rebalanced cost limit');
+
+my $filler_running = $node->safe_psql('regress_db2',
+	"SELECT count(*) FROM pg_stat_progress_vacuum WHERE relid = 'filler'::regclass");
+is($filler_running, '1', 'second autovacuum worker was still running');
+
+$node->stop;
+done_testing();
-- 
2.54.0



^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-08-24 12:31  Daniel Gustafsson <daniel@yesql.se>
  parent: Zsolt Parragi <zsolt.parragi@percona.com>
  0 siblings, 1 reply; 43+ messages in thread

From: Daniel Gustafsson @ 2026-08-24 12:31 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: pgsql-bugs@lists.postgresql.org

> On 24 Jul 2026, at 10:33, Zsolt Parragi <zsolt.parragi@percona.com> wrote:
> 
> From the documentation:
> 
> Parallel workers launched for Parallel Vacuum are using the same cost
> delay parameters as the leader worker. If any of these parameters are
> changed in the leader worker, it will propagate the new parameter
> values to all of its parallel workers.
> 
> But in practice, parallel_vacuum_propagate_shared_delay_params was
> only called during config reload. Otherwise when the leader adjusted
> the parameters, it didn't share them with the other workers.
> 
> See the attached patch which adds a test case about this and adds an
> additional parallel_vacuum_propagate_shared_delay_params call to the
> update logic.

I reviewed this today and I agree with the proposed fix.  The alternative would
be to update the documentation to match the reality of requiring a configuration
reload, but that brings on other baggage so I think fixing the code is the
better option here.

The part that worry me is the below testcode.  The relation 'filler' is sized
large enough to outlast the test:

	+# second worker's table: no indexes, no cost reloptions (participates in
	+# balancing), sized to outlast the test
	+$node->safe_psql(
	+	'regress_db2', qq{
	+	CREATE TABLE filler (id int, pad text) WITH (autovacuum_enabled = false);
	+	INSERT INTO filler SELECT g, repeat('x', 100) FROM generate_series(1, 200000) g;
	+});

This is then used for two test cases, the first one has legitimate value:

	+my $log = slurp_file($node->logfile, $log_offset);
	+my @limits =
	+  $log =~ /parallel autovacuum worker updated cost params: cost_limit=(\d+),/g;
	+note("parallel worker cost_limit sequence: @limits");
	+is($limits[0], '100',
	+	'parallel workers see the rebalanced cost limit');

The second seems less exciting.

	+my $filler_running = $node->safe_psql('regress_db2',
	+	"SELECT count(*) FROM pg_stat_progress_vacuum WHERE relid = 'filler'::regclass");
	+is($filler_running, '1', 'second autovacuum worker was still running');

My worry is that this seems very timing dependent and risk being flaky in the
buildfarm where every unsuspected timing window known to man tends to happen on
a regular basis.  Given the first test for the DEBUG2 log output, do we lose
all that much coverage if we cut the suite short after that, and reduce the
size of filler?  Reducing the resources needed to run the test and the
potential for red builds in the BF seems like a win.

--
Daniel Gustafsson






^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-08-24 19:42  Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
  parent: Daniel Gustafsson <daniel@yesql.se>
  0 siblings, 1 reply; 43+ messages in thread

From: Bharath Rupireddy @ 2026-08-24 19:42 UTC (permalink / raw)
  To: Daniel Gustafsson <daniel@yesql.se>; +Cc: Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

Hi,

On Mon, Aug 24, 2026 at 5:31 AM Daniel Gustafsson <daniel@yesql.se> wrote:
>
> > On 24 Jul 2026, at 10:33, Zsolt Parragi <zsolt.parragi@percona.com> wrote:
> >
> > Parallel workers launched for Parallel Vacuum are using the same cost
> > delay parameters as the leader worker. If any of these parameters are
> > changed in the leader worker, it will propagate the new parameter
> > values to all of its parallel workers.
> >
> > But in practice, parallel_vacuum_propagate_shared_delay_params was
> > only called during config reload. Otherwise when the leader adjusted
> > the parameters, it didn't share them with the other workers.
> >
> > See the attached patch which adds a test case about this and adds an
> > additional parallel_vacuum_propagate_shared_delay_params call to the
> > update logic.
>
> I reviewed this today and I agree with the proposed fix.  The alternative would
> be to update the documentation to match the reality of requiring a configuration
> reload, but that brings on other baggage so I think fixing the code is the
> better option here.

Nice catch! Yes, this needs to be fixed and backpatched to PG19
(1ff3180ca01). It misses propagating the changes made after the delay
update and requires one to reload the config. Mostly, the caller
doesn't know from outside whether the params have changed at all,
making the parallel workers for autovacuum not honor the cost params
at all.

How about propagating the params to parallel workers launched by
autovacuum right after updating cost limits in
AutoVacuumUpdateCostLimit()? With this change, the existing
propagate-upon-config-reload path could be removed. This looks
centralized and future-proof, unless I'm missing something. Thoughts?

> The part that worry me is the below testcode.  The relation 'filler' is sized
> large enough to outlast the test:

I understand that having a test for this in the first place could have
helped catch this issue. However, I don't see a strong point in having
one now. Can we test it manually and get the fix alone?

--
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com





^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-08-24 22:06  Zsolt Parragi <zsolt.parragi@percona.com>
  parent: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
  0 siblings, 1 reply; 43+ messages in thread

From: Zsolt Parragi @ 2026-08-24 22:06 UTC (permalink / raw)
  To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: pgsql-bugs@lists.postgresql.org, Daniel Gustafsson <daniel@yesql.se>

> I understand that having a test for this in the first place could have
> helped catch this issue. However, I don't see a strong point in having
> one now. Can we test it manually and get the fix alone?

I mainly added the testcase as a repro, I am not sure how useful it is
as an actual test case. But if we want to keep it, I can certainly
reduce it.

> How about propagating the params to parallel workers launched by
> autovacuum right after updating cost limits in
> AutoVacuumUpdateCostLimit()?

That could work too, I used the current location because most of the
parallel related things, including the other call to
parallel_vacuum_propagate_shared_delay_params is in vacuum.c, not in
autovacuum.c





^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-08-24 22:16  Daniel Gustafsson <daniel@yesql.se>
  parent: Zsolt Parragi <zsolt.parragi@percona.com>
  0 siblings, 1 reply; 43+ messages in thread

From: Daniel Gustafsson @ 2026-08-24 22:16 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; pgsql-bugs@lists.postgresql.org

> On 25 Aug 2026, at 00:06, Zsolt Parragi <zsolt.parragi@percona.com> wrote:
> 
>> I understand that having a test for this in the first place could have
>> helped catch this issue. However, I don't see a strong point in having
>> one now. Can we test it manually and get the fix alone?
> 
> I mainly added the testcase as a repro, I am not sure how useful it is
> as an actual test case. But if we want to keep it, I can certainly
> reduce it.

I think there is value in having a test, especially if we can roll it into 001
as quick step.

>> How about propagating the params to parallel workers launched by
>> autovacuum right after updating cost limits in
>> AutoVacuumUpdateCostLimit()?
> 
> That could work too, I used the current location because most of the
> parallel related things, including the other call to
> parallel_vacuum_propagate_shared_delay_params is in vacuum.c, not in
> autovacuum.c

There is also strong value in not changing things too drastic at this point in
the cycle.

--
Daniel Gustafsson






^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-08-24 22:26  Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
  parent: Daniel Gustafsson <daniel@yesql.se>
  0 siblings, 1 reply; 43+ messages in thread

From: Bharath Rupireddy @ 2026-08-24 22:26 UTC (permalink / raw)
  To: Daniel Gustafsson <daniel@yesql.se>; +Cc: Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

Hi,

On Mon, Aug 24, 2026 at 3:16 PM Daniel Gustafsson <daniel@yesql.se> wrote:
>
> > On 25 Aug 2026, at 00:06, Zsolt Parragi <zsolt.parragi@percona.com> wrote:
> >
> >> I understand that having a test for this in the first place could have
> >> helped catch this issue. However, I don't see a strong point in having
> >> one now. Can we test it manually and get the fix alone?
> >
> > I mainly added the testcase as a repro, I am not sure how useful it is
> > as an actual test case. But if we want to keep it, I can certainly
> > reduce it.
>
> I think there is value in having a test, especially if we can roll it into 001
> as quick step.

+1. How about adding the test in the existing
001_parallel_autovacuum.pl instead of a new TAP test file?

> >> How about propagating the params to parallel workers launched by
> >> autovacuum right after updating cost limits in
> >> AutoVacuumUpdateCostLimit()?
> >
> > That could work too, I used the current location because most of the
> > parallel related things, including the other call to
> > parallel_vacuum_propagate_shared_delay_params is in vacuum.c, not in
> > autovacuum.c
>
> There is also strong value in not changing things too drastic at this point in
> the cycle.

Agreed. I'm fine with the v1 approach and backpatching to PG19.

-- 
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com






^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-08-25 11:16  Zsolt Parragi <zsolt.parragi@percona.com>
  parent: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
  0 siblings, 1 reply; 43+ messages in thread

From: Zsolt Parragi @ 2026-08-25 11:16 UTC (permalink / raw)
  To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; pgsql-bugs@lists.postgresql.org

I attached v2 which moves the test case and removes the last part.

Attachments:

  [application/octet-stream] v2-0001-Propagate-rebalanced-cost-limit-to-parallel-vacuu.patch (4.9K, ../../CAN4CZFMETmka0iBJA_NjGk-+UP_t0ABwU62kVLzr0EYSzsGDBw@mail.gmail.com/2-v2-0001-Propagate-rebalanced-cost-limit-to-parallel-vacuu.patch)
  download | inline diff:
From 07f7cd5166fde9409ef7dffb788012a894b820f9 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 24 Jul 2026 08:18:54 +0000
Subject: [PATCH v2] Propagate rebalanced cost limit to parallel vacuum workers

AutoVacuumUpdateCostLimit() runs after each nap in vacuum_delay_point()
and follows av_nworkersForBalance, but the new limit never reached the
shared cost params in the vacuum DSM: propagation only happened on
config reload. Parallel workers computed their delays from the stale
limit, so a parallel autovacuum could run at up to twice the configured
budget (or half of it) until the next SIGHUP, contradicting the
propagation promise in maintenance.sgml.

Call parallel_vacuum_propagate_shared_delay_params() after rebalancing.
Gated on the leader: parallel workers take the same nap path and must
not overwrite the shared parameters.

Add a test: pause the leader at the existing injection point, start a
second autovacuum worker, check the first parameter load of the
parallel workers reports the balanced limit.
---
 src/backend/commands/vacuum.c                 |  3 +
 .../t/001_parallel_autovacuum.pl              | 70 ++++++++++++++++++-
 2 files changed, 72 insertions(+), 1 deletion(-)

diff --git a/src/backend/commands/vacuum.c b/src/backend/commands/vacuum.c
index 38539a6fd3d..c17aecce7ad 100644
--- a/src/backend/commands/vacuum.c
+++ b/src/backend/commands/vacuum.c
@@ -2577,6 +2577,9 @@ vacuum_delay_point(bool is_analyze)
 		 */
 		AutoVacuumUpdateCostLimit();
 
+		if (AmAutoVacuumWorkerProcess())
+			parallel_vacuum_propagate_shared_delay_params();
+
 		/* Might have gotten an interrupt while sleeping */
 		CHECK_FOR_INTERRUPTS();
 	}
diff --git a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
index 22f40cb1d50..c199337d449 100644
--- a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
+++ b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
@@ -36,7 +36,7 @@ $node->init;
 $node->append_conf(
 	'postgresql.conf', qq{
 autovacuum_max_workers = 1
-autovacuum_worker_slots = 1
+autovacuum_worker_slots = 2
 autovacuum_max_parallel_workers = 2
 max_worker_processes = 10
 max_parallel_workers = 10
@@ -167,5 +167,73 @@ ok(1,
 	"vacuum delay parameter changes are propagated to parallel vacuum workers"
 );
 
+# Test 3:
+# Check whether a cost limit rebalance reaches the parallel workers. The
+# leader pauses right after taking the shared cost param snapshot
+# (balance = 1, limit 500), then a second autovacuum worker starts
+# (balance = 2, limit 250). After resume, the parallel workers' first
+# parameter load must show the rebalanced 250, not the snapshotted 500.
+
+# Second worker's table lives in another database: no indexes, no cost
+# reloptions (participates in balancing)
+$node->safe_psql('postgres', 'CREATE DATABASE regress_db2');
+$node->safe_psql(
+	'regress_db2', qq{
+	CREATE TABLE filler (id int, pad text) WITH (autovacuum_enabled = false);
+	INSERT INTO filler SELECT g, repeat('x', 100) FROM generate_series(1, 200000) g;
+});
+
+# Allow a second autovacuum worker.
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_max_workers = 2;
+	SELECT pg_reload_conf();
+});
+
+# Quiesce catalogs so no extra worker skews the balance.
+$node->safe_psql($_, 'VACUUM ANALYZE')
+  for ('postgres', 'regress_db2', 'template1');
+
+prepare_for_next_test($node, 3);
+$node->safe_psql('regress_db2', 'UPDATE filler SET id = id + 1');
+
+my $db2oid = $node->safe_psql('postgres',
+	"SELECT oid FROM pg_database WHERE datname = 'regress_db2'");
+my $filleroid =
+  $node->safe_psql('regress_db2', "SELECT 'filler'::regclass::oid");
+
+$log_offset = -s $node->logfile;
+
+# Pause the leader after the shared cost param snapshot.
+$node->safe_psql('postgres',
+	"SELECT injection_points_attach('autovacuum-start-parallel-vacuum', 'wait')"
+);
+$node->safe_psql('postgres',
+	'ALTER TABLE test_autovac SET (autovacuum_enabled = true)');
+$node->wait_for_event('autovacuum worker',
+	'autovacuum-start-parallel-vacuum');
+
+# Second worker -> balance = 2
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = true)');
+$node->wait_for_log(
+	qr/VacuumUpdateCosts\(db=$db2oid, rel=$filleroid, dobalance=yes, cost_limit=250,/,
+	$log_offset);
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('autovacuum-start-parallel-vacuum')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('autovacuum-start-parallel-vacuum')");
+
+# First param load must show the rebalanced limit.
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=\d+,/,
+	$log_offset);
+my $log = slurp_file($node->logfile, $log_offset);
+my @limits =
+  $log =~ /parallel autovacuum worker updated cost params: cost_limit=(\d+),/g;
+note("parallel worker cost_limit sequence: @limits");
+is($limits[0], '250', 'parallel workers see the rebalanced cost limit');
+
 $node->stop;
 done_testing();
-- 
2.55.0



^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-08-27 00:07  Masahiko Sawada <sawada.mshk@gmail.com>
  parent: Zsolt Parragi <zsolt.parragi@percona.com>
  0 siblings, 1 reply; 43+ messages in thread

From: Masahiko Sawada @ 2026-08-27 00:07 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Daniel Gustafsson <daniel@yesql.se>; pgsql-bugs@lists.postgresql.org

On Tue, Aug 25, 2026 at 4:17 AM Zsolt Parragi <zsolt.parragi@percona.com> wrote:
>
> I attached v2 which moves the test case and removes the last part.

Thank you for the report and making the patch.

The fix looks good to me. As for the regression tests, since we don't
stop the second av worker, the first av worker needs to resume and
update its cost limit before the second worker finishes:

+# Second worker -> balance = 2
+$node->safe_psql('regress_db2',
+   'ALTER TABLE filler SET (autovacuum_enabled = true)');
+$node->wait_for_log(
+   qr/VacuumUpdateCosts\(db=$db2oid, rel=$filleroid, dobalance=yes,
cost_limit=250,/,
+   $log_offset);
+
+$node->safe_psql('postgres',
+   "SELECT injection_points_wakeup('autovacuum-start-parallel-vacuum')");
+$node->safe_psql('postgres',
+   "SELECT injection_points_detach('autovacuum-start-parallel-vacuum')");
+
+# First param load must show the rebalanced limit.
+$node->wait_for_log(
+   qr/parallel autovacuum worker updated cost params: cost_limit=\d+,/,
+   $log_offset);
+my $log = slurp_file($node->logfile, $log_offset);
+my @limits =
+  $log =~ /parallel autovacuum worker updated cost params: cost_limit=(\d+),/g;
+note("parallel worker cost_limit sequence: @limits");
+is($limits[0], '250', 'parallel workers see the rebalanced cost limit');

Which seems to be unstable to me. In order to ensure that the second
worker lives when the first worker resumes its job, we can stop the
second worker at the injection point "vacuum-truncate-enabled" for
example. While it works and we can reduce the filler table size, it
would add an unclear dependency as vacuum-truncate-enabled is not
related to parallel autovacuum.

Regards,

-- 
Masahiko Sawada
Amazon Web Services: https://aws.amazon.com






^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-08-27 08:37  Zsolt Parragi <zsolt.parragi@percona.com>
  parent: Masahiko Sawada <sawada.mshk@gmail.com>
  0 siblings, 1 reply; 43+ messages in thread

From: Zsolt Parragi @ 2026-08-27 08:37 UTC (permalink / raw)
  To: Masahiko Sawada <sawada.mshk@gmail.com>; +Cc: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Daniel Gustafsson <daniel@yesql.se>; pgsql-bugs@lists.postgresql.org

> we can stop the
> second worker at the injection point "vacuum-truncate-enabled" for
> example. While it works and we can reduce the filler table size, it
> would add an unclear dependency as vacuum-truncate-enabled is not
> related to parallel autovacuum.

I added a new injection point and stopped it there in v3. If you think
that's unnecessary, we can replace it to vacuum-truncate-enabled in
the test and remove the injection point, and it works the same way.
This just seemed cleaner to me.

Attachments:

  [application/octet-stream] v3-0001-Propagate-rebalanced-cost-limit-to-parallel-vacuu.patch (6.2K, ../../CAN4CZFPoQrT2ddLUiJXmJ5o80YF9LN_3m9PL4QCX_bddj44uGA@mail.gmail.com/2-v3-0001-Propagate-rebalanced-cost-limit-to-parallel-vacuu.patch)
  download | inline diff:
From 6f7386b5a86cd2ea12ac1e8b7e8fb3052c4fcee4 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 24 Jul 2026 08:18:54 +0000
Subject: [PATCH v3] Propagate rebalanced cost limit to parallel vacuum workers

AutoVacuumUpdateCostLimit() runs after each nap in vacuum_delay_point()
and follows av_nworkersForBalance, but the new limit never reached the
shared cost params in the vacuum DSM: propagation only happened on
config reload. Parallel workers computed their delays from the stale
limit, so a parallel autovacuum could run at up to twice the configured
budget (or half of it) until the next SIGHUP, contradicting the
propagation promise in maintenance.sgml.

Call parallel_vacuum_propagate_shared_delay_params() after rebalancing.
Gated on the leader: parallel workers take the same nap path and must
not overwrite the shared parameters.

Add a test: pause the leader at the existing injection point, start a
second autovacuum worker and hold it at a new injection point placed
after it joined the balance, then check the first parameter load of the
parallel workers reports the balanced limit. The hold is needed because
a second worker left running can finish its own vacuum before the leader
resumes, which puts the balance back where it started.
---
 src/backend/commands/vacuum.c                 |  3 +
 src/backend/postmaster/autovacuum.c           |  1 +
 .../t/001_parallel_autovacuum.pl              | 82 ++++++++++++++++++-
 3 files changed, 85 insertions(+), 1 deletion(-)

diff --git a/src/backend/commands/vacuum.c b/src/backend/commands/vacuum.c
index 38539a6fd3d..c17aecce7ad 100644
--- a/src/backend/commands/vacuum.c
+++ b/src/backend/commands/vacuum.c
@@ -2577,6 +2577,9 @@ vacuum_delay_point(bool is_analyze)
 		 */
 		AutoVacuumUpdateCostLimit();
 
+		if (AmAutoVacuumWorkerProcess())
+			parallel_vacuum_propagate_shared_delay_params();
+
 		/* Might have gotten an interrupt while sleeping */
 		CHECK_FOR_INTERRUPTS();
 	}
diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c
index 45abf48768a..b06da6bce1d 100644
--- a/src/backend/postmaster/autovacuum.c
+++ b/src/backend/postmaster/autovacuum.c
@@ -2487,6 +2487,7 @@ do_autovacuum(void)
 		 */
 		VacuumUpdateCosts();
 
+		INJECTION_POINT("autovacuum-worker-cost-balanced", NULL);
 
 		/* clean up memory before each iteration */
 		MemoryContextReset(PortalContext);
diff --git a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
index 22f40cb1d50..ca7ff727b5e 100644
--- a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
+++ b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
@@ -36,7 +36,7 @@ $node->init;
 $node->append_conf(
 	'postgresql.conf', qq{
 autovacuum_max_workers = 1
-autovacuum_worker_slots = 1
+autovacuum_worker_slots = 2
 autovacuum_max_parallel_workers = 2
 max_worker_processes = 10
 max_parallel_workers = 10
@@ -167,5 +167,85 @@ ok(1,
 	"vacuum delay parameter changes are propagated to parallel vacuum workers"
 );
 
+# Test 3:
+# Check whether a cost limit rebalance reaches the parallel workers. The
+# leader pauses right after taking the shared cost param snapshot
+# (balance = 1, limit 500), then a second autovacuum worker joins the
+# balance (balance = 2, limit 250) and is held there for the rest of the
+# test. After resume, the parallel workers' first parameter load must show
+# the rebalanced 250, not the snapshotted 500.
+
+# Second worker's table lives in another database: no cost reloptions, so it
+# participates in balancing.
+$node->safe_psql('postgres', 'CREATE DATABASE regress_db2');
+$node->safe_psql(
+	'regress_db2', qq{
+	CREATE TABLE filler (id int) WITH (autovacuum_enabled = false);
+	INSERT INTO filler SELECT g FROM generate_series(1, 1000) g;
+});
+
+# Allow a second autovacuum worker.
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_max_workers = 2;
+	SELECT pg_reload_conf();
+});
+
+# Quiesce catalogs so no extra worker skews the balance.
+$node->safe_psql($_, 'VACUUM ANALYZE')
+  for ('postgres', 'regress_db2', 'template1');
+
+prepare_for_next_test($node, 3);
+$node->safe_psql('regress_db2', 'UPDATE filler SET id = id + 1');
+
+my $db2oid = $node->safe_psql('postgres',
+	"SELECT oid FROM pg_database WHERE datname = 'regress_db2'");
+my $filleroid =
+  $node->safe_psql('regress_db2', "SELECT 'filler'::regclass::oid");
+
+$log_offset = -s $node->logfile;
+
+# Pause the leader after the shared cost param snapshot. The leader is past
+# the hold point below by then, so that one only catches the second worker.
+$node->safe_psql('postgres',
+	"SELECT injection_points_attach('autovacuum-start-parallel-vacuum', 'wait')"
+);
+$node->safe_psql('postgres',
+	'ALTER TABLE test_autovac SET (autovacuum_enabled = true)');
+$node->wait_for_event('autovacuum worker',
+	'autovacuum-start-parallel-vacuum');
+
+# Second worker -> balance = 2. Hold it there: if it were allowed to finish,
+# the balance would drop back to 1 before the leader resumes.
+$node->safe_psql('postgres',
+	"SELECT injection_points_attach('autovacuum-worker-cost-balanced', 'wait')"
+);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = true)');
+$node->wait_for_log(
+	qr/VacuumUpdateCosts\(db=$db2oid, rel=$filleroid, dobalance=yes, cost_limit=250,/,
+	$log_offset);
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('autovacuum-start-parallel-vacuum')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('autovacuum-start-parallel-vacuum')");
+
+# First param load must show the rebalanced limit.
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=\d+,/,
+	$log_offset);
+my $log = slurp_file($node->logfile, $log_offset);
+my @limits =
+  $log =~ /parallel autovacuum worker updated cost params: cost_limit=(\d+),/g;
+note("parallel worker cost_limit sequence: @limits");
+is($limits[0], '250', 'parallel workers see the rebalanced cost limit');
+
+# Release the second worker.
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('autovacuum-worker-cost-balanced')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('autovacuum-worker-cost-balanced')");
+
 $node->stop;
 done_testing();
-- 
2.55.0



^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-08-27 11:57  Daniel Gustafsson <daniel@yesql.se>
  parent: Zsolt Parragi <zsolt.parragi@percona.com>
  0 siblings, 1 reply; 43+ messages in thread

From: Daniel Gustafsson @ 2026-08-27 11:57 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Masahiko Sawada <sawada.mshk@gmail.com>; Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; pgsql-bugs@lists.postgresql.org

> On 27 Aug 2026, at 10:37, Zsolt Parragi <zsolt.parragi@percona.com> wrote:
> 
>> we can stop the
>> second worker at the injection point "vacuum-truncate-enabled" for
>> example. While it works and we can reduce the filler table size, it
>> would add an unclear dependency as vacuum-truncate-enabled is not
>> related to parallel autovacuum.
> 
> I added a new injection point and stopped it there in v3. If you think
> that's unnecessary, we can replace it to vacuum-truncate-enabled in
> the test and remove the injection point, and it works the same way.
> This just seemed cleaner to me.

I prefer this approach, injection points are cheap enough that we don't need to
reuse and cause undefined dependencies. I'll try to get this applied later today.

--
Daniel Gustafsson







^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-08-27 21:15  Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
  parent: Daniel Gustafsson <daniel@yesql.se>
  0 siblings, 1 reply; 43+ messages in thread

From: Bharath Rupireddy @ 2026-08-27 21:15 UTC (permalink / raw)
  To: Daniel Gustafsson <daniel@yesql.se>; +Cc: Zsolt Parragi <zsolt.parragi@percona.com>; Masahiko Sawada <sawada.mshk@gmail.com>; pgsql-bugs@lists.postgresql.org

Hi,

On Thu, Aug 27, 2026 at 4:57 AM Daniel Gustafsson <daniel@yesql.se> wrote:
>
> >> we can stop the
> >> second worker at the injection point "vacuum-truncate-enabled" for
> >> example. While it works and we can reduce the filler table size, it
> >> would add an unclear dependency as vacuum-truncate-enabled is not
> >> related to parallel autovacuum.
> >
> > I added a new injection point and stopped it there in v3. If you think
> > that's unnecessary, we can replace it to vacuum-truncate-enabled in
> > the test and remove the injection point, and it works the same way.
> > This just seemed cleaner to me.
>
> I prefer this approach, injection points are cheap enough that we don't need to
> reuse and cause undefined dependencies. I'll try to get this applied later today.

+1, a separate injection point makes sense here. The attached v3 patch
looks good to me. pgindent and tests are happy. This needs to be
backpatched to PG19.

--
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com






^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-08-27 22:02  Daniel Gustafsson <daniel@yesql.se>
  parent: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
  0 siblings, 1 reply; 43+ messages in thread

From: Daniel Gustafsson @ 2026-08-27 22:02 UTC (permalink / raw)
  To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: Zsolt Parragi <zsolt.parragi@percona.com>; Masahiko Sawada <sawada.mshk@gmail.com>; pgsql-bugs@lists.postgresql.org

> On 27 Aug 2026, at 23:15, Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com> wrote:
> 
> Hi,
> 
> On Thu, Aug 27, 2026 at 4:57 AM Daniel Gustafsson <daniel@yesql.se> wrote:
>> 
>>>> we can stop the
>>>> second worker at the injection point "vacuum-truncate-enabled" for
>>>> example. While it works and we can reduce the filler table size, it
>>>> would add an unclear dependency as vacuum-truncate-enabled is not
>>>> related to parallel autovacuum.
>>> 
>>> I added a new injection point and stopped it there in v3. If you think
>>> that's unnecessary, we can replace it to vacuum-truncate-enabled in
>>> the test and remove the injection point, and it works the same way.
>>> This just seemed cleaner to me.
>> 
>> I prefer this approach, injection points are cheap enough that we don't need to
>> reuse and cause undefined dependencies. I'll try to get this applied later today.
> 
> +1, a separate injection point makes sense here. The attached v3 patch
> looks good to me. pgindent and tests are happy. This needs to be
> backpatched to PG19.

Thanks for review, I have it scheduled for commit and backpatch tomorrow morning.

--
Daniel Gustafsson






^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-08-27 22:43  Masahiko Sawada <sawada.mshk@gmail.com>
  parent: Daniel Gustafsson <daniel@yesql.se>
  0 siblings, 2 replies; 43+ messages in thread

From: Masahiko Sawada @ 2026-08-27 22:43 UTC (permalink / raw)
  To: Daniel Gustafsson <daniel@yesql.se>; +Cc: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

On Thu, Aug 27, 2026 at 3:02 PM Daniel Gustafsson <daniel@yesql.se> wrote:
>
> > On 27 Aug 2026, at 23:15, Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com> wrote:
> >
> > Hi,
> >
> > On Thu, Aug 27, 2026 at 4:57 AM Daniel Gustafsson <daniel@yesql.se> wrote:
> >>
> >>>> we can stop the
> >>>> second worker at the injection point "vacuum-truncate-enabled" for
> >>>> example. While it works and we can reduce the filler table size, it
> >>>> would add an unclear dependency as vacuum-truncate-enabled is not
> >>>> related to parallel autovacuum.
> >>>
> >>> I added a new injection point and stopped it there in v3. If you think
> >>> that's unnecessary, we can replace it to vacuum-truncate-enabled in
> >>> the test and remove the injection point, and it works the same way.
> >>> This just seemed cleaner to me.
> >>
> >> I prefer this approach, injection points are cheap enough that we don't need to
> >> reuse and cause undefined dependencies. I'll try to get this applied later today.
> >
> > +1, a separate injection point makes sense here. The attached v3 patch
> > looks good to me. pgindent and tests are happy. This needs to be
> > backpatched to PG19.
>
> Thanks for review, I have it scheduled for commit and backpatch tomorrow morning.

Thank you for taking care of it. The v3 patch looks good to me.

Regards,

-- 
Masahiko Sawada
Amazon Web Services: https://aws.amazon.com





^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-08-27 22:52  Zsolt Parragi <zsolt.parragi@percona.com>
  parent: Masahiko Sawada <sawada.mshk@gmail.com>
  1 sibling, 1 reply; 43+ messages in thread

From: Zsolt Parragi @ 2026-08-27 22:52 UTC (permalink / raw)
  To: Masahiko Sawada <sawada.mshk@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; pgsql-bugs@lists.postgresql.org

Unfortunately v3 failed on CI, and after some more local testing it
was slightly unstable even on my machine.

v4 aims to fix that by changing autovacuum thresholds so that it
effectively only runs for the test tables

Attachments:

  [application/octet-stream] v4-0001-Propagate-rebalanced-cost-limit-to-parallel-vacuu.patch (7.3K, ../../CAN4CZFP-A=teXQPcq5+NXgDYYtL8D1B172kafB67WuhmOXqy+Q@mail.gmail.com/2-v4-0001-Propagate-rebalanced-cost-limit-to-parallel-vacuu.patch)
  download | inline diff:
From cb70ca048b95a0ca41f07d6ff2b220002c2ed826 Mon Sep 17 00:00:00 2001
From: Zsolt Parragi <zsolt.parragi@percona.com>
Date: Fri, 24 Jul 2026 08:18:54 +0000
Subject: [PATCH v4] Propagate rebalanced cost limit to parallel vacuum workers

AutoVacuumUpdateCostLimit() runs after each nap in vacuum_delay_point()
and follows av_nworkersForBalance, but the new limit never reached the
shared cost params in the vacuum DSM: propagation only happened on
config reload. Parallel workers computed their delays from the stale
limit, so a parallel autovacuum could run at up to twice the configured
budget (or half of it) until the next SIGHUP, contradicting the
propagation promise in maintenance.sgml.

Call parallel_vacuum_propagate_shared_delay_params() after rebalancing.
Gated on the leader: parallel workers take the same nap path and must
not overwrite the shared parameters.

Add a test: pause the leader at the existing injection point, start a
second autovacuum worker and hold it at a new injection point placed
after it joined the balance, then check the first parameter load of the
parallel workers reports the balanced limit. The hold is needed because
a second worker left running can finish its own vacuum before the leader
resumes, which puts the balance back where it started. Autovacuum is
disabled for everything but the two test tables via thresholds, as a
worker spawned by catalog churn would get trapped at the hold point and
starve the test of its second worker slot.
---
 src/backend/commands/vacuum.c                 |  3 +
 src/backend/postmaster/autovacuum.c           |  1 +
 .../t/001_parallel_autovacuum.pl              | 88 ++++++++++++++++++-
 3 files changed, 91 insertions(+), 1 deletion(-)

diff --git a/src/backend/commands/vacuum.c b/src/backend/commands/vacuum.c
index 38539a6fd3d..c17aecce7ad 100644
--- a/src/backend/commands/vacuum.c
+++ b/src/backend/commands/vacuum.c
@@ -2577,6 +2577,9 @@ vacuum_delay_point(bool is_analyze)
 		 */
 		AutoVacuumUpdateCostLimit();
 
+		if (AmAutoVacuumWorkerProcess())
+			parallel_vacuum_propagate_shared_delay_params();
+
 		/* Might have gotten an interrupt while sleeping */
 		CHECK_FOR_INTERRUPTS();
 	}
diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c
index 45abf48768a..b06da6bce1d 100644
--- a/src/backend/postmaster/autovacuum.c
+++ b/src/backend/postmaster/autovacuum.c
@@ -2487,6 +2487,7 @@ do_autovacuum(void)
 		 */
 		VacuumUpdateCosts();
 
+		INJECTION_POINT("autovacuum-worker-cost-balanced", NULL);
 
 		/* clean up memory before each iteration */
 		MemoryContextReset(PortalContext);
diff --git a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
index 22f40cb1d50..33c86bbdc94 100644
--- a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
+++ b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
@@ -33,10 +33,15 @@ $node->init;
 # Limit to one autovacuum worker and disable autovacuum logging globally
 # (enabled only on the test table) so that log checks below match only
 # activity on the expected table.
+#
+# Effectively disable autovacuum for all tables except the ones the test
+# re-enables via reloptions.  A worker spawned by catalog churn would skew
+# the cost balance, and an injection point attached below would trap it,
+# eating the only free worker slot.
 $node->append_conf(
 	'postgresql.conf', qq{
 autovacuum_max_workers = 1
-autovacuum_worker_slots = 1
+autovacuum_worker_slots = 2
 autovacuum_max_parallel_workers = 2
 max_worker_processes = 10
 max_parallel_workers = 10
@@ -44,6 +49,9 @@ log_min_messages = debug2
 autovacuum_naptime = '1s'
 min_parallel_index_scan_size = 0
 log_autovacuum_min_duration = -1
+autovacuum_vacuum_threshold = 100000
+autovacuum_analyze_threshold = 100000
+autovacuum_vacuum_insert_threshold = -1
 });
 $node->start;
 
@@ -72,6 +80,7 @@ $node->safe_psql(
 		id SERIAL PRIMARY KEY,
 		col_1 INTEGER,  col_2 INTEGER,  col_3 INTEGER,  col_4 INTEGER
 	) WITH (autovacuum_parallel_workers = $autovacuum_parallel_workers,
+			autovacuum_vacuum_threshold = 50,
 			log_autovacuum_min_duration = 0);
 
 	INSERT INTO test_autovac
@@ -167,5 +176,82 @@ ok(1,
 	"vacuum delay parameter changes are propagated to parallel vacuum workers"
 );
 
+# Test 3:
+# Check whether a cost limit rebalance reaches the parallel workers. The
+# leader pauses right after taking the shared cost param snapshot
+# (balance = 1, limit 500), then a second autovacuum worker joins the
+# balance (balance = 2, limit 250) and is held there for the rest of the
+# test. After resume, the parallel workers' first parameter load must show
+# the rebalanced 250, not the snapshotted 500.
+
+# Second worker's table lives in another database: no cost reloptions, so it
+# participates in balancing.
+$node->safe_psql('postgres', 'CREATE DATABASE regress_db2');
+$node->safe_psql(
+	'regress_db2', qq{
+	CREATE TABLE filler (id int)
+		WITH (autovacuum_enabled = false, autovacuum_vacuum_threshold = 50);
+	INSERT INTO filler SELECT g FROM generate_series(1, 1000) g;
+});
+
+# Allow a second autovacuum worker.
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_max_workers = 2;
+	SELECT pg_reload_conf();
+});
+
+prepare_for_next_test($node, 3);
+$node->safe_psql('regress_db2', 'UPDATE filler SET id = id + 1');
+
+my $db2oid = $node->safe_psql('postgres',
+	"SELECT oid FROM pg_database WHERE datname = 'regress_db2'");
+my $filleroid =
+  $node->safe_psql('regress_db2', "SELECT 'filler'::regclass::oid");
+
+$log_offset = -s $node->logfile;
+
+# Pause the leader after the shared cost param snapshot. The leader is past
+# the hold point below by then, so that one only catches the second worker.
+$node->safe_psql('postgres',
+	"SELECT injection_points_attach('autovacuum-start-parallel-vacuum', 'wait')"
+);
+$node->safe_psql('postgres',
+	'ALTER TABLE test_autovac SET (autovacuum_enabled = true)');
+$node->wait_for_event('autovacuum worker',
+	'autovacuum-start-parallel-vacuum');
+
+# Second worker -> balance = 2. Hold it there: if it were allowed to finish,
+# the balance would drop back to 1 before the leader resumes.
+$node->safe_psql('postgres',
+	"SELECT injection_points_attach('autovacuum-worker-cost-balanced', 'wait')"
+);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = true)');
+$node->wait_for_log(
+	qr/VacuumUpdateCosts\(db=$db2oid, rel=$filleroid, dobalance=yes, cost_limit=250,/,
+	$log_offset);
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('autovacuum-start-parallel-vacuum')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('autovacuum-start-parallel-vacuum')");
+
+# First param load must show the rebalanced limit.
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=\d+,/,
+	$log_offset);
+my $log = slurp_file($node->logfile, $log_offset);
+my @limits =
+  $log =~ /parallel autovacuum worker updated cost params: cost_limit=(\d+),/g;
+note("parallel worker cost_limit sequence: @limits");
+is($limits[0], '250', 'parallel workers see the rebalanced cost limit');
+
+# Release the second worker.
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('autovacuum-worker-cost-balanced')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('autovacuum-worker-cost-balanced')");
+
 $node->stop;
 done_testing();
-- 
2.55.0



^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-08-28 07:13  Daniel Gustafsson <daniel@yesql.se>
  parent: Zsolt Parragi <zsolt.parragi@percona.com>
  0 siblings, 1 reply; 43+ messages in thread

From: Daniel Gustafsson @ 2026-08-28 07:13 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Masahiko Sawada <sawada.mshk@gmail.com>; Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; pgsql-bugs@lists.postgresql.org

> On 28 Aug 2026, at 00:52, Zsolt Parragi <zsolt.parragi@percona.com> wrote:
> 
> Unfortunately v3 failed on CI, and after some more local testing it
> was slightly unstable even on my machine.
> 
> v4 aims to fix that by changing autovacuum thresholds so that it
> effectively only runs for the test tables

Thanks for the additional testing and update, I hadn't seen it fail in CI (or
locally) so it was good that you caught it.  Will run v4 in CI and locally
today before going ahead.

--
Daniel Gustafsson







^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-08-28 22:31  Daniel Gustafsson <daniel@yesql.se>
  parent: Daniel Gustafsson <daniel@yesql.se>
  0 siblings, 0 replies; 43+ messages in thread

From: Daniel Gustafsson @ 2026-08-28 22:31 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Masahiko Sawada <sawada.mshk@gmail.com>; Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; pgsql-bugs@lists.postgresql.org

> On 28 Aug 2026, at 09:13, Daniel Gustafsson <daniel@yesql.se> wrote:
> 
>> On 28 Aug 2026, at 00:52, Zsolt Parragi <zsolt.parragi@percona.com> wrote:
>> 
>> Unfortunately v3 failed on CI, and after some more local testing it
>> was slightly unstable even on my machine.
>> 
>> v4 aims to fix that by changing autovacuum thresholds so that it
>> effectively only runs for the test tables
> 
> Thanks for the additional testing and update, I hadn't seen it fail in CI (or
> locally) so it was good that you caught it.  Will run v4 in CI and locally
> today before going ahead.

Done.

--
Daniel Gustafsson







^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-21 18:28  Nikolay Samokhvalov <nik@postgres.ai>
  parent: Masahiko Sawada <sawada.mshk@gmail.com>
  1 sibling, 1 reply; 43+ messages in thread

From: Nikolay Samokhvalov @ 2026-09-21 18:28 UTC (permalink / raw)
  To: Masahiko Sawada <sawada.mshk@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

On Thu, Aug 27, 2026 at 3:43 PM Masahiko Sawada <sawada.mshk@gmail.com> wrote:
> Thank you for taking care of it. The v3 patch looks good to me.

AI found one more gap while testing the committed fix on REL_19_STABLE at
b368bdd2. I haven't manually reviewed the code yet.

Once an autovacuum leader enters WaitForParallelWorkersToFinish(), SIGHUP
wakes its latch, but the loop only runs CHECK_FOR_INTERRUPTS(), leaving
ConfigReloadPending set. A cost-limit rebalance is not signalled at all. In
both cases, the leader does not publish changed cost parameters until the
parallel worker finishes.

The standalone reproducer uses 100 rows and two indexes, with no fixed
sleeps. Apply it to b368bdd2 and run:

  make -C src/test/modules/test_autovacuum check \
    PROVE_TESTS=t/002_cost_reload_while_waiting.pl

It fails with:

  got: 'pending'
  expected: 'processed'

The other attached patch adds an optional callback to the worker-finish
wait. Parallel autovacuum uses it to handle reloads and poll cost-limit
rebalancing every 100 ms; other callers keep the existing behavior. It also
adds tests for both cases to 001_parallel_autovacuum.pl.

With the fix, test_autovacuum passes, as do the core regression (239 tests)
and isolation (133 tests) suites.

The reproducer and fix patches are separate alternatives against b368bdd2,
not a series.

Nik

Attachments:

  [application/x-patch] 0001-Add-reproducer-for-parallel-autovacuum-reload-wait.patch (5.9K, ../../CAM527d-GL=Jp2EJXBSnVBGPK-4XEZwWof5Cv8P0hghS_og6oAg@mail.gmail.com/2-0001-Add-reproducer-for-parallel-autovacuum-reload-wait.patch)
  download | inline diff:
From a6387345d0e71056442cc08b687f3c5965b8632f Mon Sep 17 00:00:00 2001
From: Nik Samokhvalov <nik@postgres.ai>
Date: Mon, 21 Sep 2026 09:01:50 -0700
Subject: [REPRODUCER] Add reproducer for parallel autovacuum reload wait

---
 src/backend/commands/vacuumparallel.c         |  14 +++
 .../t/002_cost_reload_while_waiting.pl        | 109 ++++++++++++++++++
 2 files changed, 123 insertions(+)
 create mode 100644 src/test/modules/test_autovacuum/t/002_cost_reload_while_waiting.pl

diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 767d162e57..918508518c 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -41,11 +41,14 @@
 #include "commands/progress.h"
 #include "commands/vacuum.h"
 #include "executor/instrument.h"
+#include "miscadmin.h"
 #include "optimizer/paths.h"
 #include "pgstat.h"
+#include "postmaster/interrupt.h"
 #include "storage/bufmgr.h"
 #include "storage/proc.h"
 #include "tcop/tcopprot.h"
+#include "utils/injection_point.h"
 #include "utils/lsyscache.h"
 #include "utils/rel.h"
 
@@ -929,6 +932,9 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 	/* Vacuum the indexes that can be processed by only leader process */
 	parallel_vacuum_process_unsafe_indexes(pvs);
 
+	if (pvs->shared->is_autovacuum)
+		INJECTION_POINT("parallel-autovacuum-leader-before-index", NULL);
+
 	/*
 	 * Join as a parallel worker.  The leader vacuums alone processes all
 	 * parallel-safe indexes in the case where no workers are launched.
@@ -944,6 +950,11 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 		/* Wait for all vacuum workers to finish */
 		WaitForParallelWorkersToFinish(pvs->pcxt);
 
+		if (pvs->shared->is_autovacuum)
+			INJECTION_POINT("parallel-autovacuum-leader-after-worker-wait",
+							ConfigReloadPending ? "reload pending" :
+							"reload processed");
+
 		for (int i = 0; i < pvs->pcxt->nworkers_launched; i++)
 			InstrAccumParallelQuery(&pvs->buffer_usage[i], &pvs->wal_usage[i]);
 	}
@@ -1010,6 +1021,9 @@ parallel_vacuum_process_safe_indexes(ParallelVacuumState *pvs)
 		if (!indstats->parallel_workers_can_process)
 			continue;
 
+		if (IsParallelWorker())
+			INJECTION_POINT("parallel-autovacuum-worker-before-index", NULL);
+
 		/* Do vacuum or cleanup of the index */
 		parallel_vacuum_process_one_index(pvs, pvs->indrels[idx], indstats);
 	}
diff --git a/src/test/modules/test_autovacuum/t/002_cost_reload_while_waiting.pl b/src/test/modules/test_autovacuum/t/002_cost_reload_while_waiting.pl
new file mode 100644
index 0000000000..afc3a73367
--- /dev/null
+++ b/src/test/modules/test_autovacuum/t/002_cost_reload_while_waiting.pl
@@ -0,0 +1,109 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+use strict;
+use warnings FATAL => 'all';
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+	plan skip_all => 'Injection points not supported by this build';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+$node->append_conf(
+	'postgresql.conf', qq{
+autovacuum_max_workers = 1
+autovacuum_max_parallel_workers = 1
+autovacuum_naptime = '1s'
+autovacuum_vacuum_cost_delay = '20ms'
+autovacuum_vacuum_cost_limit = 200
+log_min_messages = debug2
+min_parallel_index_scan_size = 0
+});
+$node->start;
+
+if (!$node->check_extension('injection_points'))
+{
+	plan skip_all => 'Extension injection_points not installed';
+}
+
+$node->safe_psql(
+	'postgres', q{
+	CREATE EXTENSION injection_points;
+	CREATE TABLE test_autovac (id int, a int)
+		WITH (autovacuum_enabled = false,
+			autovacuum_parallel_workers = 1,
+			autovacuum_vacuum_threshold = 0,
+			autovacuum_vacuum_scale_factor = 0);
+	INSERT INTO test_autovac
+		SELECT g, g FROM generate_series(1, 100) g;
+	CREATE INDEX test_autovac_id_idx ON test_autovac (id);
+	CREATE INDEX test_autovac_a_idx ON test_autovac (a);
+	UPDATE test_autovac SET a = a + 1;
+	SELECT injection_points_attach(
+		'parallel-autovacuum-worker-before-index', 'wait');
+	SELECT injection_points_attach(
+		'parallel-autovacuum-leader-before-index', 'wait');
+	SELECT injection_points_attach(
+		'parallel-autovacuum-leader-after-worker-wait', 'notice');
+	ALTER TABLE test_autovac SET (autovacuum_enabled = true);
+});
+
+$node->wait_for_event('autovacuum worker',
+	'parallel-autovacuum-leader-before-index');
+$node->wait_for_event('parallel worker',
+	'parallel-autovacuum-worker-before-index');
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_wakeup(
+		'parallel-autovacuum-leader-before-index');
+	SELECT injection_points_detach(
+		'parallel-autovacuum-leader-before-index');
+});
+$node->poll_query_until(
+	'postgres', q{
+	SELECT EXISTS (
+		SELECT 1
+		FROM pg_stat_activity
+		WHERE backend_type = 'autovacuum worker'
+		AND wait_event = 'ParallelFinish')
+}) or die "autovacuum leader did not reach ParallelFinish";
+
+my $log_offset = -s $node->logfile;
+if (!$ENV{NO_RELOAD_CONTROL})
+{
+	$node->safe_psql(
+		'postgres', q{
+		ALTER SYSTEM SET autovacuum_vacuum_cost_delay = 0;
+		SELECT pg_reload_conf();
+	});
+}
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_wakeup(
+		'parallel-autovacuum-worker-before-index');
+	SELECT injection_points_detach(
+		'parallel-autovacuum-worker-before-index');
+});
+
+$node->wait_for_log(
+	qr/parallel-autovacuum-leader-after-worker-wait \(reload (?:pending|processed)\)/,
+	$log_offset);
+
+my $log = slurp_file($node->logfile, $log_offset);
+my ($reload_state) =
+	$log =~ /parallel-autovacuum-leader-after-worker-wait \(reload (pending|processed)\)/;
+is($reload_state, 'processed',
+	'autovacuum leader processes a configuration reload while waiting');
+
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_detach(
+		'parallel-autovacuum-leader-after-worker-wait');
+});
+$node->stop;
+
+done_testing();

base-commit: b368bdd230181c60e085267fe65d42fd372d3a24
-- 
2.50.1 (Apple Git-155)



  [application/x-patch] 0001-Refresh-autovacuum-costs-while-waiting-for-parallel-.patch (14.2K, ../../CAM527d-GL=Jp2EJXBSnVBGPK-4XEZwWof5Cv8P0hghS_og6oAg@mail.gmail.com/3-0001-Refresh-autovacuum-costs-while-waiting-for-parallel-.patch)
  download | inline diff:
From 6e7d8a2bf027ffbbda1287264b3e88dc64917dea Mon Sep 17 00:00:00 2001
From: Nik Samokhvalov <nik@postgres.ai>
Date: Mon, 21 Sep 2026 08:57:49 -0700
Subject: [PATCH] Refresh autovacuum costs while waiting for parallel workers

---
 src/backend/access/transam/parallel.c         |  37 +++-
 src/backend/commands/vacuumparallel.c         |  47 ++++-
 src/include/access/parallel.h                 |   4 +
 .../t/001_parallel_autovacuum.pl              | 163 ++++++++++++++++++
 4 files changed, 246 insertions(+), 5 deletions(-)

diff --git a/src/backend/access/transam/parallel.c b/src/backend/access/transam/parallel.c
index 89e9d224ee..b222588e66 100644
--- a/src/backend/access/transam/parallel.c
+++ b/src/backend/access/transam/parallel.c
@@ -801,14 +801,17 @@ WaitForParallelWorkersToAttach(ParallelContext *pcxt)
  * Also, we want to update our notion of XactLastRecEnd based on worker
  * feedback.
  */
-void
-WaitForParallelWorkersToFinish(ParallelContext *pcxt)
+static void
+WaitForParallelWorkersToFinishInternal(ParallelContext *pcxt,
+									   ParallelWorkerWaitCallback callback,
+									   void *arg, long wait_interval)
 {
 	for (;;)
 	{
 		bool		anyone_alive = false;
 		int			nfinished = 0;
 		int			i;
+		uint32		wait_events = WL_LATCH_SET | WL_EXIT_ON_PM_DEATH;
 
 		/*
 		 * This will process any parallel messages that are pending, which may
@@ -816,6 +819,8 @@ WaitForParallelWorkersToFinish(ParallelContext *pcxt)
 		 * error propagated from a worker.
 		 */
 		CHECK_FOR_INTERRUPTS();
+		if (callback != NULL)
+			callback(arg);
 
 		for (i = 0; i < pcxt->nworkers_launched; ++i)
 		{
@@ -892,7 +897,11 @@ WaitForParallelWorkersToFinish(ParallelContext *pcxt)
 			}
 		}
 
-		(void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, -1,
+		if (wait_interval >= 0)
+			wait_events |= WL_TIMEOUT;
+
+		(void) WaitLatch(MyLatch, wait_events,
+						 wait_interval,
 						 WAIT_EVENT_PARALLEL_FINISH);
 		ResetLatch(MyLatch);
 	}
@@ -907,6 +916,28 @@ WaitForParallelWorkersToFinish(ParallelContext *pcxt)
 	}
 }
 
+void
+WaitForParallelWorkersToFinish(ParallelContext *pcxt)
+{
+	WaitForParallelWorkersToFinishInternal(pcxt, NULL, NULL, -1);
+}
+
+/*
+ * As above, but run callback before checking worker state and after each
+ * wakeup.  While workers remain active, wake at least once every
+ * wait_interval milliseconds.
+ */
+void
+WaitForParallelWorkersToFinishWithCallback(ParallelContext *pcxt,
+										   ParallelWorkerWaitCallback callback,
+										   void *arg, long wait_interval)
+{
+	Assert(callback != NULL);
+	Assert(wait_interval > 0);
+
+	WaitForParallelWorkersToFinishInternal(pcxt, callback, arg, wait_interval);
+}
+
 /*
  * Wait for all workers to exit.
  *
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 767d162e57..b5d30f2152 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -43,9 +43,11 @@
 #include "executor/instrument.h"
 #include "optimizer/paths.h"
 #include "pgstat.h"
+#include "postmaster/interrupt.h"
 #include "storage/bufmgr.h"
 #include "storage/proc.h"
 #include "tcop/tcopprot.h"
+#include "utils/injection_point.h"
 #include "utils/lsyscache.h"
 #include "utils/rel.h"
 
@@ -60,6 +62,8 @@
 #define PARALLEL_VACUUM_KEY_WAL_USAGE		4
 #define PARALLEL_VACUUM_KEY_INDEX_STATS		5
 
+#define PARALLEL_VACUUM_COST_UPDATE_INTERVAL_MS 100
+
 /*
  * Struct for cost-based vacuum delay related parameters to share among an
  * autovacuum worker and its parallel vacuum workers.
@@ -285,6 +289,7 @@ static int	parallel_vacuum_compute_workers(Relation *indrels, int nindexes, int
 											bool *will_parallel_vacuum);
 static void parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scans,
 												bool vacuum, PVWorkerStats *wstats);
+static void parallel_vacuum_update_leader_cost_params(void *arg);
 static void parallel_vacuum_process_safe_indexes(ParallelVacuumState *pvs);
 static void parallel_vacuum_process_unsafe_indexes(ParallelVacuumState *pvs);
 static void parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
@@ -929,6 +934,10 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 	/* Vacuum the indexes that can be processed by only leader process */
 	parallel_vacuum_process_unsafe_indexes(pvs);
 
+	if (pvs->shared->is_autovacuum &&
+		pvs->pcxt->nworkers_launched > 0)
+		INJECTION_POINT("parallel-vacuum-leader-before-index", NULL);
+
 	/*
 	 * Join as a parallel worker.  The leader vacuums alone processes all
 	 * parallel-safe indexes in the case where no workers are launched.
@@ -941,8 +950,14 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 	 */
 	if (nworkers > 0)
 	{
-		/* Wait for all vacuum workers to finish */
-		WaitForParallelWorkersToFinish(pvs->pcxt);
+		/* Wait for all vacuum workers to finish. */
+		if (AmAutoVacuumWorkerProcess())
+			WaitForParallelWorkersToFinishWithCallback(pvs->pcxt,
+													   parallel_vacuum_update_leader_cost_params,
+													   NULL,
+													   PARALLEL_VACUUM_COST_UPDATE_INTERVAL_MS);
+		else
+			WaitForParallelWorkersToFinish(pvs->pcxt);
 
 		for (int i = 0; i < pvs->pcxt->nworkers_launched; i++)
 			InstrAccumParallelQuery(&pvs->buffer_usage[i], &pvs->wal_usage[i]);
@@ -974,6 +989,29 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 	}
 }
 
+/*
+ * Refresh the leader's cost parameters while it waits for parallel vacuum
+ * workers, and propagate any changes to them.
+ */
+static void
+parallel_vacuum_update_leader_cost_params(void *arg)
+{
+	Assert(AmAutoVacuumWorkerProcess());
+	Assert(arg == NULL);
+
+	if (ConfigReloadPending)
+	{
+		ConfigReloadPending = false;
+		ProcessConfigFile(PGC_SIGHUP);
+		VacuumUpdateCosts();
+	}
+	else
+		AutoVacuumUpdateCostLimit();
+
+	parallel_vacuum_propagate_shared_delay_params();
+	INJECTION_POINT("parallel-vacuum-leader-cost-updated", NULL);
+}
+
 /*
  * Index vacuum/cleanup routine used by the leader process and parallel
  * vacuum worker processes to vacuum the indexes in parallel.
@@ -1097,6 +1135,11 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
 	pvs->indname = pstrdup(RelationGetRelationName(indrel));
 	pvs->status = indstats->status;
 
+#ifdef USE_INJECTION_POINTS
+	if (IsParallelWorker())
+		INJECTION_POINT("parallel-vacuum-worker-before-index", NULL);
+#endif
+
 	switch (indstats->status)
 	{
 		case PARALLEL_INDVAC_STATUS_NEED_BULKDELETE:
diff --git a/src/include/access/parallel.h b/src/include/access/parallel.h
index 60f857675e..4ca5a80b1a 100644
--- a/src/include/access/parallel.h
+++ b/src/include/access/parallel.h
@@ -23,6 +23,7 @@
 #include "storage/shm_toc.h"
 
 typedef void (*parallel_worker_main_type) (dsm_segment *seg, shm_toc *toc);
+typedef void (*ParallelWorkerWaitCallback) (void *arg);
 
 typedef struct ParallelWorkerInfo
 {
@@ -69,6 +70,9 @@ extern void ReinitializeParallelWorkers(ParallelContext *pcxt, int nworkers_to_l
 extern void LaunchParallelWorkers(ParallelContext *pcxt);
 extern void WaitForParallelWorkersToAttach(ParallelContext *pcxt);
 extern void WaitForParallelWorkersToFinish(ParallelContext *pcxt);
+extern void WaitForParallelWorkersToFinishWithCallback(ParallelContext *pcxt,
+													   ParallelWorkerWaitCallback callback,
+													   void *arg, long wait_interval);
 extern void DestroyParallelContext(ParallelContext *pcxt);
 extern bool ParallelContextActive(void);
 
diff --git a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
index 33c86bbdc9..ad84e66758 100644
--- a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
+++ b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
@@ -252,6 +252,169 @@ $node->safe_psql('postgres',
 	"SELECT injection_points_wakeup('autovacuum-worker-cost-balanced')");
 $node->safe_psql('postgres',
 	"SELECT injection_points_detach('autovacuum-worker-cost-balanced')");
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) = 0 FROM pg_stat_activity
+	WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
+}) or die "second autovacuum worker did not finish";
+
+# Test 4:
+# Check whether a config reload is serviced while the autovacuum leader waits
+# for a parallel worker to finish an index.  Hold the worker after it claims
+# an index, so the leader can process the remaining indexes and enter
+# ParallelFinish.
+my $postgresoid = $node->safe_psql('postgres',
+	"SELECT oid FROM pg_database WHERE datname = 'postgres'");
+my $testautovacid =
+  $node->safe_psql('postgres', "SELECT 'test_autovac'::regclass::oid");
+
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_max_workers = 1;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 700;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_delay = 0;
+	SELECT pg_reload_conf();
+});
+
+prepare_for_next_test($node, 4);
+$log_offset = -s $node->logfile;
+
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_attach('parallel-vacuum-worker-before-index', 'wait');
+	SELECT injection_points_attach('parallel-vacuum-leader-before-index', 'wait');
+});
+$node->safe_psql('postgres',
+	'ALTER TABLE test_autovac SET (autovacuum_enabled = true)');
+$node->wait_for_event('autovacuum worker',
+	'parallel-vacuum-leader-before-index');
+$node->wait_for_event('parallel worker',
+	'parallel-vacuum-worker-before-index');
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_wakeup('parallel-vacuum-leader-before-index');
+	SELECT injection_points_detach('parallel-vacuum-leader-before-index');
+});
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) > 0
+	FROM pg_stat_activity worker
+	JOIN pg_stat_activity leader ON leader.pid = worker.leader_pid
+	WHERE worker.backend_type = 'parallel worker'
+	  AND worker.wait_event = 'parallel-vacuum-worker-before-index'
+	  AND leader.wait_event = 'ParallelFinish'
+}) or die "autovacuum leader did not enter ParallelFinish";
+
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 800;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_delay = 8;
+	ALTER SYSTEM SET vacuum_cost_page_miss = 11;
+	ALTER SYSTEM SET vacuum_cost_page_dirty = 12;
+	ALTER SYSTEM SET vacuum_cost_page_hit = 13;
+	SELECT pg_reload_conf();
+});
+
+$node->wait_for_log(
+	qr/Autovacuum VacuumUpdateCosts\(db=$postgresoid, rel=$testautovacid, dobalance=yes, cost_limit=800, cost_delay=8 /,
+	$log_offset);
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('parallel-vacuum-worker-before-index')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('parallel-vacuum-worker-before-index')");
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=800, cost_delay=8, cost_page_miss=11, cost_page_dirty=12, cost_page_hit=13/,
+	$log_offset);
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+ok(1, "config reload is propagated while the leader waits for workers");
+
+# Test 5:
+# Check the same wait path for an un-signalled cost-limit rebalance.  A second
+# autovacuum worker joins the balance while the first leader and its parallel
+# worker remain held.
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_max_workers = 2;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 600;
+	SELECT pg_reload_conf();
+});
+
+prepare_for_next_test($node, 5);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = false)');
+$node->safe_psql('regress_db2', 'UPDATE filler SET id = id + 1');
+
+$log_offset = -s $node->logfile;
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_attach('parallel-vacuum-worker-before-index', 'wait');
+	SELECT injection_points_attach('parallel-vacuum-leader-before-index', 'wait');
+});
+$node->safe_psql('postgres',
+	'ALTER TABLE test_autovac SET (autovacuum_enabled = true)');
+$node->wait_for_event('autovacuum worker',
+	'parallel-vacuum-leader-before-index');
+$node->wait_for_event('parallel worker',
+	'parallel-vacuum-worker-before-index');
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_wakeup('parallel-vacuum-leader-before-index');
+	SELECT injection_points_detach('parallel-vacuum-leader-before-index');
+});
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) > 0
+	FROM pg_stat_activity worker
+	JOIN pg_stat_activity leader ON leader.pid = worker.leader_pid
+	WHERE worker.backend_type = 'parallel worker'
+	  AND worker.wait_event = 'parallel-vacuum-worker-before-index'
+	  AND leader.wait_event = 'ParallelFinish'
+}) or die "autovacuum leader did not enter ParallelFinish";
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_attach('autovacuum-worker-cost-balanced', 'wait')"
+);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = true)');
+$node->wait_for_log(
+	qr/VacuumUpdateCosts\(db=$db2oid, rel=$filleroid, dobalance=yes, cost_limit=300,/,
+	$log_offset);
+$node->safe_psql('postgres',
+	"SELECT injection_points_attach('parallel-vacuum-leader-cost-updated', 'notice')"
+);
+$node->wait_for_log(
+	qr/notice triggered for injection point parallel-vacuum-leader-cost-updated/,
+	$log_offset);
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('parallel-vacuum-leader-cost-updated')");
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('parallel-vacuum-worker-before-index')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('parallel-vacuum-worker-before-index')");
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=300,/,
+	$log_offset);
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('autovacuum-worker-cost-balanced')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('autovacuum-worker-cost-balanced')");
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) = 0 FROM pg_stat_activity
+	WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
+}) or die "second autovacuum worker did not finish";
+ok(1, "cost rebalance is propagated while the leader waits for workers");
 
 $node->stop;
 done_testing();

base-commit: b368bdd230181c60e085267fe65d42fd372d3a24
-- 
2.50.1 (Apple Git-155)



^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-24 00:38  Masahiko Sawada <sawada.mshk@gmail.com>
  parent: Nikolay Samokhvalov <nik@postgres.ai>
  0 siblings, 2 replies; 43+ messages in thread

From: Masahiko Sawada @ 2026-09-24 00:38 UTC (permalink / raw)
  To: Nikolay Samokhvalov <nik@postgres.ai>; +Cc: Daniel Gustafsson <daniel@yesql.se>; Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

Hi,

Sorry for the late reply. I didn't realize your report until now.

On Mon, Sep 21, 2026 at 11:28 AM Nikolay Samokhvalov <nik@postgres.ai> wrote:
>
> On Thu, Aug 27, 2026 at 3:43 PM Masahiko Sawada <sawada.mshk@gmail.com> wrote:
> > Thank you for taking care of it. The v3 patch looks good to me.
>
> AI found one more gap while testing the committed fix on REL_19_STABLE at
> b368bdd2. I haven't manually reviewed the code yet.
>
> Once an autovacuum leader enters WaitForParallelWorkersToFinish(), SIGHUP
> wakes its latch, but the loop only runs CHECK_FOR_INTERRUPTS(), leaving
> ConfigReloadPending set. A cost-limit rebalance is not signalled at all. In
> both cases, the leader does not publish changed cost parameters until the
> parallel worker finishes.

Good catch, we should fix it.

> The other attached patch adds an optional callback to the worker-finish
> wait. Parallel autovacuum uses it to handle reloads and poll cost-limit
> rebalancing every 100 ms; other callers keep the existing behavior. It also
> adds tests for both cases to 001_parallel_autovacuum.pl.

I've confirmed that the patch fixes the issue. While it works fine,
I'm a bit concerned that adding
WaitForParallelWorkersToFinishWithCallback() with a callback and a
timeout might be overkill, as I don't see any usecase other than
parallel autovacuum that needs to pass a callback.

An alternative approach would be to have a function in
vacuumparallel.c that waits for all index statuses to become
PARALLEL_INDVAC_STATUS_COMPLETED while periodically checking for cost
parameter updates. We still need the timeout there since nothing wakes
up the leader on a cost limit rebalance. While it adds another wait
loop before WaitForParallelWorkersToFinish(), that call should return
almost immediately. We can consider adding a callback to
WaitForParallelWorkersToFinish() when we find other use cases in the
future.

Regards,

-- 
Masahiko Sawada
Amazon Web Services: https://aws.amazon.com






^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-24 04:44  Nikolay Samokhvalov <nik@postgres.ai>
  parent: Masahiko Sawada <sawada.mshk@gmail.com>
  1 sibling, 0 replies; 43+ messages in thread

From: Nikolay Samokhvalov @ 2026-09-24 04:44 UTC (permalink / raw)
  To: Masahiko Sawada <sawada.mshk@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

On Wed, Sep 23, 2026 at 5:39 PM Masahiko Sawada <sawada.mshk@gmail.com> wrote:
> I'm a bit concerned that adding
> WaitForParallelWorkersToFinishWithCallback() with a callback and a
> timeout might be overkill, as I don't see any usecase other than
> parallel autovacuum that needs to pass a callback.
>
> An alternative approach would be to have a function in
> vacuumparallel.c that waits for all index statuses to become
> PARALLEL_INDVAC_STATUS_COMPLETED while periodically checking for cost
> parameter updates.

Thanks for testing it. Agreed; attached v2 keeps the timed wait in
vacuumparallel.c. It uses an atomic completion count rather than polling
the index statuses while workers update them, then calls the existing
WaitForParallelWorkersToFinish() for final worker errors and WAL feedback.

On REL_19_STABLE at b73d13c3, test_autovacuum, core regression, and
isolation pass.

Nik

Attachments:

  [application/x-patch] v2-0001-Refresh-autovacuum-costs-while-parallel-indexes-finish.patch (13.8K, ../../CAM527d8V7RYT3iXr2Zs8OaeKXprTVpxT3mjsynk_OA7MkMc_LA@mail.gmail.com/2-v2-0001-Refresh-autovacuum-costs-while-parallel-indexes-finish.patch)
  download | inline diff:
From 7815bb3ba93acfd0a7ad3899be726a886d6db4dd Mon Sep 17 00:00:00 2001
From: Nik Samokhvalov <nik@postgres.ai>
Date: Wed, 23 Sep 2026 18:34:32 -0700
Subject: [PATCH] Refresh autovacuum costs while parallel indexes finish

Use an atomic completion count to keep the leader responsive while parallel index workers run. Preserve the normal parallel worker wait for final errors and WAL feedback. Add injection-point tests for reload and cost rebalance.
---
 src/backend/commands/vacuumparallel.c         | 102 ++++++++++-
 .../t/001_parallel_autovacuum.pl              | 163 ++++++++++++++++++
 2 files changed, 264 insertions(+), 1 deletion(-)

diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 767d162e57..9124caf5b9 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -43,11 +43,14 @@
 #include "executor/instrument.h"
 #include "optimizer/paths.h"
 #include "pgstat.h"
+#include "postmaster/interrupt.h"
 #include "storage/bufmgr.h"
 #include "storage/proc.h"
 #include "tcop/tcopprot.h"
+#include "utils/injection_point.h"
 #include "utils/lsyscache.h"
 #include "utils/rel.h"
+#include "utils/wait_event.h"
 
 /*
  * DSM keys for parallel vacuum.  Unlike other parallel execution code, since
@@ -60,6 +63,8 @@
 #define PARALLEL_VACUUM_KEY_WAL_USAGE		4
 #define PARALLEL_VACUUM_KEY_INDEX_STATS		5
 
+#define PARALLEL_VACUUM_COST_UPDATE_INTERVAL_MS 100
+
 /*
  * Struct for cost-based vacuum delay related parameters to share among an
  * autovacuum worker and its parallel vacuum workers.
@@ -148,6 +153,9 @@ typedef struct PVShared
 	/* Counter for vacuuming and cleanup */
 	pg_atomic_uint32 idx;
 
+	/* Number of indexes completed in the current phase */
+	pg_atomic_uint32 completed_indexes;
+
 	/* DSA handle where the TidStore lives */
 	dsa_handle	dead_items_dsa_handle;
 
@@ -285,6 +293,8 @@ static int	parallel_vacuum_compute_workers(Relation *indrels, int nindexes, int
 											bool *will_parallel_vacuum);
 static void parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scans,
 												bool vacuum, PVWorkerStats *wstats);
+static void parallel_vacuum_update_leader_cost_params(void);
+static void parallel_vacuum_wait_for_indexes(ParallelVacuumState *pvs);
 static void parallel_vacuum_process_safe_indexes(ParallelVacuumState *pvs);
 static void parallel_vacuum_process_unsafe_indexes(ParallelVacuumState *pvs);
 static void parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
@@ -453,6 +463,7 @@ parallel_vacuum_init(Relation rel, Relation *indrels, int nindexes,
 	pg_atomic_init_u32(&(shared->cost_balance), 0);
 	pg_atomic_init_u32(&(shared->active_nworkers), 0);
 	pg_atomic_init_u32(&(shared->idx), 0);
+	pg_atomic_init_u32(&(shared->completed_indexes), 0);
 
 	shared->is_autovacuum = AmAutoVacuumWorkerProcess();
 
@@ -869,6 +880,7 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 
 	/* Reset the parallel index processing and progress counters */
 	pg_atomic_write_u32(&(pvs->shared->idx), 0);
+	pg_atomic_write_u32(&(pvs->shared->completed_indexes), 0);
 
 	/* Setup the shared cost-based vacuum delay and launch workers */
 	if (nworkers > 0)
@@ -929,6 +941,10 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 	/* Vacuum the indexes that can be processed by only leader process */
 	parallel_vacuum_process_unsafe_indexes(pvs);
 
+	if (pvs->shared->is_autovacuum &&
+		pvs->pcxt->nworkers_launched > 0)
+		INJECTION_POINT("parallel-vacuum-leader-before-index", NULL);
+
 	/*
 	 * Join as a parallel worker.  The leader vacuums alone processes all
 	 * parallel-safe indexes in the case where no workers are launched.
@@ -941,7 +957,9 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 	 */
 	if (nworkers > 0)
 	{
-		/* Wait for all vacuum workers to finish */
+		/* Wait for all vacuum workers to finish. */
+		if (AmAutoVacuumWorkerProcess())
+			parallel_vacuum_wait_for_indexes(pvs);
 		WaitForParallelWorkersToFinish(pvs->pcxt);
 
 		for (int i = 0; i < pvs->pcxt->nworkers_launched; i++)
@@ -974,6 +992,82 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 	}
 }
 
+/*
+ * Refresh the leader's cost parameters while it waits for parallel vacuum
+ * workers, and propagate any changes to them.
+ */
+static void
+parallel_vacuum_update_leader_cost_params(void)
+{
+	Assert(AmAutoVacuumWorkerProcess());
+
+	if (ConfigReloadPending)
+	{
+		ConfigReloadPending = false;
+		ProcessConfigFile(PGC_SIGHUP);
+		VacuumUpdateCosts();
+	}
+	else
+		AutoVacuumUpdateCostLimit();
+
+	parallel_vacuum_propagate_shared_delay_params();
+	INJECTION_POINT("parallel-vacuum-leader-cost-updated", NULL);
+}
+
+/*
+ * Keep autovacuum cost parameters current while workers vacuum indexes.
+ * Worker completion is counted atomically after publishing index results;
+ * the ordinary parallel-worker wait still handles final errors and cleanup.
+ */
+static void
+parallel_vacuum_wait_for_indexes(ParallelVacuumState *pvs)
+{
+	ParallelContext *pcxt = pvs->pcxt;
+
+	while (pg_atomic_read_membarrier_u32(&(pvs->shared->completed_indexes)) <
+		   pvs->nindexes)
+	{
+		int			nfinished = 0;
+
+		CHECK_FOR_INTERRUPTS();
+		parallel_vacuum_update_leader_cost_params();
+
+		for (int i = 0; i < pcxt->nworkers_launched; i++)
+		{
+			pid_t		pid;
+			shm_mq	   *mq;
+
+			if (pcxt->worker[i].error_mqh == NULL)
+			{
+				nfinished++;
+				continue;
+			}
+
+			if (pcxt->worker[i].bgwhandle == NULL ||
+				GetBackgroundWorkerPid(pcxt->worker[i].bgwhandle, &pid) !=
+				BGWH_STOPPED)
+				continue;
+
+			mq = shm_mq_get_queue(pcxt->worker[i].error_mqh);
+			if (shm_mq_get_sender(mq) == NULL)
+				ereport(ERROR,
+						(errcode(ERRCODE_OBJECT_NOT_IN_PREREQUISITE_STATE),
+						 errmsg("parallel worker failed to initialize"),
+						 errhint("More details may be available in the server log.")));
+		}
+
+		/* Let the final index-status check report any missing index. */
+		if (nfinished == pcxt->nworkers_launched)
+			break;
+
+		(void) WaitLatch(MyLatch,
+					 WL_LATCH_SET | WL_TIMEOUT | WL_EXIT_ON_PM_DEATH,
+					 PARALLEL_VACUUM_COST_UPDATE_INTERVAL_MS,
+					 WAIT_EVENT_PARALLEL_FINISH);
+		ResetLatch(MyLatch);
+	}
+}
+
 /*
  * Index vacuum/cleanup routine used by the leader process and parallel
  * vacuum worker processes to vacuum the indexes in parallel.
@@ -1097,6 +1191,11 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
 	pvs->indname = pstrdup(RelationGetRelationName(indrel));
 	pvs->status = indstats->status;
 
+#ifdef USE_INJECTION_POINTS
+	if (IsParallelWorker())
+		INJECTION_POINT("parallel-vacuum-worker-before-index", NULL);
+#endif
+
 	switch (indstats->status)
 	{
 		case PARALLEL_INDVAC_STATUS_NEED_BULKDELETE:
@@ -1138,6 +1237,7 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
 	 * touches different indexes.
 	 */
 	indstats->status = PARALLEL_INDVAC_STATUS_COMPLETED;
+	pg_atomic_fetch_add_u32(&(pvs->shared->completed_indexes), 1);
 
 	/* Reset error traceback information */
 	pvs->status = PARALLEL_INDVAC_STATUS_COMPLETED;
diff --git a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
index 33c86bbdc9..ad84e66758 100644
--- a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
+++ b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
@@ -252,6 +252,169 @@ $node->safe_psql('postgres',
 	"SELECT injection_points_wakeup('autovacuum-worker-cost-balanced')");
 $node->safe_psql('postgres',
 	"SELECT injection_points_detach('autovacuum-worker-cost-balanced')");
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) = 0 FROM pg_stat_activity
+	WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
+}) or die "second autovacuum worker did not finish";
+
+# Test 4:
+# Check whether a config reload is serviced while the autovacuum leader waits
+# for a parallel worker to finish an index.  Hold the worker after it claims
+# an index, so the leader can process the remaining indexes and enter
+# ParallelFinish.
+my $postgresoid = $node->safe_psql('postgres',
+	"SELECT oid FROM pg_database WHERE datname = 'postgres'");
+my $testautovacid =
+  $node->safe_psql('postgres', "SELECT 'test_autovac'::regclass::oid");
+
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_max_workers = 1;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 700;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_delay = 0;
+	SELECT pg_reload_conf();
+});
+
+prepare_for_next_test($node, 4);
+$log_offset = -s $node->logfile;
+
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_attach('parallel-vacuum-worker-before-index', 'wait');
+	SELECT injection_points_attach('parallel-vacuum-leader-before-index', 'wait');
+});
+$node->safe_psql('postgres',
+	'ALTER TABLE test_autovac SET (autovacuum_enabled = true)');
+$node->wait_for_event('autovacuum worker',
+	'parallel-vacuum-leader-before-index');
+$node->wait_for_event('parallel worker',
+	'parallel-vacuum-worker-before-index');
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_wakeup('parallel-vacuum-leader-before-index');
+	SELECT injection_points_detach('parallel-vacuum-leader-before-index');
+});
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) > 0
+	FROM pg_stat_activity worker
+	JOIN pg_stat_activity leader ON leader.pid = worker.leader_pid
+	WHERE worker.backend_type = 'parallel worker'
+	  AND worker.wait_event = 'parallel-vacuum-worker-before-index'
+	  AND leader.wait_event = 'ParallelFinish'
+}) or die "autovacuum leader did not enter ParallelFinish";
+
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 800;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_delay = 8;
+	ALTER SYSTEM SET vacuum_cost_page_miss = 11;
+	ALTER SYSTEM SET vacuum_cost_page_dirty = 12;
+	ALTER SYSTEM SET vacuum_cost_page_hit = 13;
+	SELECT pg_reload_conf();
+});
+
+$node->wait_for_log(
+	qr/Autovacuum VacuumUpdateCosts\(db=$postgresoid, rel=$testautovacid, dobalance=yes, cost_limit=800, cost_delay=8 /,
+	$log_offset);
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('parallel-vacuum-worker-before-index')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('parallel-vacuum-worker-before-index')");
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=800, cost_delay=8, cost_page_miss=11, cost_page_dirty=12, cost_page_hit=13/,
+	$log_offset);
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+ok(1, "config reload is propagated while the leader waits for workers");
+
+# Test 5:
+# Check the same wait path for an un-signalled cost-limit rebalance.  A second
+# autovacuum worker joins the balance while the first leader and its parallel
+# worker remain held.
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_max_workers = 2;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 600;
+	SELECT pg_reload_conf();
+});
+
+prepare_for_next_test($node, 5);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = false)');
+$node->safe_psql('regress_db2', 'UPDATE filler SET id = id + 1');
+
+$log_offset = -s $node->logfile;
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_attach('parallel-vacuum-worker-before-index', 'wait');
+	SELECT injection_points_attach('parallel-vacuum-leader-before-index', 'wait');
+});
+$node->safe_psql('postgres',
+	'ALTER TABLE test_autovac SET (autovacuum_enabled = true)');
+$node->wait_for_event('autovacuum worker',
+	'parallel-vacuum-leader-before-index');
+$node->wait_for_event('parallel worker',
+	'parallel-vacuum-worker-before-index');
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_wakeup('parallel-vacuum-leader-before-index');
+	SELECT injection_points_detach('parallel-vacuum-leader-before-index');
+});
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) > 0
+	FROM pg_stat_activity worker
+	JOIN pg_stat_activity leader ON leader.pid = worker.leader_pid
+	WHERE worker.backend_type = 'parallel worker'
+	  AND worker.wait_event = 'parallel-vacuum-worker-before-index'
+	  AND leader.wait_event = 'ParallelFinish'
+}) or die "autovacuum leader did not enter ParallelFinish";
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_attach('autovacuum-worker-cost-balanced', 'wait')"
+);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = true)');
+$node->wait_for_log(
+	qr/VacuumUpdateCosts\(db=$db2oid, rel=$filleroid, dobalance=yes, cost_limit=300,/,
+	$log_offset);
+$node->safe_psql('postgres',
+	"SELECT injection_points_attach('parallel-vacuum-leader-cost-updated', 'notice')"
+);
+$node->wait_for_log(
+	qr/notice triggered for injection point parallel-vacuum-leader-cost-updated/,
+	$log_offset);
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('parallel-vacuum-leader-cost-updated')");
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('parallel-vacuum-worker-before-index')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('parallel-vacuum-worker-before-index')");
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=300,/,
+	$log_offset);
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('autovacuum-worker-cost-balanced')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('autovacuum-worker-cost-balanced')");
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) = 0 FROM pg_stat_activity
+	WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
+}) or die "second autovacuum worker did not finish";
+ok(1, "cost rebalance is propagated while the leader waits for workers");
 
 $node->stop;
 done_testing();


^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-24 06:34  Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
  parent: Masahiko Sawada <sawada.mshk@gmail.com>
  1 sibling, 1 reply; 43+ messages in thread

From: Bharath Rupireddy @ 2026-09-24 06:34 UTC (permalink / raw)
  To: Masahiko Sawada <sawada.mshk@gmail.com>; +Cc: Nikolay Samokhvalov <nik@postgres.ai>; Daniel Gustafsson <daniel@yesql.se>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

Hi,

On Wed, Sep 23, 2026 at 5:39 PM Masahiko Sawada <sawada.mshk@gmail.com>
wrote:
>
> Good catch, we should fix it.
>
> > The other attached patch adds an optional callback to the worker-finish
> > wait. Parallel autovacuum uses it to handle reloads and poll cost-limit
> > rebalancing every 100 ms; other callers keep the existing behavior. It
also
> > adds tests for both cases to 001_parallel_autovacuum.pl.
>
> I've confirmed that the patch fixes the issue. While it works fine,
> I'm a bit concerned that adding
> WaitForParallelWorkersToFinishWithCallback() with a callback and a
> timeout might be overkill, as I don't see any usecase other than
> parallel autovacuum that needs to pass a callback.
>
> An alternative approach would be to have a function in
> vacuumparallel.c that waits for all index statuses to become
> PARALLEL_INDVAC_STATUS_COMPLETED while periodically checking for cost
> parameter updates. We still need the timeout there since nothing wakes
> up the leader on a cost limit rebalance. While it adds another wait
> loop before WaitForParallelWorkersToFinish(), that call should return
> almost immediately. We can consider adding a callback to
> WaitForParallelWorkersToFinish() when we find other use cases in the
> future.

Thanks for reporting the issue. I believe this issue can happen fairly
often in practice, especially when the autovacuum worker (leader) gets
smaller or fewer indexes than the parallel workers for index vacuuming
(workers), and the workers take longer than the leader (for example,
non-core indexes that can take a while). So I think we do need to fix it.

A separate wait function would lose all the cases that the existing wait
loop for parallel workers, WaitForParallelWorkersToFinish(), already
handles, like detecting error messages reported by the workers, their
liveness checks, attach/detach, and so on. It would also likely end up
duplicating that same logic along with the cost param update code at the
end.

I also don't like the wait-100ms-wakeup approach that a separate wait
function would run all the time, since it wastes power and CPU cycles.
Imagine a worker vacuuming an index that is hundreds of GBs or even TBs,
while the leader has only a small index and finishes first. The leader then
sits in the wait loop for longer until the large index is done, waking up
every 100ms the whole time. So I would prefer not to go that route.

How about doing the check inside the existing wait loop, only when the
process is an autovacuum worker, along the lines of the attached WIP? This
is simple, the pattern already exists elsewhere in the code, and it looks
safe. I also believe the wait loop is not in a performance-critical hot
path, and this check is not costly. I checked that it still fixes the
reported issues.

What I am less sure about is the fix for the second issue in the attached
WIP patch, which needs any leader waiting for its workers to be woken up
after the cost limit is rebalanced. I haven't found a better one yet.

PS: I noticed the attached v2 patch posted upthread only after finishing my
response.

Thoughts?

--
Bharath Rupireddy

Amazon Web Services: https://aws.amazon.com

Attachments:

  [application/x-patch] nocfbot-0001-WIP-Refresh-autovacuum-cost-parameters-while-waiting-for.patch (5.8K, ../../CALj2ACWr98Q_Wt3Zus73dJ5DpJttsnaw0-1W2pqjRN9SPpbgfw@mail.gmail.com/3-nocfbot-0001-WIP-Refresh-autovacuum-cost-parameters-while-waiting-for.patch)
  download | inline diff:
From 6f37d18f2145ccd86f5deb4daf94f0ba6536517f Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Wed, 23 Sep 2026 22:02:08 -0700
Subject: [PATCH] Refresh autovacuum cost parameters while waiting for parallel
 workers

An autovacuum leader publishes its cost-based delay parameters to its
parallel vacuum workers at its cost delay points. Once it has finished
its own share of the indexes and is only waiting for its workers, it
has no delay points left, so it stops publishing for as long as that
wait lasts, which can be as long as the largest index takes. Two kinds
of changes were lost for that whole time.

1. A config reload. The signal wakes the leader, but the wait does not
   act on it, so the reload stays pending until the wait ends.

2. A change in the number of autovacuum workers sharing the cost limit.
   Nothing signals that at all, and a waiting leader never re-reads the
   count because it no longer naps.

Fix the first by refreshing and publishing the parameters from the wait
itself, on every wakeup, when the process is an autovacuum worker. Fix
the second by waking the workers that share the limit whenever the count
changes, so a waiting leader picks it up the same way. A worker that is
busy vacuuming ignores the extra wakeup.

Reported-by: Nikolay Samokhvalov <nik@postgres.ai>
Discussion: https://postgr.es/m/CAM527d-GL%3DJp2EJXBSnVBGPK-4XEZwWof5Cv8P0hghS_og6oAg%40mail.gmail.com
---
 src/backend/access/transam/parallel.c |  8 ++++++
 src/backend/commands/vacuumparallel.c | 40 +++++++++++++++++++++++++++
 src/backend/postmaster/autovacuum.c   | 19 +++++++++++++
 src/include/commands/vacuum.h         |  1 +
 4 files changed, 68 insertions(+)

diff --git a/src/backend/access/transam/parallel.c b/src/backend/access/transam/parallel.c
index e1806a9a28a..89e45bdeb9c 100644
--- a/src/backend/access/transam/parallel.c
+++ b/src/backend/access/transam/parallel.c
@@ -812,6 +812,14 @@ WaitForParallelWorkersToFinish(ParallelContext *pcxt)
 		 */
 		CHECK_FOR_INTERRUPTS();
 
+		/*
+		 * An autovacuum leader publishes cost parameter changes to its
+		 * parallel workers at its cost delay points, which it no longer
+		 * reaches while waiting here. Do it here instead, on every wakeup.
+		 */
+		if (AmAutoVacuumWorkerProcess())
+			parallel_vacuum_refresh_cost_params();
+
 		for (i = 0; i < pcxt->nworkers_launched; ++i)
 		{
 			/*
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 767d162e578..b2e87bb9433 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -43,6 +43,7 @@
 #include "executor/instrument.h"
 #include "optimizer/paths.h"
 #include "pgstat.h"
+#include "postmaster/interrupt.h"
 #include "storage/bufmgr.h"
 #include "storage/proc.h"
 #include "tcop/tcopprot.h"
@@ -725,6 +726,45 @@ parallel_vacuum_propagate_shared_delay_params(void)
 	pg_atomic_fetch_add_u32(&pv_shared_cost_params->generation, 1);
 }
 
+/*
+ * Refresh the leader's cost-based vacuum delay parameters and propagate them
+ * to its parallel vacuum workers.
+ *
+ * The leader normally does this at its cost delay points, which it no longer
+ * reaches once it is only waiting for its workers to finish. It calls this
+ * from that wait instead, on every wakeup, to pick up a config reload and a
+ * change in the number of autovacuum workers sharing the cost limit. Both of
+ * those wake the leader, the first by signal and the second when the count is
+ * recalculated.
+ */
+void
+parallel_vacuum_refresh_cost_params(void)
+{
+	Assert(AmAutoVacuumWorkerProcess());
+
+	/*
+	 * Quick return if the leader process is not sharing the delay parameters.
+	 */
+	if (pv_shared_cost_params == NULL)
+		return;
+
+	if (ConfigReloadPending)
+	{
+		ConfigReloadPending = false;
+		ProcessConfigFile(PGC_SIGHUP);
+
+		/* This rebalances the cost limit too */
+		VacuumUpdateCosts();
+	}
+	else
+	{
+		/* The number of workers sharing the cost limit may have changed */
+		AutoVacuumUpdateCostLimit();
+	}
+
+	parallel_vacuum_propagate_shared_delay_params();
+}
+
 /*
  * Compute the number of parallel worker processes to request.  Both index
  * vacuum and index cleanup can be executed with parallel workers.
diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c
index 60ebe828900..d45f361ac0d 100644
--- a/src/backend/postmaster/autovacuum.c
+++ b/src/backend/postmaster/autovacuum.c
@@ -1814,8 +1814,27 @@ autovac_recalculate_workers_for_balance(void)
 	}
 
 	if (nworkers_for_balance != orig_nworkers_for_balance)
+	{
 		pg_atomic_write_u32(&AutoVacuumShmem->av_nworkersForBalance,
 							nworkers_for_balance);
+
+		/*
+		 * Wake up the workers that share the limit. One that is vacuuming
+		 * picks up the new count on its next nap, but one that is only
+		 * waiting for its parallel vacuum workers never naps, and nothing
+		 * else wakes it while they keep running at the old share.
+		 */
+		dlist_foreach(iter, &AutoVacuumShmem->av_runningWorkers)
+		{
+			WorkerInfo	worker = dlist_container(WorkerInfoData, wi_links, iter.cur);
+
+			if (worker->wi_proc == NULL ||
+				pg_atomic_unlocked_test_flag(&worker->wi_dobalance))
+				continue;
+
+			SetLatch(&worker->wi_proc->procLatch);
+		}
+	}
 }
 
 /*
diff --git a/src/include/commands/vacuum.h b/src/include/commands/vacuum.h
index 6e3c912bf5c..89fa1fcd261 100644
--- a/src/include/commands/vacuum.h
+++ b/src/include/commands/vacuum.h
@@ -432,6 +432,7 @@ extern void parallel_vacuum_cleanup_all_indexes(ParallelVacuumState *pvs,
 												PVWorkerStats *wstats);
 extern void parallel_vacuum_update_shared_delay_params(void);
 extern void parallel_vacuum_propagate_shared_delay_params(void);
+extern void parallel_vacuum_refresh_cost_params(void);
 extern void parallel_vacuum_main(dsm_segment *seg, shm_toc *toc);
 
 /* in commands/analyze.c */

base-commit: 4545cee303c257e58195e3d033c05bf38e2cd4d6
-- 
2.43.0



^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-24 09:04  Daniel Gustafsson <daniel@yesql.se>
  parent: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
  0 siblings, 2 replies; 43+ messages in thread

From: Daniel Gustafsson @ 2026-09-24 09:04 UTC (permalink / raw)
  To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: Masahiko Sawada <sawada.mshk@gmail.com>; Nikolay Samokhvalov <nik@postgres.ai>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

> On 24 Sep 2026, at 08:34, Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com> wrote:
> On Wed, Sep 23, 2026 at 5:39 PM Masahiko Sawada <sawada.mshk@gmail.com> wrote:

> > I've confirmed that the patch fixes the issue. While it works fine,
> > I'm a bit concerned that adding
> > WaitForParallelWorkersToFinishWithCallback() with a callback and a
> > timeout might be overkill, as I don't see any usecase other than
> > parallel autovacuum that needs to pass a callback.

I was also looking at this thread over the past few days and I agree that this
seems too invasive for the issue at hand given where we are in the cycle.

> I also don't like the wait-100ms-wakeup approach that a separate wait function would run all the time, since it wastes power and CPU cycles. Imagine a worker vacuuming an index that is hundreds of GBs or even TBs, while the leader has only a small index and finishes first. The leader then sits in the wait loop for longer until the large index is done, waking up every 100ms the whole time. So I would prefer not to go that route.

Agree, we should avoid polling loops like that as much as possible.

> How about doing the check inside the existing wait loop, only when the process is an autovacuum worker, along the lines of the attached WIP? This is simple, the pattern already exists elsewhere in the code, and it looks safe. I also believe the wait loop is not in a performance-critical hot path, and this check is not costly. I checked that it still fixes the reported issues.

It's not great to sprinkle in worker specific code in the generic parallel
handling.  My original thinking was to reject the idea, but the vacuum costing
is already used for non-vacuum purposes (and there have been discussions to
rename and generalize it) so with that in mind I am less concerned for this
particular case.

I can verify that the posted reproducer is still fixed with this patch applied
(and it survives a check-world).

> What I am less sure about is the fix for the second issue in the attached WIP patch, which needs any leader waiting for its workers to be woken up after the cost limit is rebalanced. I haven't found a better one yet.

Not sure I see a better solution either, and we are running short of time
before 19 RC1.

--
Daniel Gustafsson







^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-24 20:41  Zsolt Parragi <zsolt.parragi@percona.com>
  parent: Daniel Gustafsson <daniel@yesql.se>
  1 sibling, 1 reply; 43+ messages in thread

From: Zsolt Parragi @ 2026-09-24 20:41 UTC (permalink / raw)
  To: Daniel Gustafsson <daniel@yesql.se>; +Cc: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Masahiko Sawada <sawada.mshk@gmail.com>; Nikolay Samokhvalov <nik@postgres.ai>; pgsql-bugs@lists.postgresql.org

> I was also looking at this thread over the past few days and I agree that this
> seems too invasive for the issue at hand given where we are in the cycle.

I had the same issue, I tried to figure out something better but
couldn't. I still don't have a better solution. I have a somewhat
simpler version of v2, but the alternative WIP patch overall looks
better to me.

I attached an updated version of that with the tests integrated, and a
few code changes related to that (new injection points and one
simplification in parallel_vacuum_refresh_cost_params), so it's easier
to test.

> What I am less sure about is the fix for the second issue in the attached WIP patch, which needs any leader waiting for its workers to be woken up after the cost limit is rebalanced. I haven't found a better one yet.

Any alternative solution seems more complex to me, with no clear advantage.

Attachments:

  [application/octet-stream] nocfbot-v3-0001-Refresh-autovacuum-cost-parameters-while-waiting-.patch (12.3K, ../../CAN4CZFP-zBeCPWRHshj+10ofF0-NvPmG47qgP9fGhnS4AAf5LA@mail.gmail.com/2-nocfbot-v3-0001-Refresh-autovacuum-cost-parameters-while-waiting-.patch)
  download | inline diff:
From 85fdedf8d4f4183e7eb8b02b6d55a482730a4718 Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Thu, 24 Sep 2026 18:42:47 +0000
Subject: [PATCH v3] Refresh autovacuum cost parameters while waiting for
 parallel workers

An autovacuum leader publishes its cost-based delay parameters to its
parallel vacuum workers at its cost delay points. Once it has finished
its own share of the indexes and is only waiting for its workers, it
has no delay points left, so it stops publishing for as long as that
wait lasts, which can be as long as the largest index takes. Two kinds
of changes were lost for that whole time.

1. A config reload. The signal wakes the leader, but the wait does not
   act on it, so the reload stays pending until the wait ends.

2. A change in the number of autovacuum workers sharing the cost limit.
   Nothing signals that at all, and a waiting leader never re-reads the
   count because it no longer naps.

Fix the first by refreshing and publishing the parameters from the wait
itself, on every wakeup, when the process is an autovacuum worker. Fix
the second by waking the workers that share the limit whenever the count
changes, so a waiting leader picks it up the same way. A worker that is
busy vacuuming ignores the extra wakeup.

Reported-by: Nikolay Samokhvalov <nik@postgres.ai>
Discussion: https://postgr.es/m/CAM527d-GL%3DJp2EJXBSnVBGPK-4XEZwWof5Cv8P0hghS_og6oAg%40mail.gmail.com
---
 src/backend/access/transam/parallel.c         |   8 +
 src/backend/commands/vacuumparallel.c         |  44 ++++++
 src/backend/postmaster/autovacuum.c           |  19 +++
 src/include/commands/vacuum.h                 |   1 +
 .../t/001_parallel_autovacuum.pl              | 142 ++++++++++++++++++
 5 files changed, 214 insertions(+)

diff --git a/src/backend/access/transam/parallel.c b/src/backend/access/transam/parallel.c
index e1806a9a28a..89e45bdeb9c 100644
--- a/src/backend/access/transam/parallel.c
+++ b/src/backend/access/transam/parallel.c
@@ -812,6 +812,14 @@ WaitForParallelWorkersToFinish(ParallelContext *pcxt)
 		 */
 		CHECK_FOR_INTERRUPTS();
 
+		/*
+		 * An autovacuum leader publishes cost parameter changes to its
+		 * parallel workers at its cost delay points, which it no longer
+		 * reaches while waiting here. Do it here instead, on every wakeup.
+		 */
+		if (AmAutoVacuumWorkerProcess())
+			parallel_vacuum_refresh_cost_params();
+
 		for (i = 0; i < pcxt->nworkers_launched; ++i)
 		{
 			/*
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 767d162e578..07720609aea 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -43,9 +43,11 @@
 #include "executor/instrument.h"
 #include "optimizer/paths.h"
 #include "pgstat.h"
+#include "postmaster/interrupt.h"
 #include "storage/bufmgr.h"
 #include "storage/proc.h"
 #include "tcop/tcopprot.h"
+#include "utils/injection_point.h"
 #include "utils/lsyscache.h"
 #include "utils/rel.h"
 
@@ -725,6 +727,42 @@ parallel_vacuum_propagate_shared_delay_params(void)
 	pg_atomic_fetch_add_u32(&pv_shared_cost_params->generation, 1);
 }
 
+/*
+ * Refresh the leader's cost-based vacuum delay parameters and propagate them
+ * to its parallel vacuum workers.
+ *
+ * The leader normally does this at its cost delay points, which it no longer
+ * reaches once it is only waiting for its workers to finish. It calls this
+ * from that wait instead, on every wakeup, to pick up a config reload and a
+ * change in the number of autovacuum workers sharing the cost limit. Both of
+ * those wake the leader, the first by signal and the second when the count is
+ * recalculated.
+ */
+void
+parallel_vacuum_refresh_cost_params(void)
+{
+	Assert(AmAutoVacuumWorkerProcess());
+
+	/*
+	 * Quick return if the leader process is not sharing the delay parameters.
+	 */
+	if (pv_shared_cost_params == NULL)
+		return;
+
+	if (ConfigReloadPending)
+	{
+		ConfigReloadPending = false;
+		ProcessConfigFile(PGC_SIGHUP);
+	}
+
+	/*
+	 * Recompute from the (possibly reloaded) GUCs and the current number of
+	 * workers sharing the cost limit, then publish any change.
+	 */
+	VacuumUpdateCosts();
+	parallel_vacuum_propagate_shared_delay_params();
+}
+
 /*
  * Compute the number of parallel worker processes to request.  Both index
  * vacuum and index cleanup can be executed with parallel workers.
@@ -929,6 +967,9 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 	/* Vacuum the indexes that can be processed by only leader process */
 	parallel_vacuum_process_unsafe_indexes(pvs);
 
+	if (pvs->shared->is_autovacuum)
+		INJECTION_POINT("parallel-autovacuum-leader-before-index", NULL);
+
 	/*
 	 * Join as a parallel worker.  The leader vacuums alone processes all
 	 * parallel-safe indexes in the case where no workers are launched.
@@ -1010,6 +1051,9 @@ parallel_vacuum_process_safe_indexes(ParallelVacuumState *pvs)
 		if (!indstats->parallel_workers_can_process)
 			continue;
 
+		if (IsParallelWorker() && pvs->shared->is_autovacuum)
+			INJECTION_POINT("parallel-autovacuum-worker-before-index", NULL);
+
 		/* Do vacuum or cleanup of the index */
 		parallel_vacuum_process_one_index(pvs, pvs->indrels[idx], indstats);
 	}
diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c
index 60ebe828900..d45f361ac0d 100644
--- a/src/backend/postmaster/autovacuum.c
+++ b/src/backend/postmaster/autovacuum.c
@@ -1814,8 +1814,27 @@ autovac_recalculate_workers_for_balance(void)
 	}
 
 	if (nworkers_for_balance != orig_nworkers_for_balance)
+	{
 		pg_atomic_write_u32(&AutoVacuumShmem->av_nworkersForBalance,
 							nworkers_for_balance);
+
+		/*
+		 * Wake up the workers that share the limit. One that is vacuuming
+		 * picks up the new count on its next nap, but one that is only
+		 * waiting for its parallel vacuum workers never naps, and nothing
+		 * else wakes it while they keep running at the old share.
+		 */
+		dlist_foreach(iter, &AutoVacuumShmem->av_runningWorkers)
+		{
+			WorkerInfo	worker = dlist_container(WorkerInfoData, wi_links, iter.cur);
+
+			if (worker->wi_proc == NULL ||
+				pg_atomic_unlocked_test_flag(&worker->wi_dobalance))
+				continue;
+
+			SetLatch(&worker->wi_proc->procLatch);
+		}
+	}
 }
 
 /*
diff --git a/src/include/commands/vacuum.h b/src/include/commands/vacuum.h
index 6e3c912bf5c..89fa1fcd261 100644
--- a/src/include/commands/vacuum.h
+++ b/src/include/commands/vacuum.h
@@ -432,6 +432,7 @@ extern void parallel_vacuum_cleanup_all_indexes(ParallelVacuumState *pvs,
 												PVWorkerStats *wstats);
 extern void parallel_vacuum_update_shared_delay_params(void);
 extern void parallel_vacuum_propagate_shared_delay_params(void);
+extern void parallel_vacuum_refresh_cost_params(void);
 extern void parallel_vacuum_main(dsm_segment *seg, shm_toc *toc);
 
 /* in commands/analyze.c */
diff --git a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
index 33c86bbdc94..10232795e02 100644
--- a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
+++ b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
@@ -247,6 +247,148 @@ my @limits =
 note("parallel worker cost_limit sequence: @limits");
 is($limits[0], '250', 'parallel workers see the rebalanced cost limit');
 
+# Release the second worker.
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('autovacuum-worker-cost-balanced')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('autovacuum-worker-cost-balanced')");
+
+# Wait for the second worker from test 3 to finish, and go back to a single
+# autovacuum worker so that the balance below is known.
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) = 0 FROM pg_stat_activity
+	WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
+}) or die "second autovacuum worker did not finish";
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_max_workers = 1;
+	SELECT pg_reload_conf();
+});
+
+my $postgresoid = $node->safe_psql('postgres',
+	"SELECT oid FROM pg_database WHERE datname = 'postgres'");
+my $testautovacoid =
+  $node->safe_psql('postgres', "SELECT 'test_autovac'::regclass::oid");
+
+# Start a parallel autovacuum on test_autovac and leave it with the leader
+# idle in WaitForParallelWorkersToFinish() while its parallel worker is held
+# on the first index it claimed.  The leader is held before it starts on the
+# indexes, so that the worker gets to claim one before the leader takes them
+# all.
+sub start_leader_waiting_for_worker
+{
+	my ($node) = @_;
+
+	$node->safe_psql(
+		'postgres', q{
+		SELECT injection_points_attach('parallel-autovacuum-worker-before-index', 'wait');
+		SELECT injection_points_attach('parallel-autovacuum-leader-before-index', 'wait');
+		ALTER TABLE test_autovac SET (autovacuum_enabled = true);
+	});
+	$node->wait_for_event('autovacuum worker',
+		'parallel-autovacuum-leader-before-index');
+	$node->wait_for_event('parallel worker',
+		'parallel-autovacuum-worker-before-index');
+	$node->safe_psql(
+		'postgres', q{
+		SELECT injection_points_wakeup('parallel-autovacuum-leader-before-index');
+		SELECT injection_points_detach('parallel-autovacuum-leader-before-index');
+	});
+	$node->poll_query_until(
+		'postgres', q{
+		SELECT count(*) > 0 FROM pg_stat_activity
+		WHERE backend_type = 'autovacuum worker'
+		  AND wait_event = 'ParallelFinish'
+	}) or die "autovacuum leader did not enter ParallelFinish";
+}
+
+sub release_worker
+{
+	my ($node) = @_;
+
+	$node->safe_psql(
+		'postgres', q{
+		SELECT injection_points_wakeup('parallel-autovacuum-worker-before-index');
+		SELECT injection_points_detach('parallel-autovacuum-worker-before-index');
+	});
+}
+
+# Test 4:
+# A config reload arriving while the leader only waits for its parallel
+# worker must still reach the worker before it finishes.
+
+prepare_for_next_test($node, 4);
+$log_offset = -s $node->logfile;
+
+start_leader_waiting_for_worker($node);
+
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 800;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_delay = 8;
+	ALTER SYSTEM SET vacuum_cost_page_miss = 11;
+	ALTER SYSTEM SET vacuum_cost_page_dirty = 12;
+	ALTER SYSTEM SET vacuum_cost_page_hit = 13;
+	SELECT pg_reload_conf();
+});
+
+# The waiting leader must pick up the reload on its own.
+$node->wait_for_log(
+	qr/Autovacuum VacuumUpdateCosts\(db=$postgresoid, rel=$testautovacoid, dobalance=yes, cost_limit=800, cost_delay=8 /,
+	$log_offset);
+
+release_worker($node);
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=800, cost_delay=8, cost_page_miss=11, cost_page_dirty=12, cost_page_hit=13/,
+	$log_offset);
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+ok(1, "config reload reaches parallel workers while the leader waits");
+
+# Test 5:
+# Same for a cost limit rebalance, which unlike a reload sends no signal:
+# a second autovacuum worker joins the balance while the leader waits.
+
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_max_workers = 2;
+	SELECT pg_reload_conf();
+});
+
+prepare_for_next_test($node, 5);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = false)');
+$node->safe_psql('regress_db2', 'UPDATE filler SET id = id + 1');
+$log_offset = -s $node->logfile;
+
+start_leader_waiting_for_worker($node);
+
+# Second worker -> balance = 2, held there as in test 3.
+$node->safe_psql('postgres',
+	"SELECT injection_points_attach('autovacuum-worker-cost-balanced', 'wait')"
+);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = true)');
+$node->wait_for_log(
+	qr/VacuumUpdateCosts\(db=$db2oid, rel=$filleroid, dobalance=yes, cost_limit=400,/,
+	$log_offset);
+
+# The waiting leader must notice the rebalance on its own.
+$node->wait_for_log(
+	qr/Autovacuum VacuumUpdateCosts\(db=$postgresoid, rel=$testautovacoid, dobalance=yes, cost_limit=400,/,
+	$log_offset);
+
+release_worker($node);
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=400,/,
+	$log_offset);
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+ok(1, "cost limit rebalance reaches parallel workers while the leader waits");
+
 # Release the second worker.
 $node->safe_psql('postgres',
 	"SELECT injection_points_wakeup('autovacuum-worker-cost-balanced')");
-- 
2.55.0



^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-24 21:35  Manu <manuelreyesbravo@gmail.com>
  parent: Zsolt Parragi <zsolt.parragi@percona.com>
  0 siblings, 0 replies; 43+ messages in thread

From: Manu @ 2026-09-24 21:35 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Masahiko Sawada <sawada.mshk@gmail.com>; Nikolay Samokhvalov <nik@postgres.ai>; pgsql-bugs@lists.postgresql.org

Hi,

I ran v3 on REL_19_STABLE (e60ee52841d) with --enable-cassert and
--enable-injection-points, next to Nikolay's v2 and two cut-down copies
of v3, to check the new tests and the polling question from upthread.

1. Each part of the fix has a test that fails without it

  v3 tests and injection points, no fix:   tests 1-3 pass, 4 fails
  v3 without the SetLatch() in
  autovac_recalculate_workers_for_balance(): tests 1-4 pass, 5 fails
  v3:                                        all pass

2. Stability

test_autovacuum passed 30 of 30 runs on an idle machine, and 20 of 20
with every CPU busy (one busy loop per core), at about 10 s and 16 s a
run.

3. A case the tests don't cover: a worker leaving the balance

Test 5 checks a worker joining while the leader waits.  A worker
leaving goes through another path: FreeWorkerInfo() sets
AutoVacRebalance, and the launcher recalculates the count and so does
the SetLatch().  I added it as test 6 (attached, on top of v3): the
second worker joins and is held as in test 5, the leader goes to 400,
then the second worker is released and exits.  With v3 the leader,
still in ParallelFinish, logs cost_limit=800 by the first check after
the worker is gone, and the parallel worker ends at 800.

Without the SetLatch() the test stops earlier, at the join, like test
5, so it does not tell the two paths apart.  It does show that the
launcher path works, which the current tests don't exercise.

4. Polling vs. waking, measured

With the parallel worker held before its index and the leader waiting
for it for 20 s, the leader's voluntary context switches were:

  v3:            0
  Nikolay's v2:  199 (9.9 per second)

The time from pg_reload_conf() to the leader logging the new limit was
7 ms with v3 and 8 ms with v2.  So the timed wait costs about ten
wakeups a second for as long as the largest index takes, and makes no
difference to how fast a reload is picked up.

5. Things I checked in the code, in case they come up

- The extra SetLatch() does not cut short a busy worker's cost delay:
  vacuum_delay_point() sleeps with pg_usleep(), not on the latch, so
  "a worker that is busy vacuuming ignores the extra wakeup" holds.
- No wakeup is lost in the wait.  The refresh runs at the top of each
  pass, before WaitLatch().  A SetLatch() after it makes WaitLatch()
  return at once, and one between WaitLatch() and ResetLatch() is
  followed by another pass that reads the current count anyway, since
  the count is written before the SetLatch().
- The leader also wakes for its workers' messages, but
  parallel_vacuum_propagate_shared_delay_params() bumps the generation
  only when a value changed, so those wakeups make no worker re-read
  its parameters.

The scripts and all outputs are in the second attachment.

Regards,
Manu
diff --git a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
--- a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
+++ b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
@@ -8,6 +8,7 @@
 use PostgreSQL::Test::Cluster;
 use PostgreSQL::Test::Utils;
 use Test::More;
+use Time::HiRes qw(usleep);
 
 if ($ENV{enable_injection_points} ne 'yes')
 {
@@ -395,5 +396,78 @@
 $node->safe_psql('postgres',
 	"SELECT injection_points_detach('autovacuum-worker-cost-balanced')");
 
+
+# Test 6 (added for review):
+# The reverse of test 5: a worker LEAVES the balance while the leader waits.
+# That rebalance takes another path: the leaving worker sets AutoVacRebalance
+# in FreeWorkerInfo(), and the launcher recalculates the count.
+
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) = 0 FROM pg_stat_activity
+	WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
+}) or die "second autovacuum worker from test 5 did not finish";
+
+prepare_for_next_test($node, 6);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = false)');
+$node->safe_psql('regress_db2', 'UPDATE filler SET id = id + 1 WHERE true');
+$log_offset = -s $node->logfile;
+
+start_leader_waiting_for_worker($node);
+
+# A second worker joins and is held, as in test 5: the leader goes to 400.
+$node->safe_psql('postgres',
+	"SELECT injection_points_attach('autovacuum-worker-cost-balanced', 'wait')"
+);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = true)');
+$node->wait_for_log(
+	qr/Autovacuum VacuumUpdateCosts\(db=$postgresoid, rel=$testautovacoid, dobalance=yes, cost_limit=400,/,
+	$log_offset);
+
+# Now let the second worker finish and leave, with the leader still waiting.
+my $leave_offset = -s $node->logfile;
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('autovacuum-worker-cost-balanced')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('autovacuum-worker-cost-balanced')");
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) = 0 FROM pg_stat_activity
+	WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
+}) or die "second autovacuum worker did not finish";
+my $left_at = time;
+
+# The waiting leader should go back to the whole limit.  Poll the log for up
+# to 30 s instead of the default timeout.
+my $seen = 0;
+for (my $i = 0; $i < 300; $i++)
+{
+	if (slurp_file($node->logfile, $leave_offset) =~
+		/Autovacuum VacuumUpdateCosts\(db=$postgresoid, rel=$testautovacoid, dobalance=yes, cost_limit=800,/)
+	{
+		$seen = 1;
+		last;
+	}
+	usleep(100_000);
+}
+my $still_waiting = $node->safe_psql(
+	'postgres', q{
+	SELECT count(*) FROM pg_stat_activity
+	WHERE backend_type = 'autovacuum worker' AND wait_event = 'ParallelFinish'});
+note("leader still in ParallelFinish when checked: $still_waiting; waited "
+	  . (time - $left_at) . " s");
+ok($seen, 'waiting leader returns to the whole limit when a worker leaves');
+
+release_worker($node);
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$leave_offset);
+my @limits6 =
+  slurp_file($node->logfile, $leave_offset) =~
+  /parallel autovacuum worker updated cost params: cost_limit=(\d+),/g;
+note("parallel worker cost_limit sequence after the worker left: @limits6");
+is($limits6[-1] // 'none', '800', 'parallel worker ends at the whole limit');
 $node->stop;
 done_testing();

# Review of Zsolt v3 (Bharath approach + tests) for the parallel autovacuum
# cost-parameter wait.  All builds from REL_19_STABLE e60ee52841d,
# --enable-cassert --enable-injection-points --enable-tap-tests.
#   zsolt   + v3
#   nikv2   + Nikolay v2 (timed wait)
#   nofix   v3 tests and injection points, without the parallel.c and autovacuum.c changes
#   nowake  v3 without the autovacuum.c change (the SetLatch)


======== stability.sh ========

#!/bin/bash
# Run the whole test_autovacuum TAP suite N times on one build and count
# failures; keep the log of every failing run.
#   stability.sh <build> <runs> [label]
# LOAD=1 runs it while every CPU is busy (one busy loop per core), the
# kind of machine where timing-dependent tests break.
set -u
B=$1; N=$2; L=${3:-idle}
A=$(cd "$(dirname "$0")" && pwd)
OUT=$A/stability.$B.$L.txt
: > $OUT
if [ "${LOAD:-0}" = 1 ]; then
  for _ in $(seq "$(nproc)"); do ( while :; do :; done ) & done
  trap 'kill $(jobs -p) 2>/dev/null' EXIT
fi
fail=0
for i in $(seq 1 $N); do
  t0=$(date +%s.%N)
  if make -C $HOME/pgav/b-$B/src/test/modules/test_autovacuum check > /tmp/claude-1000/stab.$B.log 2>&1; then r=ok; else r=FAIL; fail=$((fail+1));
    mkdir -p $A/stability-fail; cp -r $HOME/pgav/b-$B/src/test/modules/test_autovacuum/tmp_check/log $A/stability-fail/$B.$L.run$i 2>/dev/null
    cp /tmp/claude-1000/stab.$B.log $A/stability-fail/$B.$L.run$i.make.log; fi
  printf '%s run %2d %s %.1fs\n' "$B/$L" $i $r "$(echo "$(date +%s.%N) - $t0" | bc)" | tee -a $OUT
done
echo "$B/$L: $fail failures in $N runs" | tee -a $OUT


======== wakeups.sh ========

#!/bin/bash
# How often the autovacuum leader wakes up while it only waits for a parallel
# worker, and how fast it picks up a config reload, with each proposed fix.
#   zsolt  v3 (wait in WaitForParallelWorkersToFinish(), woken by the reload
#          signal or by SetLatch() on a rebalance)
#   nikv2  v2 (timed wait in vacuumparallel.c, 100 ms)
# The parallel worker is held at the build's own injection point before its
# index; the leader goes through its own indexes and then waits.  Wakeups are
# the leader's voluntary context switches over HOLD seconds (/proc).
#   wakeups.sh [HOLD]
set -u
HOLD=${1:-20}
A=$(cd "$(dirname "$0")" && pwd)
for B in zsolt nikv2; do
  I=$HOME/pgav/i-$B/bin
  case $B in
    zsolt) WPT=parallel-autovacuum-worker-before-index; LPT=parallel-autovacuum-leader-before-index ;;
    nikv2) WPT=parallel-vacuum-worker-before-index;     LPT=parallel-vacuum-leader-before-index ;;
  esac
  D=$(mktemp -d /tmp/claude-1000/wk.XXXX); P=55440
  "$I/initdb" -D $D -A trust --no-sync -U postgres >/dev/null
  cat >> $D/postgresql.conf <<EOF
port = $P
unix_socket_directories = '/tmp'
autovacuum_max_workers = 1
autovacuum_worker_slots = 2
autovacuum_max_parallel_workers = 2
max_worker_processes = 10
max_parallel_workers = 10
log_min_messages = debug2
log_line_prefix = '%m [%p] '
autovacuum_naptime = '1s'
min_parallel_index_scan_size = 0
autovacuum_vacuum_threshold = 100000
autovacuum_analyze_threshold = 100000
autovacuum_vacuum_insert_threshold = -1
EOF
  "$I/pg_ctl" -D $D -l $D/log -w start >/dev/null
  q() { "$I/psql" -X -qAt -h /tmp -p $P -U postgres -c "$1"; }
  q "CREATE EXTENSION injection_points"
  q "CREATE TABLE test_autovac (id serial primary key, c1 int, c2 int, c3 int)
       WITH (autovacuum_parallel_workers = 1, autovacuum_vacuum_threshold = 50, autovacuum_enabled = false)"
  q "INSERT INTO test_autovac (c1, c2, c3) SELECT g, g, g FROM generate_series(1, 10000) g"
  q "CREATE INDEX i1 ON test_autovac (c1)"; q "CREATE INDEX i2 ON test_autovac (c2)"; q "CREATE INDEX i3 ON test_autovac (c3)"
  q "UPDATE test_autovac SET c1 = c1 + 1"
  q "SELECT injection_points_attach('$WPT', 'wait')"
  q "SELECT injection_points_attach('$LPT', 'wait')"
  q "ALTER TABLE test_autovac SET (autovacuum_enabled = true)"
  for _ in $(seq 60); do [ "$(q "SELECT count(*) FROM pg_stat_activity WHERE wait_event = '$LPT'")" = 1 ] && break; sleep 0.5; done
  for _ in $(seq 60); do [ "$(q "SELECT count(*) FROM pg_stat_activity WHERE wait_event = '$WPT'")" = 1 ] && break; sleep 0.5; done
  q "SELECT injection_points_wakeup('$LPT')"; q "SELECT injection_points_detach('$LPT')"
  sleep 2
  LEADER=$(q "SELECT pid FROM pg_stat_activity WHERE backend_type = 'autovacuum worker'")
  WEV=$(q "SELECT wait_event_type || '/' || wait_event FROM pg_stat_activity WHERE pid = $LEADER")
  v0=$(awk '/^voluntary_ctxt_switches/{print $2}' /proc/$LEADER/status)
  sleep $HOLD
  v1=$(awk '/^voluntary_ctxt_switches/{print $2}' /proc/$LEADER/status)
  # reload latency: from pg_reload_conf() to the leader's VacuumUpdateCosts with the new limit
  off=$(stat -c %s $D/log)
  q "ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 777"
  t0=$(date +%s.%N); q "SELECT pg_reload_conf()"
  for _ in $(seq 300); do tail -c +$((off+1)) $D/log | grep -q "pid\|cost_limit=777" && tail -c +$((off+1)) $D/log | grep -q "VacuumUpdateCosts.*cost_limit=777" && break; sleep 0.01; done
  t1=$(date +%s.%N)
  lat=$(echo "($t1 - $t0) * 1000" | bc)
  printf '%-6s leader %s waiting as %-28s voluntary wakeups in %ss: %6d (%.1f/s)   reload seen after %5.0f ms\n' \
    $B $LEADER "$WEV" $HOLD $((v1 - v0)) "$(echo "($v1 - $v0) / $HOLD" | bc -l)" "$lat"
  q "SELECT injection_points_wakeup('$WPT')"; q "SELECT injection_points_detach('$WPT')"
  "$I/pg_ctl" -D $D -m fast -w stop >/dev/null
  cp $D/log $A/wakeups.$B.server.log
  rm -rf $D
done


======== build_partial.sh ========

#!/bin/bash
# Zsolt's v3 with parts of the fix taken out, to check that each test fails
# without the part it is meant to test (the tests and injection points stay):
#   nofix    without the refresh in WaitForParallelWorkersToFinish() and
#            without the SetLatch() in autovac_recalculate_workers_for_balance()
#   nowake   only without the SetLatch()
set -eu
BASE=${BASE:-e60ee52841d}
SRC=$HOME/Proyectos/postgresql
W=$HOME/pgav
A=$(cd "$(dirname "$0")" && pwd)
P=$(ls $A/zsolt-latest/nocfbot-v3-0001-*.patch)
mk() {  # name  files-to-exclude...
  local name=$1; shift
  local tree=$W/src-$name
  if [ ! -d $tree ]; then
    git -C $SRC worktree add -q --detach $tree $BASE
    local ex=(); for f in "$@"; do ex+=(--exclude="$f"); done
    git -C $tree apply --whitespace=nowarn "${ex[@]}" "$P"
    # parallel_vacuum_refresh_cost_params() stays defined but unused when
    # parallel.c is excluded; that is fine for a test-only build.
  fi
  mkdir -p $W/b-$name && cd $W/b-$name
  $tree/configure --prefix=$W/i-$name --enable-cassert --enable-injection-points \
    --enable-tap-tests --quiet > configure.log 2>&1
  make -j"$(nproc)" -s > build.log 2>&1
  make -s install > install.log 2>&1
  make -C src/test/modules/injection_points -s install >> install.log 2>&1
  echo "$name: warnings=$(grep -c 'warning:' build.log) | $(git -C $tree diff --stat | tr '\n' ' ' | sed 's/  */ /g')"
}
mk nofix src/backend/access/transam/parallel.c src/backend/postmaster/autovacuum.c
mk nowake src/backend/postmaster/autovacuum.c
echo PARTIAL-DONE


======== run_test6.sh ========

#!/bin/bash
# Run 003_rebalance_worker_leaves.pl (Zsolt's tests + test 6) on one build.
# The file is copied into the build's source tree for this run only and is
# removed afterwards, whatever the outcome, so it never leaks into other runs.
#   run_test6.sh <build>
set -u
B=$1
A=$(cd "$(dirname "$0")" && pwd)
T=$HOME/pgav/src-$B/src/test/modules/test_autovacuum/t
cp $A/003_rebalance_worker_leaves.pl $T/
trap 'rm -f $T/003_rebalance_worker_leaves.pl' EXIT
export PG_TEST_TIMEOUT_DEFAULT=60
make -C $HOME/pgav/b-$B/src/test/modules/test_autovacuum check \
  PROVE_TESTS=t/003_rebalance_worker_leaves.pl > $A/test6.$B.log 2>&1
echo "== $B exit=$?"
L=$HOME/pgav/b-$B/src/test/modules/test_autovacuum/tmp_check/log
grep -hE '(ok|not ok) [0-9]+ - |# leader still|# parallel worker cost_limit sequence|die:|Looks like' \
  $L/regress_log_003_rebalance_worker_leaves | sed 's/^\[[^]]*\]//' | cut -c1-170
mkdir -p $A/test6-logs/$B && cp $L/* $A/test6-logs/$B/


======== results: stability, idle ========
zsolt/idle run  1 ok 9.9s
zsolt/idle run  2 ok 9.4s
zsolt/idle run  3 ok 9.5s
zsolt/idle run  4 ok 9.9s
zsolt/idle run  5 ok 9.5s
zsolt/idle run  6 ok 9.5s
zsolt/idle run  7 FAIL 9.6s
zsolt/idle run  8 ok 9.7s
zsolt/idle run  9 ok 9.4s
zsolt/idle run 10 ok 9.7s
zsolt/idle run 11 ok 9.7s
zsolt/idle run 12 ok 10.7s
zsolt/idle run 13 ok 9.8s
zsolt/idle run 14 ok 9.8s
zsolt/idle run 15 ok 9.4s
zsolt/idle run 16 ok 9.6s
zsolt/idle run 17 ok 9.6s
zsolt/idle run 18 ok 9.6s
zsolt/idle run 19 ok 9.5s
zsolt/idle run 20 ok 9.3s
zsolt/idle run 21 ok 9.7s
zsolt/idle run 22 ok 10.0s
zsolt/idle run 23 ok 9.8s
zsolt/idle run 24 ok 9.7s
zsolt/idle run 25 ok 9.2s
zsolt/idle run 26 ok 9.5s
zsolt/idle run 27 ok 9.6s
zsolt/idle run 28 ok 9.6s
zsolt/idle run 29 ok 9.4s
zsolt/idle run 30 ok 9.4s
zsolt/idle: 1 failures in 30 runs
NOTE: run 7 above is not a test failure: prove globbed t/003_rebalance_worker_leaves.pl, a file I placed in the source tree by mistake and removed during the run ('Cannot detect source'). t/001_parallel_autovacuum.pl itself passed in run 7 (see stability-fail/zsolt.idle.run7.make.log). Valid result: 001 passed 30 of 30.

======== results: stability, every CPU busy ========
zsolt/load run  1 ok 16.6s
zsolt/load run  2 ok 16.7s
zsolt/load run  3 ok 16.1s
zsolt/load run  4 ok 16.4s
zsolt/load run  5 ok 15.7s
zsolt/load run  6 ok 16.9s
zsolt/load run  7 ok 16.8s
zsolt/load run  8 ok 16.5s
zsolt/load run  9 ok 16.3s
zsolt/load run 10 ok 18.5s
zsolt/load run 11 ok 16.9s
zsolt/load run 12 ok 18.3s
zsolt/load run 13 ok 19.0s
zsolt/load run 14 ok 16.5s
zsolt/load run 15 ok 16.8s
zsolt/load run 16 ok 16.9s
zsolt/load run 17 ok 16.8s
zsolt/load run 18 ok 16.3s
zsolt/load run 19 ok 17.5s
zsolt/load run 20 ok 17.2s
zsolt/load: 0 failures in 20 runs

======== results: v3 tests on the cut-down builds ========
-- nofix
(1.395s) ok 1 - parallel autovacuum on test_autovac table
(1.159s) ok 2 - vacuum delay parameter changes are propagated to parallel vacuum workers
(0.000s) ok 3 - parallel workers see the rebalanced cost limit
(30.943s) # die: timed out waiting for file pgav/b-nofix/src/test/modules/test_autovacuum/tmp_check/log/001_parallel_autovacuum_main.log co
(0.104s) # Looks like your test exited with 255 just after 3.
-- nowake
(1.321s) ok 1 - parallel autovacuum on test_autovac table
(1.144s) ok 2 - vacuum delay parameter changes are propagated to parallel vacuum workers
(0.000s) ok 3 - parallel workers see the rebalanced cost limit
(0.872s) ok 4 - config reload reaches parallel workers while the leader waits
(31.539s) # die: timed out waiting for file pgav/b-nowake/src/test/modules/test_autovacuum/tmp_check/log/001_parallel_autovacuum_main.log c
(0.103s) # Looks like your test exited with 255 just after 4.

======== results: v3 tests + test 6 ========
== zsolt exit=0
(1.444s) ok 1 - parallel autovacuum on test_autovac table
(1.147s) ok 2 - vacuum delay parameter changes are propagated to parallel vacuum workers
(1.759s) # parallel worker cost_limit sequence: 250
(0.000s) ok 3 - parallel workers see the rebalanced cost limit
(1.194s) ok 4 - config reload reaches parallel workers while the leader waits
(1.684s) ok 5 - cost limit rebalance reaches parallel workers while the leader waits
(0.983s) # leader still in ParallelFinish when checked: 1; waited 0 s
(0.000s) ok 6 - waiting leader returns to the whole limit when a worker leaves
(0.105s) # parallel worker cost_limit sequence after the worker left: 800
(0.000s) ok 7 - parallel worker ends at the whole limit
== nowake exit=2
(1.332s) ok 1 - parallel autovacuum on test_autovac table
(1.256s) ok 2 - vacuum delay parameter changes are propagated to parallel vacuum workers
(1.657s) # parallel worker cost_limit sequence: 250
(0.000s) ok 3 - parallel workers see the rebalanced cost limit
(1.091s) ok 4 - config reload reaches parallel workers while the leader waits
(61.605s) # die: timed out waiting for file pgav/b-nowake/src/test/modules/test_autovacuum/tmp_check/log/003_rebalance_worker_leaves_main.log contents to match
(0.104s) # Looks like your test exited with 255 just after 4.

======== results: leader wakeups while waiting 20 s ========
zsolt  leader 272238 waiting as IPC/ParallelFinish           voluntary wakeups in 20s:      0 (0.0/s)   reload seen after     7 ms
nikv2  leader 278985 waiting as IPC/ParallelFinish           voluntary wakeups in 20s:    199 (9.9/s)   reload seen after     8 ms


Attachments:

  [text/plain] nocfbot-v3-add-test-worker-leaves.diff.txt (3.4K, ../../179028573649.493682.4321152895319402838@gmail.com/2-nocfbot-v3-add-test-worker-leaves.diff.txt)
  download | inline diff:
diff --git a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
--- a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
+++ b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
@@ -8,6 +8,7 @@
 use PostgreSQL::Test::Cluster;
 use PostgreSQL::Test::Utils;
 use Test::More;
+use Time::HiRes qw(usleep);
 
 if ($ENV{enable_injection_points} ne 'yes')
 {
@@ -395,5 +396,78 @@
 $node->safe_psql('postgres',
 	"SELECT injection_points_detach('autovacuum-worker-cost-balanced')");
 
+
+# Test 6 (added for review):
+# The reverse of test 5: a worker LEAVES the balance while the leader waits.
+# That rebalance takes another path: the leaving worker sets AutoVacRebalance
+# in FreeWorkerInfo(), and the launcher recalculates the count.
+
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) = 0 FROM pg_stat_activity
+	WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
+}) or die "second autovacuum worker from test 5 did not finish";
+
+prepare_for_next_test($node, 6);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = false)');
+$node->safe_psql('regress_db2', 'UPDATE filler SET id = id + 1 WHERE true');
+$log_offset = -s $node->logfile;
+
+start_leader_waiting_for_worker($node);
+
+# A second worker joins and is held, as in test 5: the leader goes to 400.
+$node->safe_psql('postgres',
+	"SELECT injection_points_attach('autovacuum-worker-cost-balanced', 'wait')"
+);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = true)');
+$node->wait_for_log(
+	qr/Autovacuum VacuumUpdateCosts\(db=$postgresoid, rel=$testautovacoid, dobalance=yes, cost_limit=400,/,
+	$log_offset);
+
+# Now let the second worker finish and leave, with the leader still waiting.
+my $leave_offset = -s $node->logfile;
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('autovacuum-worker-cost-balanced')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('autovacuum-worker-cost-balanced')");
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) = 0 FROM pg_stat_activity
+	WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
+}) or die "second autovacuum worker did not finish";
+my $left_at = time;
+
+# The waiting leader should go back to the whole limit.  Poll the log for up
+# to 30 s instead of the default timeout.
+my $seen = 0;
+for (my $i = 0; $i < 300; $i++)
+{
+	if (slurp_file($node->logfile, $leave_offset) =~
+		/Autovacuum VacuumUpdateCosts\(db=$postgresoid, rel=$testautovacoid, dobalance=yes, cost_limit=800,/)
+	{
+		$seen = 1;
+		last;
+	}
+	usleep(100_000);
+}
+my $still_waiting = $node->safe_psql(
+	'postgres', q{
+	SELECT count(*) FROM pg_stat_activity
+	WHERE backend_type = 'autovacuum worker' AND wait_event = 'ParallelFinish'});
+note("leader still in ParallelFinish when checked: $still_waiting; waited "
+	  . (time - $left_at) . " s");
+ok($seen, 'waiting leader returns to the whole limit when a worker leaves');
+
+release_worker($node);
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$leave_offset);
+my @limits6 =
+  slurp_file($node->logfile, $leave_offset) =~
+  /parallel autovacuum worker updated cost params: cost_limit=(\d+),/g;
+note("parallel worker cost_limit sequence after the worker left: @limits6");
+is($limits6[-1] // 'none', '800', 'parallel worker ends at the whole limit');
 $node->stop;
 done_testing();


  [text/plain] nocfbot-av-cost-wait-review.txt (12.2K, ../../179028573649.493682.4321152895319402838@gmail.com/3-nocfbot-av-cost-wait-review.txt)
  download | inline:
# Review of Zsolt v3 (Bharath approach + tests) for the parallel autovacuum
# cost-parameter wait.  All builds from REL_19_STABLE e60ee52841d,
# --enable-cassert --enable-injection-points --enable-tap-tests.
#   zsolt   + v3
#   nikv2   + Nikolay v2 (timed wait)
#   nofix   v3 tests and injection points, without the parallel.c and autovacuum.c changes
#   nowake  v3 without the autovacuum.c change (the SetLatch)


======== stability.sh ========

#!/bin/bash
# Run the whole test_autovacuum TAP suite N times on one build and count
# failures; keep the log of every failing run.
#   stability.sh <build> <runs> [label]
# LOAD=1 runs it while every CPU is busy (one busy loop per core), the
# kind of machine where timing-dependent tests break.
set -u
B=$1; N=$2; L=${3:-idle}
A=$(cd "$(dirname "$0")" && pwd)
OUT=$A/stability.$B.$L.txt
: > $OUT
if [ "${LOAD:-0}" = 1 ]; then
  for _ in $(seq "$(nproc)"); do ( while :; do :; done ) & done
  trap 'kill $(jobs -p) 2>/dev/null' EXIT
fi
fail=0
for i in $(seq 1 $N); do
  t0=$(date +%s.%N)
  if make -C $HOME/pgav/b-$B/src/test/modules/test_autovacuum check > /tmp/claude-1000/stab.$B.log 2>&1; then r=ok; else r=FAIL; fail=$((fail+1));
    mkdir -p $A/stability-fail; cp -r $HOME/pgav/b-$B/src/test/modules/test_autovacuum/tmp_check/log $A/stability-fail/$B.$L.run$i 2>/dev/null
    cp /tmp/claude-1000/stab.$B.log $A/stability-fail/$B.$L.run$i.make.log; fi
  printf '%s run %2d %s %.1fs\n' "$B/$L" $i $r "$(echo "$(date +%s.%N) - $t0" | bc)" | tee -a $OUT
done
echo "$B/$L: $fail failures in $N runs" | tee -a $OUT


======== wakeups.sh ========

#!/bin/bash
# How often the autovacuum leader wakes up while it only waits for a parallel
# worker, and how fast it picks up a config reload, with each proposed fix.
#   zsolt  v3 (wait in WaitForParallelWorkersToFinish(), woken by the reload
#          signal or by SetLatch() on a rebalance)
#   nikv2  v2 (timed wait in vacuumparallel.c, 100 ms)
# The parallel worker is held at the build's own injection point before its
# index; the leader goes through its own indexes and then waits.  Wakeups are
# the leader's voluntary context switches over HOLD seconds (/proc).
#   wakeups.sh [HOLD]
set -u
HOLD=${1:-20}
A=$(cd "$(dirname "$0")" && pwd)
for B in zsolt nikv2; do
  I=$HOME/pgav/i-$B/bin
  case $B in
    zsolt) WPT=parallel-autovacuum-worker-before-index; LPT=parallel-autovacuum-leader-before-index ;;
    nikv2) WPT=parallel-vacuum-worker-before-index;     LPT=parallel-vacuum-leader-before-index ;;
  esac
  D=$(mktemp -d /tmp/claude-1000/wk.XXXX); P=55440
  "$I/initdb" -D $D -A trust --no-sync -U postgres >/dev/null
  cat >> $D/postgresql.conf <<EOF
port = $P
unix_socket_directories = '/tmp'
autovacuum_max_workers = 1
autovacuum_worker_slots = 2
autovacuum_max_parallel_workers = 2
max_worker_processes = 10
max_parallel_workers = 10
log_min_messages = debug2
log_line_prefix = '%m [%p] '
autovacuum_naptime = '1s'
min_parallel_index_scan_size = 0
autovacuum_vacuum_threshold = 100000
autovacuum_analyze_threshold = 100000
autovacuum_vacuum_insert_threshold = -1
EOF
  "$I/pg_ctl" -D $D -l $D/log -w start >/dev/null
  q() { "$I/psql" -X -qAt -h /tmp -p $P -U postgres -c "$1"; }
  q "CREATE EXTENSION injection_points"
  q "CREATE TABLE test_autovac (id serial primary key, c1 int, c2 int, c3 int)
       WITH (autovacuum_parallel_workers = 1, autovacuum_vacuum_threshold = 50, autovacuum_enabled = false)"
  q "INSERT INTO test_autovac (c1, c2, c3) SELECT g, g, g FROM generate_series(1, 10000) g"
  q "CREATE INDEX i1 ON test_autovac (c1)"; q "CREATE INDEX i2 ON test_autovac (c2)"; q "CREATE INDEX i3 ON test_autovac (c3)"
  q "UPDATE test_autovac SET c1 = c1 + 1"
  q "SELECT injection_points_attach('$WPT', 'wait')"
  q "SELECT injection_points_attach('$LPT', 'wait')"
  q "ALTER TABLE test_autovac SET (autovacuum_enabled = true)"
  for _ in $(seq 60); do [ "$(q "SELECT count(*) FROM pg_stat_activity WHERE wait_event = '$LPT'")" = 1 ] && break; sleep 0.5; done
  for _ in $(seq 60); do [ "$(q "SELECT count(*) FROM pg_stat_activity WHERE wait_event = '$WPT'")" = 1 ] && break; sleep 0.5; done
  q "SELECT injection_points_wakeup('$LPT')"; q "SELECT injection_points_detach('$LPT')"
  sleep 2
  LEADER=$(q "SELECT pid FROM pg_stat_activity WHERE backend_type = 'autovacuum worker'")
  WEV=$(q "SELECT wait_event_type || '/' || wait_event FROM pg_stat_activity WHERE pid = $LEADER")
  v0=$(awk '/^voluntary_ctxt_switches/{print $2}' /proc/$LEADER/status)
  sleep $HOLD
  v1=$(awk '/^voluntary_ctxt_switches/{print $2}' /proc/$LEADER/status)
  # reload latency: from pg_reload_conf() to the leader's VacuumUpdateCosts with the new limit
  off=$(stat -c %s $D/log)
  q "ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 777"
  t0=$(date +%s.%N); q "SELECT pg_reload_conf()"
  for _ in $(seq 300); do tail -c +$((off+1)) $D/log | grep -q "pid\|cost_limit=777" && tail -c +$((off+1)) $D/log | grep -q "VacuumUpdateCosts.*cost_limit=777" && break; sleep 0.01; done
  t1=$(date +%s.%N)
  lat=$(echo "($t1 - $t0) * 1000" | bc)
  printf '%-6s leader %s waiting as %-28s voluntary wakeups in %ss: %6d (%.1f/s)   reload seen after %5.0f ms\n' \
    $B $LEADER "$WEV" $HOLD $((v1 - v0)) "$(echo "($v1 - $v0) / $HOLD" | bc -l)" "$lat"
  q "SELECT injection_points_wakeup('$WPT')"; q "SELECT injection_points_detach('$WPT')"
  "$I/pg_ctl" -D $D -m fast -w stop >/dev/null
  cp $D/log $A/wakeups.$B.server.log
  rm -rf $D
done


======== build_partial.sh ========

#!/bin/bash
# Zsolt's v3 with parts of the fix taken out, to check that each test fails
# without the part it is meant to test (the tests and injection points stay):
#   nofix    without the refresh in WaitForParallelWorkersToFinish() and
#            without the SetLatch() in autovac_recalculate_workers_for_balance()
#   nowake   only without the SetLatch()
set -eu
BASE=${BASE:-e60ee52841d}
SRC=$HOME/Proyectos/postgresql
W=$HOME/pgav
A=$(cd "$(dirname "$0")" && pwd)
P=$(ls $A/zsolt-latest/nocfbot-v3-0001-*.patch)
mk() {  # name  files-to-exclude...
  local name=$1; shift
  local tree=$W/src-$name
  if [ ! -d $tree ]; then
    git -C $SRC worktree add -q --detach $tree $BASE
    local ex=(); for f in "$@"; do ex+=(--exclude="$f"); done
    git -C $tree apply --whitespace=nowarn "${ex[@]}" "$P"
    # parallel_vacuum_refresh_cost_params() stays defined but unused when
    # parallel.c is excluded; that is fine for a test-only build.
  fi
  mkdir -p $W/b-$name && cd $W/b-$name
  $tree/configure --prefix=$W/i-$name --enable-cassert --enable-injection-points \
    --enable-tap-tests --quiet > configure.log 2>&1
  make -j"$(nproc)" -s > build.log 2>&1
  make -s install > install.log 2>&1
  make -C src/test/modules/injection_points -s install >> install.log 2>&1
  echo "$name: warnings=$(grep -c 'warning:' build.log) | $(git -C $tree diff --stat | tr '\n' ' ' | sed 's/  */ /g')"
}
mk nofix src/backend/access/transam/parallel.c src/backend/postmaster/autovacuum.c
mk nowake src/backend/postmaster/autovacuum.c
echo PARTIAL-DONE


======== run_test6.sh ========

#!/bin/bash
# Run 003_rebalance_worker_leaves.pl (Zsolt's tests + test 6) on one build.
# The file is copied into the build's source tree for this run only and is
# removed afterwards, whatever the outcome, so it never leaks into other runs.
#   run_test6.sh <build>
set -u
B=$1
A=$(cd "$(dirname "$0")" && pwd)
T=$HOME/pgav/src-$B/src/test/modules/test_autovacuum/t
cp $A/003_rebalance_worker_leaves.pl $T/
trap 'rm -f $T/003_rebalance_worker_leaves.pl' EXIT
export PG_TEST_TIMEOUT_DEFAULT=60
make -C $HOME/pgav/b-$B/src/test/modules/test_autovacuum check \
  PROVE_TESTS=t/003_rebalance_worker_leaves.pl > $A/test6.$B.log 2>&1
echo "== $B exit=$?"
L=$HOME/pgav/b-$B/src/test/modules/test_autovacuum/tmp_check/log
grep -hE '(ok|not ok) [0-9]+ - |# leader still|# parallel worker cost_limit sequence|die:|Looks like' \
  $L/regress_log_003_rebalance_worker_leaves | sed 's/^\[[^]]*\]//' | cut -c1-170
mkdir -p $A/test6-logs/$B && cp $L/* $A/test6-logs/$B/


======== results: stability, idle ========
zsolt/idle run  1 ok 9.9s
zsolt/idle run  2 ok 9.4s
zsolt/idle run  3 ok 9.5s
zsolt/idle run  4 ok 9.9s
zsolt/idle run  5 ok 9.5s
zsolt/idle run  6 ok 9.5s
zsolt/idle run  7 FAIL 9.6s
zsolt/idle run  8 ok 9.7s
zsolt/idle run  9 ok 9.4s
zsolt/idle run 10 ok 9.7s
zsolt/idle run 11 ok 9.7s
zsolt/idle run 12 ok 10.7s
zsolt/idle run 13 ok 9.8s
zsolt/idle run 14 ok 9.8s
zsolt/idle run 15 ok 9.4s
zsolt/idle run 16 ok 9.6s
zsolt/idle run 17 ok 9.6s
zsolt/idle run 18 ok 9.6s
zsolt/idle run 19 ok 9.5s
zsolt/idle run 20 ok 9.3s
zsolt/idle run 21 ok 9.7s
zsolt/idle run 22 ok 10.0s
zsolt/idle run 23 ok 9.8s
zsolt/idle run 24 ok 9.7s
zsolt/idle run 25 ok 9.2s
zsolt/idle run 26 ok 9.5s
zsolt/idle run 27 ok 9.6s
zsolt/idle run 28 ok 9.6s
zsolt/idle run 29 ok 9.4s
zsolt/idle run 30 ok 9.4s
zsolt/idle: 1 failures in 30 runs
NOTE: run 7 above is not a test failure: prove globbed t/003_rebalance_worker_leaves.pl, a file I placed in the source tree by mistake and removed during the run ('Cannot detect source'). t/001_parallel_autovacuum.pl itself passed in run 7 (see stability-fail/zsolt.idle.run7.make.log). Valid result: 001 passed 30 of 30.

======== results: stability, every CPU busy ========
zsolt/load run  1 ok 16.6s
zsolt/load run  2 ok 16.7s
zsolt/load run  3 ok 16.1s
zsolt/load run  4 ok 16.4s
zsolt/load run  5 ok 15.7s
zsolt/load run  6 ok 16.9s
zsolt/load run  7 ok 16.8s
zsolt/load run  8 ok 16.5s
zsolt/load run  9 ok 16.3s
zsolt/load run 10 ok 18.5s
zsolt/load run 11 ok 16.9s
zsolt/load run 12 ok 18.3s
zsolt/load run 13 ok 19.0s
zsolt/load run 14 ok 16.5s
zsolt/load run 15 ok 16.8s
zsolt/load run 16 ok 16.9s
zsolt/load run 17 ok 16.8s
zsolt/load run 18 ok 16.3s
zsolt/load run 19 ok 17.5s
zsolt/load run 20 ok 17.2s
zsolt/load: 0 failures in 20 runs

======== results: v3 tests on the cut-down builds ========
-- nofix
(1.395s) ok 1 - parallel autovacuum on test_autovac table
(1.159s) ok 2 - vacuum delay parameter changes are propagated to parallel vacuum workers
(0.000s) ok 3 - parallel workers see the rebalanced cost limit
(30.943s) # die: timed out waiting for file pgav/b-nofix/src/test/modules/test_autovacuum/tmp_check/log/001_parallel_autovacuum_main.log co
(0.104s) # Looks like your test exited with 255 just after 3.
-- nowake
(1.321s) ok 1 - parallel autovacuum on test_autovac table
(1.144s) ok 2 - vacuum delay parameter changes are propagated to parallel vacuum workers
(0.000s) ok 3 - parallel workers see the rebalanced cost limit
(0.872s) ok 4 - config reload reaches parallel workers while the leader waits
(31.539s) # die: timed out waiting for file pgav/b-nowake/src/test/modules/test_autovacuum/tmp_check/log/001_parallel_autovacuum_main.log c
(0.103s) # Looks like your test exited with 255 just after 4.

======== results: v3 tests + test 6 ========
== zsolt exit=0
(1.444s) ok 1 - parallel autovacuum on test_autovac table
(1.147s) ok 2 - vacuum delay parameter changes are propagated to parallel vacuum workers
(1.759s) # parallel worker cost_limit sequence: 250
(0.000s) ok 3 - parallel workers see the rebalanced cost limit
(1.194s) ok 4 - config reload reaches parallel workers while the leader waits
(1.684s) ok 5 - cost limit rebalance reaches parallel workers while the leader waits
(0.983s) # leader still in ParallelFinish when checked: 1; waited 0 s
(0.000s) ok 6 - waiting leader returns to the whole limit when a worker leaves
(0.105s) # parallel worker cost_limit sequence after the worker left: 800
(0.000s) ok 7 - parallel worker ends at the whole limit
== nowake exit=2
(1.332s) ok 1 - parallel autovacuum on test_autovac table
(1.256s) ok 2 - vacuum delay parameter changes are propagated to parallel vacuum workers
(1.657s) # parallel worker cost_limit sequence: 250
(0.000s) ok 3 - parallel workers see the rebalanced cost limit
(1.091s) ok 4 - config reload reaches parallel workers while the leader waits
(61.605s) # die: timed out waiting for file pgav/b-nowake/src/test/modules/test_autovacuum/tmp_check/log/003_rebalance_worker_leaves_main.log contents to match
(0.104s) # Looks like your test exited with 255 just after 4.

======== results: leader wakeups while waiting 20 s ========
zsolt  leader 272238 waiting as IPC/ParallelFinish           voluntary wakeups in 20s:      0 (0.0/s)   reload seen after     7 ms
nikv2  leader 278985 waiting as IPC/ParallelFinish           voluntary wakeups in 20s:    199 (9.9/s)   reload seen after     8 ms

^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-24 23:46  Masahiko Sawada <sawada.mshk@gmail.com>
  parent: Daniel Gustafsson <daniel@yesql.se>
  1 sibling, 3 replies; 43+ messages in thread

From: Masahiko Sawada @ 2026-09-24 23:46 UTC (permalink / raw)
  To: Daniel Gustafsson <daniel@yesql.se>; +Cc: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Nikolay Samokhvalov <nik@postgres.ai>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

On Thu, Sep 24, 2026 at 2:04 AM Daniel Gustafsson <daniel@yesql.se> wrote:
>
> > On 24 Sep 2026, at 08:34, Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com> wrote:
> > On Wed, Sep 23, 2026 at 5:39 PM Masahiko Sawada <sawada.mshk@gmail.com> wrote:
>
> > I also don't like the wait-100ms-wakeup approach that a separate wait function would run all the time, since it wastes power and CPU cycles. Imagine a worker vacuuming an index that is hundreds of GBs or even TBs, while the leader has only a small index and finishes first. The leader then sits in the wait loop for longer until the large index is done, waking up every 100ms the whole time. So I would prefer not to go that route.
>
> Agree, we should avoid polling loops like that as much as possible.

Agreed.

>
> > How about doing the check inside the existing wait loop, only when the process is an autovacuum worker, along the lines of the attached WIP? This is simple, the pattern already exists elsewhere in the code, and it looks safe. I also believe the wait loop is not in a performance-critical hot path, and this check is not costly. I checked that it still fixes the reported issues.
>
> It's not great to sprinkle in worker specific code in the generic parallel
> handling.  My original thinking was to reject the idea, but the vacuum costing
> is already used for non-vacuum purposes (and there have been discussions to
> rename and generalize it) so with that in mind I am less concerned for this
> particular case.

Same here. After thinking on the Bharath's patch, I now think it might
make sense. I guess it's likely that when autovacuum wants to wait for
other workers to finish in other cases like parallel heap vacuum, we
would want to use WaitForParallelWorkersToFinish() and want it to
update the cost-based delay params during the wait.

> > What I am less sure about is the fix for the second issue in the attached WIP patch, which needs any leader waiting for its workers to be woken up after the cost limit is rebalanced. I haven't found a better one yet.
>
> Not sure I see a better solution either, and we are running short of time
> before 19 RC1.

While it's a good idea to wake the leader up instead of polling, I'm a
bit concerned that we call SetLatch() on all autovacuum workers
participating in cost balancing, whereas we need it only in a narrow
situation: when the worker is participating in cost balancing, using
parallel vacuum, and waiting for its parallel workers to finish.
Calling SetLatch() on a process that is not waiting doesn't send a
signal, butit still leaves the latch set, causing a spurious wakeup
the next time the process waits for something else.

I think a condition variable fits better here. The leader can prepare
to sleep on a condition variable in AutoVacuumShmemStruct while
waiting on its latch, like WalSndWait() does, and a process
recalculating the balance can wake it up by broadcasting on it. This
way, we wake up only the leaders that actually need it. That said, I
don't think it's a good idea to make WaitForParallelWorkersToFinish()
prepare to sleep on autovacuum's condition variable. So if we want to
use the condition variable, it seems better to me to have a dedicated
wait function in vacuumparallel.c that waits until all indexes are
completed, and then call WaitForParallelWorkersToFinish().

Regards,

--
Masahiko Sawada


Amazon Web Services: https://aws.amazon.com






^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-25 00:50  Manu <manuelreyesbravo@gmail.com>
  parent: Masahiko Sawada <sawada.mshk@gmail.com>
  2 siblings, 0 replies; 43+ messages in thread

From: Manu @ 2026-09-25 00:50 UTC (permalink / raw)
  To: Masahiko Sawada <sawada.mshk@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Nikolay Samokhvalov <nik@postgres.ai>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

Hi,

Masahiko Sawada <sawada.mshk@gmail.com> wrote:
> While it's a good idea to wake the leader up instead of polling, I'm a
> bit concerned that we call SetLatch() on all autovacuum workers
> participating in cost balancing, whereas we need it only in a narrow
> situation: when the worker is participating in cost balancing, using
> parallel vacuum, and waiting for its parallel workers to finish.

I wanted to know how narrow, so I logged every SetLatch() that v3 adds,
with the target's wait event and whether its latch was already set
(the log line is in the attachment, on top of v3, REL_19_STABLE
e60ee52841d).

Load: autovacuum_max_workers = 3, three databases with the same 120
small tables each (every third table with its own
autovacuum_vacuum_cost_limit, so it is not in the balance) plus one
table with three indexes and autovacuum_parallel_workers = 2, updated
in rounds for 120 s.  That gave 14506 table vacuums, 28 of them
parallel.

  SetLatch() calls:                    10066
    on the calling worker itself:       4871
    on another worker:                  5195
      waiting in ParallelFinish:         235
      in VacuumDelay:                   4635
      anything else:                     325
    other-worker latch already set:     4724
    ParallelFinish latch already set:      1

Workers made 10001 of the calls and the launcher 65.  The count moves
every time a worker starts a table that is in the balance or not, often
2 -> 3 -> 2 within a few milliseconds, and v3 turns each move into a
SetLatch() on every balanced worker.

So you are right that almost none of them are needed: 235 out of 5195
reach a leader that is waiting.  The cost is smaller than the count
suggests, though.  A worker that is vacuuming sleeps with pg_usleep()
and does not reset its latch, so 4724 of those calls found it already
set and changed nothing.  That leaves 237 new spurious sets in
120 s across all workers, each costing at most one early return from
the worker's next wait.  Workers already get the same from a config
reload, since SignalHandlerForConfigReload() does SetLatch(MyLatch).

With a single database, only one worker ran at a time and every call
was on the caller itself.  Those 4871 calls are never needed, because
the caller runs VacuumUpdateCosts() right after the recalculation.  If
the SetLatch() stays, skipping wi_proc == MyProc removes about half of
them.  With the condition variable you describe, none of this applies,
and only the 235 waiting leaders would be woken.

The script, the log line and the full results are in the attachment.

Regards,
Manu
Who receives the SetLatch() that v3 adds to autovac_recalculate_workers_for_balance()
REL_19_STABLE e60ee52841d + v3, --enable-cassert --enable-injection-points

===== instrument.diff (logging only, on top of v3) =====
--- a/src/backend/postmaster/autovacuum.c (v3)
+++ b/src/backend/postmaster/autovacuum.c (v3 + log)
@@ -1835,6 +1835,16 @@
 				pg_atomic_unlocked_test_flag(&worker->wi_dobalance))
 				continue;
 
+			{
+				const char *we = pgstat_get_wait_event(worker->wi_proc->wait_event_info);
+
+				elog(LOG, "AVLATCH caller=%s count=%d->%d target=%d self=%d already_set=%d wait=%s",
+					 AmAutoVacuumLauncherProcess() ? "launcher" : "worker",
+					 orig_nworkers_for_balance, nworkers_for_balance,
+					 worker->wi_proc->pid, worker->wi_proc == MyProc,
+					 (int) worker->wi_proc->procLatch.is_set,
+					 we ? we : "none");
+			}
 			SetLatch(&worker->wi_proc->procLatch);
 		}
 	}

===== latchcount.sh =====
#!/bin/bash
# Who receives the SetLatch() that v3 adds to
# autovac_recalculate_workers_for_balance(), and what each target was doing
# at that moment, under an ordinary autovacuum load.
#
# Build: v3 plus one elog(LOG) before that SetLatch() (instrument.diff),
# printing the caller, the old and new count, the target pid and the
# target's wait event read from its PGPROC.
#
# Load: 3 workers, 120 small tables (every third one with its own
# autovacuum_vacuum_cost_limit, so it is not in the balance and the count
# changes as workers move between tables), and one larger table with three
# indexes and autovacuum_parallel_workers = 2, so a leader sometimes waits
# in ParallelFinish.  The tables are updated in rounds for DURATION seconds.
# NDB databases get the same tables, so up to three workers run at once.
#   latchcount.sh [DURATION] [NDB]
set -u
DURATION=${1:-120}
NDB=${2:-1}
A=$(cd "$(dirname "$0")" && pwd)
I=$HOME/pgav/i-zsinst/bin
D=$(mktemp -d /tmp/claude-1000/lc.XXXX); P=55441
"$I/initdb" -D $D -A trust --no-sync -U postgres >/dev/null
cat >> $D/postgresql.conf <<EOF
port = $P
unix_socket_directories = '/tmp'
autovacuum_max_workers = 3
autovacuum_max_parallel_workers = 2
max_worker_processes = 16
max_parallel_workers = 8
autovacuum_naptime = '1s'
autovacuum_vacuum_cost_limit = 200
autovacuum_vacuum_cost_delay = '2ms'
autovacuum_vacuum_threshold = 50
autovacuum_vacuum_scale_factor = 0
autovacuum_analyze_threshold = 1000000000
min_parallel_index_scan_size = 0
log_autovacuum_min_duration = 0
log_line_prefix = '%m [%p] '
EOF
"$I/pg_ctl" -D $D -l $D/log -w start >/dev/null
q() { "$I/psql" -X -qAt -h /tmp -p $P -U postgres "$@"; }

{
  for n in $(seq 1 120); do
    if [ $((n % 3)) = 0 ]; then opt="WITH (autovacuum_vacuum_cost_limit = 500)"; else opt=""; fi
    echo "CREATE TABLE s$n (id int PRIMARY KEY, v int) $opt;"
    echo "INSERT INTO s$n SELECT g, g FROM generate_series(1, 3000) g;"
  done
  echo "CREATE TABLE big (id int PRIMARY KEY, a int, b int, c int) WITH (autovacuum_parallel_workers = 2);"
  echo "INSERT INTO big SELECT g, g, g, g FROM generate_series(1, 400000) g;"
  echo "CREATE INDEX big_a ON big (a); CREATE INDEX big_b ON big (b); CREATE INDEX big_c ON big (c);"
} > $D/setup.sql
DBS=postgres
for k in $(seq 2 $NDB); do q -c "CREATE DATABASE d$k"; DBS="$DBS d$k"; done
for db in $DBS; do q -d $db -f $D/setup.sql; done

# One round dirties every small table; every fifth round also dirties big.
for n in $(seq 1 120); do echo "UPDATE s$n SET v = v + 1 WHERE id <= 500;"; done > $D/round.sql
echo "UPDATE big SET a = a + 1 WHERE id % 4 = 0;" > $D/big.sql

off=$(stat -c %s $D/log)
end=$(( $(date +%s) + DURATION )); r=0
while [ "$(date +%s)" -lt $end ]; do
  for db in $DBS; do q -d $db -f $D/round.sql & done; wait
  if [ $((r % 5)) = 0 ]; then for db in $DBS; do q -d $db -f $D/big.sql & done; wait; fi
  r=$((r + 1)); sleep 2
done
sleep 5
tail -c +$((off + 1)) $D/log > $A/latchcount.ndb$NDB.server.log
"$I/pg_ctl" -D $D -m fast -w stop >/dev/null
rm -rf $D

L=$A/latchcount.ndb$NDB.server.log
{
  echo "duration ${DURATION}s, $NDB database(s), $r update rounds"
  echo "tables vacuumed:        $(grep -c 'automatic vacuum of table' $L)"
  echo "  of them big (parallel): $(grep -c 'automatic vacuum of table "[a-z0-9]*.public.big"' $L)"
  echo "SetLatch() calls:       $(grep -c AVLATCH $L)"
  echo "  target is the caller itself: $(grep AVLATCH $L | grep -c 'self=1')"
  echo "  target is another worker:    $(grep AVLATCH $L | grep -c 'self=0')"
  echo "  by caller:"
  grep AVLATCH $L | grep -o 'caller=[a-z]*' | sort | uniq -c | sed 's/^/    /'
  echo "  other-worker targets whose latch was already set: $(grep AVLATCH $L | grep 'self=0' | grep -c 'already_set=1')"
  echo "  ParallelFinish targets whose latch was already set: $(grep AVLATCH $L | grep 'wait=ParallelFinish' | grep -c 'already_set=1')"
  echo "  other-worker targets, by wait event at that moment:"
  grep AVLATCH $L | grep 'self=0' | grep -o 'wait=[A-Za-z]*' | sort | uniq -c | sort -rn | sed 's/^/    /'
  echo "distinct worker pids: $(grep 'automatic vacuum of table' $L | grep -o '\[[0-9]*\]' | sort -u | wc -l)"
} | tee $A/latchcount.ndb$NDB.txt

===== ./latchcount.sh 120 3 =====
duration 120s, 3 database(s), 45 update rounds
tables vacuumed:        14506
  of them big (parallel): 28
SetLatch() calls:       10066
  target is the caller itself: 4871
  target is another worker:    5195
  by caller:
         65 caller=launcher
      10001 caller=worker
  other-worker targets whose latch was already set: 4724
  ParallelFinish targets whose latch was already set: 1
  other-worker targets, by wait event at that moment:
       4635 wait=VacuumDelay
        235 wait=ParallelFinish
        117 wait=none
         96 wait=AioIoCompletion
         66 wait=WALWrite
         21 wait=WalWrite
         13 wait=BufferExclusive
         10 wait=DataFileWrite
          1 wait=WALBufMapping
          1 wait=DataFileRead
distinct worker pids: 148

(self=1 calls whose latch was already set: 814)

===== ./latchcount.sh 120 1 (earlier build: same log line without already_set) =====
duration 120s, 54 update rounds
tables vacuumed:        6613
  of them big (parallel): 12
SetLatch() calls:       2212
  target is the caller itself: 2212
  target is another worker:    0
  by caller:
       2212 caller=worker
  other-worker targets, by wait event at that moment:

===== sample of the log, 3 databases =====
2026-09-24 21:35:06.263 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3773519 self=1 already_set=0 wait=none
2026-09-24 21:35:06.263 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.263 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.268 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.268 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.268 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3773519 self=1 already_set=0 wait=none
2026-09-24 21:35:06.268 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.268 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.274 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.274 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.274 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3773519 self=1 already_set=0 wait=none
2026-09-24 21:35:06.274 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.274 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.279 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.279 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.280 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3773519 self=1 already_set=0 wait=none
2026-09-24 21:35:06.280 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.280 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.285 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.285 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.285 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3773519 self=1 already_set=0 wait=none
2026-09-24 21:35:06.285 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.285 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.290 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.290 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.290 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3773519 self=1 already_set=0 wait=none
2026-09-24 21:35:06.290 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.290 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.295 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.295 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.296 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3773519 self=1 already_set=0 wait=none


Attachments:

  [text/plain] nocfbot-av-setlatch-targets.txt (10.1K, ../../179029740211.4041151.3406196415226766626@gmail.com/2-nocfbot-av-setlatch-targets.txt)
  download | inline diff:
Who receives the SetLatch() that v3 adds to autovac_recalculate_workers_for_balance()
REL_19_STABLE e60ee52841d + v3, --enable-cassert --enable-injection-points

===== instrument.diff (logging only, on top of v3) =====
--- a/src/backend/postmaster/autovacuum.c (v3)
+++ b/src/backend/postmaster/autovacuum.c (v3 + log)
@@ -1835,6 +1835,16 @@
 				pg_atomic_unlocked_test_flag(&worker->wi_dobalance))
 				continue;
 
+			{
+				const char *we = pgstat_get_wait_event(worker->wi_proc->wait_event_info);
+
+				elog(LOG, "AVLATCH caller=%s count=%d->%d target=%d self=%d already_set=%d wait=%s",
+					 AmAutoVacuumLauncherProcess() ? "launcher" : "worker",
+					 orig_nworkers_for_balance, nworkers_for_balance,
+					 worker->wi_proc->pid, worker->wi_proc == MyProc,
+					 (int) worker->wi_proc->procLatch.is_set,
+					 we ? we : "none");
+			}
 			SetLatch(&worker->wi_proc->procLatch);
 		}
 	}

===== latchcount.sh =====
#!/bin/bash
# Who receives the SetLatch() that v3 adds to
# autovac_recalculate_workers_for_balance(), and what each target was doing
# at that moment, under an ordinary autovacuum load.
#
# Build: v3 plus one elog(LOG) before that SetLatch() (instrument.diff),
# printing the caller, the old and new count, the target pid and the
# target's wait event read from its PGPROC.
#
# Load: 3 workers, 120 small tables (every third one with its own
# autovacuum_vacuum_cost_limit, so it is not in the balance and the count
# changes as workers move between tables), and one larger table with three
# indexes and autovacuum_parallel_workers = 2, so a leader sometimes waits
# in ParallelFinish.  The tables are updated in rounds for DURATION seconds.
# NDB databases get the same tables, so up to three workers run at once.
#   latchcount.sh [DURATION] [NDB]
set -u
DURATION=${1:-120}
NDB=${2:-1}
A=$(cd "$(dirname "$0")" && pwd)
I=$HOME/pgav/i-zsinst/bin
D=$(mktemp -d /tmp/claude-1000/lc.XXXX); P=55441
"$I/initdb" -D $D -A trust --no-sync -U postgres >/dev/null
cat >> $D/postgresql.conf <<EOF
port = $P
unix_socket_directories = '/tmp'
autovacuum_max_workers = 3
autovacuum_max_parallel_workers = 2
max_worker_processes = 16
max_parallel_workers = 8
autovacuum_naptime = '1s'
autovacuum_vacuum_cost_limit = 200
autovacuum_vacuum_cost_delay = '2ms'
autovacuum_vacuum_threshold = 50
autovacuum_vacuum_scale_factor = 0
autovacuum_analyze_threshold = 1000000000
min_parallel_index_scan_size = 0
log_autovacuum_min_duration = 0
log_line_prefix = '%m [%p] '
EOF
"$I/pg_ctl" -D $D -l $D/log -w start >/dev/null
q() { "$I/psql" -X -qAt -h /tmp -p $P -U postgres "$@"; }

{
  for n in $(seq 1 120); do
    if [ $((n % 3)) = 0 ]; then opt="WITH (autovacuum_vacuum_cost_limit = 500)"; else opt=""; fi
    echo "CREATE TABLE s$n (id int PRIMARY KEY, v int) $opt;"
    echo "INSERT INTO s$n SELECT g, g FROM generate_series(1, 3000) g;"
  done
  echo "CREATE TABLE big (id int PRIMARY KEY, a int, b int, c int) WITH (autovacuum_parallel_workers = 2);"
  echo "INSERT INTO big SELECT g, g, g, g FROM generate_series(1, 400000) g;"
  echo "CREATE INDEX big_a ON big (a); CREATE INDEX big_b ON big (b); CREATE INDEX big_c ON big (c);"
} > $D/setup.sql
DBS=postgres
for k in $(seq 2 $NDB); do q -c "CREATE DATABASE d$k"; DBS="$DBS d$k"; done
for db in $DBS; do q -d $db -f $D/setup.sql; done

# One round dirties every small table; every fifth round also dirties big.
for n in $(seq 1 120); do echo "UPDATE s$n SET v = v + 1 WHERE id <= 500;"; done > $D/round.sql
echo "UPDATE big SET a = a + 1 WHERE id % 4 = 0;" > $D/big.sql

off=$(stat -c %s $D/log)
end=$(( $(date +%s) + DURATION )); r=0
while [ "$(date +%s)" -lt $end ]; do
  for db in $DBS; do q -d $db -f $D/round.sql & done; wait
  if [ $((r % 5)) = 0 ]; then for db in $DBS; do q -d $db -f $D/big.sql & done; wait; fi
  r=$((r + 1)); sleep 2
done
sleep 5
tail -c +$((off + 1)) $D/log > $A/latchcount.ndb$NDB.server.log
"$I/pg_ctl" -D $D -m fast -w stop >/dev/null
rm -rf $D

L=$A/latchcount.ndb$NDB.server.log
{
  echo "duration ${DURATION}s, $NDB database(s), $r update rounds"
  echo "tables vacuumed:        $(grep -c 'automatic vacuum of table' $L)"
  echo "  of them big (parallel): $(grep -c 'automatic vacuum of table "[a-z0-9]*.public.big"' $L)"
  echo "SetLatch() calls:       $(grep -c AVLATCH $L)"
  echo "  target is the caller itself: $(grep AVLATCH $L | grep -c 'self=1')"
  echo "  target is another worker:    $(grep AVLATCH $L | grep -c 'self=0')"
  echo "  by caller:"
  grep AVLATCH $L | grep -o 'caller=[a-z]*' | sort | uniq -c | sed 's/^/    /'
  echo "  other-worker targets whose latch was already set: $(grep AVLATCH $L | grep 'self=0' | grep -c 'already_set=1')"
  echo "  ParallelFinish targets whose latch was already set: $(grep AVLATCH $L | grep 'wait=ParallelFinish' | grep -c 'already_set=1')"
  echo "  other-worker targets, by wait event at that moment:"
  grep AVLATCH $L | grep 'self=0' | grep -o 'wait=[A-Za-z]*' | sort | uniq -c | sort -rn | sed 's/^/    /'
  echo "distinct worker pids: $(grep 'automatic vacuum of table' $L | grep -o '\[[0-9]*\]' | sort -u | wc -l)"
} | tee $A/latchcount.ndb$NDB.txt

===== ./latchcount.sh 120 3 =====
duration 120s, 3 database(s), 45 update rounds
tables vacuumed:        14506
  of them big (parallel): 28
SetLatch() calls:       10066
  target is the caller itself: 4871
  target is another worker:    5195
  by caller:
         65 caller=launcher
      10001 caller=worker
  other-worker targets whose latch was already set: 4724
  ParallelFinish targets whose latch was already set: 1
  other-worker targets, by wait event at that moment:
       4635 wait=VacuumDelay
        235 wait=ParallelFinish
        117 wait=none
         96 wait=AioIoCompletion
         66 wait=WALWrite
         21 wait=WalWrite
         13 wait=BufferExclusive
         10 wait=DataFileWrite
          1 wait=WALBufMapping
          1 wait=DataFileRead
distinct worker pids: 148

(self=1 calls whose latch was already set: 814)

===== ./latchcount.sh 120 1 (earlier build: same log line without already_set) =====
duration 120s, 54 update rounds
tables vacuumed:        6613
  of them big (parallel): 12
SetLatch() calls:       2212
  target is the caller itself: 2212
  target is another worker:    0
  by caller:
       2212 caller=worker
  other-worker targets, by wait event at that moment:

===== sample of the log, 3 databases =====
2026-09-24 21:35:06.263 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3773519 self=1 already_set=0 wait=none
2026-09-24 21:35:06.263 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.263 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.268 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.268 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.268 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3773519 self=1 already_set=0 wait=none
2026-09-24 21:35:06.268 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.268 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.274 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.274 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.274 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3773519 self=1 already_set=0 wait=none
2026-09-24 21:35:06.274 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.274 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.279 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.279 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.280 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3773519 self=1 already_set=0 wait=none
2026-09-24 21:35:06.280 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.280 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.285 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.285 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.285 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3773519 self=1 already_set=0 wait=none
2026-09-24 21:35:06.285 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.285 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.290 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.290 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.290 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3773519 self=1 already_set=0 wait=none
2026-09-24 21:35:06.290 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.290 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.295 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772425 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.295 -03 [3773519] LOG:  AVLATCH caller=worker count=3->2 target=3772327 self=0 already_set=1 wait=VacuumDelay
2026-09-24 21:35:06.296 -03 [3773519] LOG:  AVLATCH caller=worker count=2->3 target=3773519 self=1 already_set=0 wait=none


^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-25 07:45  Zsolt Parragi <zsolt.parragi@percona.com>
  parent: Masahiko Sawada <sawada.mshk@gmail.com>
  2 siblings, 1 reply; 43+ messages in thread

From: Zsolt Parragi @ 2026-09-25 07:45 UTC (permalink / raw)
  To: Masahiko Sawada <sawada.mshk@gmail.com>; +Cc: pgsql-bugs@lists.postgresql.org, Daniel Gustafsson <daniel@yesql.se>; Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>

On Thu, 24 Sep 2026, Masahiko Sawada <sawada.mshk@gmail.com> wrote:
> So if we want to
> use the condition variable, it seems better to me to have a dedicated
> wait function in vacuumparallel.c that waits until all indexes are
> completed, and then call WaitForParallelWorkersToFinish().

That'll bring back the more complicated approach of v2, and also make
this part of the code more complex.

While I didn't do proper measurements, my logic against it was similar
to what Manu already posted, that it seems harmless here and it is
significantly simpler than adding a cv.

Also, I think both issues could be solved in a more proper way with
more refactoring, but that seems excessive for 19. (And in that case,
going with a simple solution for 19 also seems like a good answer to
me)





^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-25 08:41  Daniel Gustafsson <daniel@yesql.se>
  parent: Zsolt Parragi <zsolt.parragi@percona.com>
  0 siblings, 1 reply; 43+ messages in thread

From: Daniel Gustafsson @ 2026-09-25 08:41 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Masahiko Sawada <sawada.mshk@gmail.com>; pgsql-bugs@lists.postgresql.org, Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>

> On 25 Sep 2026, at 09:45, Zsolt Parragi <zsolt.parragi@percona.com> wrote:
> 
> On Thu, 24 Sep 2026, Masahiko Sawada <sawada.mshk@gmail.com> wrote:
>> So if we want to
>> use the condition variable, it seems better to me to have a dedicated
>> wait function in vacuumparallel.c that waits until all indexes are
>> completed, and then call WaitForParallelWorkersToFinish().
> 
> That'll bring back the more complicated approach of v2, and also make
> this part of the code more complex.

I think that would be cleaner, but also more invasive at post-beta4.

> While I didn't do proper measurements, my logic against it was similar
> to what Manu already posted, that it seems harmless here and it is
> significantly simpler than adding a cv.
> 
> Also, I think both issues could be solved in a more proper way with
> more refactoring, but that seems excessive for 19. (And in that case,
> going with a simple solution for 19 also seems like a good answer to
> me)

Given where we are in the cycle I am also in favor of a simpler solution unless
it's shown to have (severe) performance regressions.

--
Daniel Gustafsson







^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-25 19:12  Manu <manuelreyesbravo@gmail.com>
  parent: Daniel Gustafsson <daniel@yesql.se>
  0 siblings, 0 replies; 43+ messages in thread

From: Manu @ 2026-09-25 19:12 UTC (permalink / raw)
  To: Daniel Gustafsson <daniel@yesql.se>; +Cc: Masahiko Sawada <sawada.mshk@gmail.com>; Zsolt Parragi <zsolt.parragi@percona.com>; Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Nikolay Samokhvalov <nik@postgres.ai>; pgsql-bugs@lists.postgresql.org

Hi,

Daniel Gustafsson <daniel@yesql.se> wrote:
> Given where we are in the cycle I am also in favor of a simpler solution unless
> it's shown to have (severe) performance regressions.

My earlier numbers counted the SetLatch() calls, not their cost, so I
measured the server's CPU with and without v3 (REL_19_STABLE
e60ee52841d, -O2, no cassert).  The load is the same as before: small
tables moving in and out of the balance plus one parallel table, with
autovacuum saturated, 120 s per run, runs alternated.

3 workers, 8 runs each: median 25.75 s CPU both with and without v3.
One v3 run used 28.4 s, and it did not repeat in 7 more.

10 workers, 4 runs each: 125.8 s without v3 (sd 0.8), 126.1 s with
it (sd 0.3).

So I see no CPU cost from v3 at either size.  One difference did show
up at 10 workers: v3 vacuumed 2.5% fewer tables in the same window, in
all four pairs.  I have not checked why.  It would fit the parallel
workers now following the rebalanced limit, but that is a guess.

The scripts and all runs are attached.

Regards,
Manu
Server CPU with and without v3 under a changing autovacuum balance
REL_19_STABLE e60ee52841d, with and without v3, both built with CFLAGS=-O2 and no cassert
24-core machine, otherwise idle (load average in each line)

===== perfcost.sh =====
#!/bin/bash
# What v3's SetLatch() in autovac_recalculate_workers_for_balance() costs,
# under the same load as latchcount.sh: NDB databases with 120 small tables
# each (every third one out of the balance, so the balance count keeps
# changing) plus one table vacuumed in parallel, and as many autovacuum
# workers as databases.  v3 sets the latch of every balanced worker on each
# change, so the number of calls grows with the square of the workers.
#
# Reports, for the DURATION-second load window:
#   server_cpu_s   CPU of the whole server, read from its cgroup
#   vacuums        tables vacuumed
# (log_autovacuum's own CPU figures are rounded to 10 ms, and most of these
# vacuums take less, so they are not used.)
#
#   perfcost.sh BUILD NDB DURATION   (BUILD = pbase | pv3)
set -u
BUILD=$1 NDB=$2 DURATION=$3
I=$HOME/pgav/i-$BUILD/bin
D=$(mktemp -d /tmp/claude-1000/pc.XXXX); P=55451
UNIT=avperf-$BUILD-$$
"$I/initdb" -D $D -A trust --no-sync -U postgres >/dev/null
cat >> $D/postgresql.conf <<EOF
port = $P
unix_socket_directories = '/tmp'
autovacuum_max_workers = $NDB
autovacuum_max_parallel_workers = 2
max_worker_processes = 32
max_parallel_workers = 16
autovacuum_naptime = '1s'
autovacuum_vacuum_cost_limit = 200
autovacuum_vacuum_cost_delay = '2ms'
autovacuum_vacuum_threshold = 50
autovacuum_vacuum_scale_factor = 0
autovacuum_analyze_threshold = 1000000000
min_parallel_index_scan_size = 0
log_autovacuum_min_duration = 0
log_line_prefix = '%m [%p] '
EOF
# The server runs in its own scope, so its CPU can be read from the cgroup.
systemd-run --user --scope --quiet --unit=$UNIT \
  "$I/pg_ctl" -D $D -l $D/log -w start >/dev/null
q() { "$I/psql" -X -qAt -h /tmp -p $P -U postgres "$@"; }
cpu() { systemctl --user show -P CPUUsageNSec $UNIT.scope; }

{
  for n in $(seq 1 120); do
    if [ $((n % 3)) = 0 ]; then opt="WITH (autovacuum_vacuum_cost_limit = 500)"; else opt=""; fi
    echo "CREATE TABLE s$n (id int PRIMARY KEY, v int) $opt;"
    echo "INSERT INTO s$n SELECT g, g FROM generate_series(1, 3000) g;"
  done
  echo "CREATE TABLE big (id int PRIMARY KEY, a int, b int, c int) WITH (autovacuum_parallel_workers = 2);"
  echo "INSERT INTO big SELECT g, g, g, g FROM generate_series(1, 400000) g;"
  echo "CREATE INDEX big_a ON big (a); CREATE INDEX big_b ON big (b); CREATE INDEX big_c ON big (c);"
} > $D/setup.sql
DBS=postgres
for k in $(seq 2 $NDB); do q -c "CREATE DATABASE d$k"; DBS="$DBS d$k"; done
for db in $DBS; do q -d $db -f $D/setup.sql; done
# 60 rows: just over autovacuum_vacuum_threshold, so the server's CPU goes
# mostly to autovacuum rather than to the updates themselves.
for n in $(seq 1 120); do echo "UPDATE s$n SET v = v + 1 WHERE id <= 60;"; done > $D/round.sql
echo "UPDATE big SET a = a + 1 WHERE id % 4 = 0;" > $D/big.sql
sleep 20   # let autovacuum finish with the setup

off=$(stat -c %s $D/log); c0=$(cpu)
end=$(( $(date +%s) + DURATION )); r=0
while [ "$(date +%s)" -lt $end ]; do
  for db in $DBS; do q -d $db -f $D/round.sql & done; wait
  if [ $((r % 5)) = 0 ]; then for db in $DBS; do q -d $db -f $D/big.sql & done; wait; fi
  r=$((r + 1)); sleep 2
done
c1=$(cpu)
L=$D/window.log; tail -c +$((off + 1)) $D/log > $L
"$I/pg_ctl" -D $D -m fast -w stop >/dev/null

vac=$(grep -c 'automatic vacuum of table' $L)
awk -v b=$BUILD -v n=$NDB -v r=$r -v c=$(( (c1 - c0) / 1000000 )) -v v=$vac 'BEGIN {
  printf "build=%s ndb=%d rounds=%d server_cpu_s=%.2f vacuums=%d\n", b, n, r, c / 1000, v }'
rm -rf $D

===== perf_all.sh =====
#!/bin/bash
# perfcost.sh for pbase and pv3, alternating, REPS times each, with 3 and 10
# workers.  Alternating spreads any drift in the machine's load over both.
#   perf_all.sh [DURATION] [REPS]
set -u
DURATION=${1:-120}
REPS=${2:-4}
A=$(cd "$(dirname "$0")" && pwd)
OUT=$A/perfcost.runs.txt
: > $OUT
for ndb in 3 10; do
  for rep in $(seq 1 $REPS); do
    for b in pbase pv3; do
      echo "load: $(cut -d' ' -f1-3 /proc/loadavg) | $(bash $A/perfcost.sh $b $ndb $DURATION)" | tee -a $OUT
    done
  done
done
echo
awk '{
  for (i = 1; i <= NF; i++) { split($i, kv, "="); f[kv[1]] = kv[2] }
  k = f["ndb"] " " f["build"]; n[k]++
  c[k] += f["server_cpu_s"]; cc[k] += f["server_cpu_s"] ^ 2
  v[k] += f["vacuums"]; r[k] += f["rounds"]
} END {
  for (k in n) {
    m = c[k] / n[k]; sd = (n[k] > 1) ? sqrt((cc[k] - n[k] * m * m) / (n[k] - 1)) : 0
    printf "ndb=%s runs=%d server_cpu_s=%.2f sd=%.2f vacuums=%.0f rounds=%.1f cpu_ms_per_vacuum=%.3f\n",
           k, n[k], m, sd, v[k] / n[k], r[k] / n[k], c[k] * 1000 / v[k]
  }
}' $OUT | sort | tee $A/perfcost.summary.txt

===== runs: perf_all.sh 120 4 (pbase first in each pair) =====
load: 1.99 2.71 6.71 | build=pbase ndb=3 rounds=57 server_cpu_s=25.20 vacuums=15522
load: 0.95 2.00 5.88 | build=pv3 ndb=3 rounds=57 server_cpu_s=28.43 vacuums=15337
load: 1.94 2.03 5.32 | build=pbase ndb=3 rounds=57 server_cpu_s=25.22 vacuums=15766
load: 1.22 1.80 4.78 | build=pv3 ndb=3 rounds=57 server_cpu_s=25.67 vacuums=15384
load: 1.19 1.62 4.28 | build=pbase ndb=3 rounds=57 server_cpu_s=25.80 vacuums=15633
load: 1.32 1.44 3.81 | build=pv3 ndb=3 rounds=57 server_cpu_s=25.69 vacuums=15404
load: 1.02 1.20 3.38 | build=pbase ndb=3 rounds=57 server_cpu_s=25.70 vacuums=15039
load: 1.18 1.22 3.07 | build=pv3 ndb=3 rounds=57 server_cpu_s=25.99 vacuums=15697
load: 1.02 1.11 2.77 | build=pbase ndb=10 rounds=53 server_cpu_s=125.34 vacuums=23643
load: 0.94 1.11 2.52 | build=pv3 ndb=10 rounds=53 server_cpu_s=126.33 vacuums=23213
load: 1.44 1.32 2.39 | build=pbase ndb=10 rounds=53 server_cpu_s=124.91 vacuums=24138
load: 1.99 1.64 2.35 | build=pv3 ndb=10 rounds=53 server_cpu_s=126.28 vacuums=23771
load: 1.66 1.62 2.24 | build=pbase ndb=10 rounds=53 server_cpu_s=126.60 vacuums=24658
load: 2.41 1.82 2.21 | build=pv3 ndb=10 rounds=53 server_cpu_s=125.69 vacuums=23555
load: 1.62 1.67 2.09 | build=pbase ndb=10 rounds=53 server_cpu_s=126.27 vacuums=24415
load: 1.82 1.83 2.09 | build=pv3 ndb=10 rounds=53 server_cpu_s=126.24 vacuums=23929

===== summary of those runs =====
ndb=10 pbase runs=4 server_cpu_s=125.78 sd=0.79 vacuums=24214 rounds=53.0 cpu_ms_per_vacuum=5.195
ndb=10 pv3 runs=4 server_cpu_s=126.14 sd=0.30 vacuums=23617 rounds=53.0 cpu_ms_per_vacuum=5.341
ndb=3 pbase runs=4 server_cpu_s=25.48 sd=0.31 vacuums=15490 rounds=57.0 cpu_ms_per_vacuum=1.645
ndb=3 pv3 runs=4 server_cpu_s=26.45 sd=1.33 vacuums=15456 rounds=57.0 cpu_ms_per_vacuum=1.711

===== 4 more pairs with 3 workers, pv3 first in each pair =====
load: 1.46 1.66 1.98 | build=pv3 ndb=3 rounds=57 server_cpu_s=25.67 vacuums=15426
load: 0.71 1.31 1.80 | build=pbase ndb=3 rounds=57 server_cpu_s=25.33 vacuums=15222
load: 0.68 1.15 1.67 | build=pv3 ndb=3 rounds=57 server_cpu_s=25.50 vacuums=15436
load: 1.05 1.07 1.56 | build=pbase ndb=3 rounds=57 server_cpu_s=25.90 vacuums=15642
load: 1.45 1.11 1.49 | build=pv3 ndb=3 rounds=57 server_cpu_s=25.81 vacuums=15755
load: 0.83 0.95 1.37 | build=pbase ndb=3 rounds=57 server_cpu_s=26.26 vacuums=15395
load: 0.97 1.01 1.33 | build=pv3 ndb=3 rounds=57 server_cpu_s=26.39 vacuums=15514
load: 0.62 0.95 1.27 | build=pbase ndb=3 rounds=57 server_cpu_s=26.63 vacuums=15988


Attachments:

  [text/plain] nocfbot-av-v3-cpu.txt (7.2K, ../../179036354055.1964197.12261940030767792554@gmail.com/2-nocfbot-av-v3-cpu.txt)
  download | inline:
Server CPU with and without v3 under a changing autovacuum balance
REL_19_STABLE e60ee52841d, with and without v3, both built with CFLAGS=-O2 and no cassert
24-core machine, otherwise idle (load average in each line)

===== perfcost.sh =====
#!/bin/bash
# What v3's SetLatch() in autovac_recalculate_workers_for_balance() costs,
# under the same load as latchcount.sh: NDB databases with 120 small tables
# each (every third one out of the balance, so the balance count keeps
# changing) plus one table vacuumed in parallel, and as many autovacuum
# workers as databases.  v3 sets the latch of every balanced worker on each
# change, so the number of calls grows with the square of the workers.
#
# Reports, for the DURATION-second load window:
#   server_cpu_s   CPU of the whole server, read from its cgroup
#   vacuums        tables vacuumed
# (log_autovacuum's own CPU figures are rounded to 10 ms, and most of these
# vacuums take less, so they are not used.)
#
#   perfcost.sh BUILD NDB DURATION   (BUILD = pbase | pv3)
set -u
BUILD=$1 NDB=$2 DURATION=$3
I=$HOME/pgav/i-$BUILD/bin
D=$(mktemp -d /tmp/claude-1000/pc.XXXX); P=55451
UNIT=avperf-$BUILD-$$
"$I/initdb" -D $D -A trust --no-sync -U postgres >/dev/null
cat >> $D/postgresql.conf <<EOF
port = $P
unix_socket_directories = '/tmp'
autovacuum_max_workers = $NDB
autovacuum_max_parallel_workers = 2
max_worker_processes = 32
max_parallel_workers = 16
autovacuum_naptime = '1s'
autovacuum_vacuum_cost_limit = 200
autovacuum_vacuum_cost_delay = '2ms'
autovacuum_vacuum_threshold = 50
autovacuum_vacuum_scale_factor = 0
autovacuum_analyze_threshold = 1000000000
min_parallel_index_scan_size = 0
log_autovacuum_min_duration = 0
log_line_prefix = '%m [%p] '
EOF
# The server runs in its own scope, so its CPU can be read from the cgroup.
systemd-run --user --scope --quiet --unit=$UNIT \
  "$I/pg_ctl" -D $D -l $D/log -w start >/dev/null
q() { "$I/psql" -X -qAt -h /tmp -p $P -U postgres "$@"; }
cpu() { systemctl --user show -P CPUUsageNSec $UNIT.scope; }

{
  for n in $(seq 1 120); do
    if [ $((n % 3)) = 0 ]; then opt="WITH (autovacuum_vacuum_cost_limit = 500)"; else opt=""; fi
    echo "CREATE TABLE s$n (id int PRIMARY KEY, v int) $opt;"
    echo "INSERT INTO s$n SELECT g, g FROM generate_series(1, 3000) g;"
  done
  echo "CREATE TABLE big (id int PRIMARY KEY, a int, b int, c int) WITH (autovacuum_parallel_workers = 2);"
  echo "INSERT INTO big SELECT g, g, g, g FROM generate_series(1, 400000) g;"
  echo "CREATE INDEX big_a ON big (a); CREATE INDEX big_b ON big (b); CREATE INDEX big_c ON big (c);"
} > $D/setup.sql
DBS=postgres
for k in $(seq 2 $NDB); do q -c "CREATE DATABASE d$k"; DBS="$DBS d$k"; done
for db in $DBS; do q -d $db -f $D/setup.sql; done
# 60 rows: just over autovacuum_vacuum_threshold, so the server's CPU goes
# mostly to autovacuum rather than to the updates themselves.
for n in $(seq 1 120); do echo "UPDATE s$n SET v = v + 1 WHERE id <= 60;"; done > $D/round.sql
echo "UPDATE big SET a = a + 1 WHERE id % 4 = 0;" > $D/big.sql
sleep 20   # let autovacuum finish with the setup

off=$(stat -c %s $D/log); c0=$(cpu)
end=$(( $(date +%s) + DURATION )); r=0
while [ "$(date +%s)" -lt $end ]; do
  for db in $DBS; do q -d $db -f $D/round.sql & done; wait
  if [ $((r % 5)) = 0 ]; then for db in $DBS; do q -d $db -f $D/big.sql & done; wait; fi
  r=$((r + 1)); sleep 2
done
c1=$(cpu)
L=$D/window.log; tail -c +$((off + 1)) $D/log > $L
"$I/pg_ctl" -D $D -m fast -w stop >/dev/null

vac=$(grep -c 'automatic vacuum of table' $L)
awk -v b=$BUILD -v n=$NDB -v r=$r -v c=$(( (c1 - c0) / 1000000 )) -v v=$vac 'BEGIN {
  printf "build=%s ndb=%d rounds=%d server_cpu_s=%.2f vacuums=%d\n", b, n, r, c / 1000, v }'
rm -rf $D

===== perf_all.sh =====
#!/bin/bash
# perfcost.sh for pbase and pv3, alternating, REPS times each, with 3 and 10
# workers.  Alternating spreads any drift in the machine's load over both.
#   perf_all.sh [DURATION] [REPS]
set -u
DURATION=${1:-120}
REPS=${2:-4}
A=$(cd "$(dirname "$0")" && pwd)
OUT=$A/perfcost.runs.txt
: > $OUT
for ndb in 3 10; do
  for rep in $(seq 1 $REPS); do
    for b in pbase pv3; do
      echo "load: $(cut -d' ' -f1-3 /proc/loadavg) | $(bash $A/perfcost.sh $b $ndb $DURATION)" | tee -a $OUT
    done
  done
done
echo
awk '{
  for (i = 1; i <= NF; i++) { split($i, kv, "="); f[kv[1]] = kv[2] }
  k = f["ndb"] " " f["build"]; n[k]++
  c[k] += f["server_cpu_s"]; cc[k] += f["server_cpu_s"] ^ 2
  v[k] += f["vacuums"]; r[k] += f["rounds"]
} END {
  for (k in n) {
    m = c[k] / n[k]; sd = (n[k] > 1) ? sqrt((cc[k] - n[k] * m * m) / (n[k] - 1)) : 0
    printf "ndb=%s runs=%d server_cpu_s=%.2f sd=%.2f vacuums=%.0f rounds=%.1f cpu_ms_per_vacuum=%.3f\n",
           k, n[k], m, sd, v[k] / n[k], r[k] / n[k], c[k] * 1000 / v[k]
  }
}' $OUT | sort | tee $A/perfcost.summary.txt

===== runs: perf_all.sh 120 4 (pbase first in each pair) =====
load: 1.99 2.71 6.71 | build=pbase ndb=3 rounds=57 server_cpu_s=25.20 vacuums=15522
load: 0.95 2.00 5.88 | build=pv3 ndb=3 rounds=57 server_cpu_s=28.43 vacuums=15337
load: 1.94 2.03 5.32 | build=pbase ndb=3 rounds=57 server_cpu_s=25.22 vacuums=15766
load: 1.22 1.80 4.78 | build=pv3 ndb=3 rounds=57 server_cpu_s=25.67 vacuums=15384
load: 1.19 1.62 4.28 | build=pbase ndb=3 rounds=57 server_cpu_s=25.80 vacuums=15633
load: 1.32 1.44 3.81 | build=pv3 ndb=3 rounds=57 server_cpu_s=25.69 vacuums=15404
load: 1.02 1.20 3.38 | build=pbase ndb=3 rounds=57 server_cpu_s=25.70 vacuums=15039
load: 1.18 1.22 3.07 | build=pv3 ndb=3 rounds=57 server_cpu_s=25.99 vacuums=15697
load: 1.02 1.11 2.77 | build=pbase ndb=10 rounds=53 server_cpu_s=125.34 vacuums=23643
load: 0.94 1.11 2.52 | build=pv3 ndb=10 rounds=53 server_cpu_s=126.33 vacuums=23213
load: 1.44 1.32 2.39 | build=pbase ndb=10 rounds=53 server_cpu_s=124.91 vacuums=24138
load: 1.99 1.64 2.35 | build=pv3 ndb=10 rounds=53 server_cpu_s=126.28 vacuums=23771
load: 1.66 1.62 2.24 | build=pbase ndb=10 rounds=53 server_cpu_s=126.60 vacuums=24658
load: 2.41 1.82 2.21 | build=pv3 ndb=10 rounds=53 server_cpu_s=125.69 vacuums=23555
load: 1.62 1.67 2.09 | build=pbase ndb=10 rounds=53 server_cpu_s=126.27 vacuums=24415
load: 1.82 1.83 2.09 | build=pv3 ndb=10 rounds=53 server_cpu_s=126.24 vacuums=23929

===== summary of those runs =====
ndb=10 pbase runs=4 server_cpu_s=125.78 sd=0.79 vacuums=24214 rounds=53.0 cpu_ms_per_vacuum=5.195
ndb=10 pv3 runs=4 server_cpu_s=126.14 sd=0.30 vacuums=23617 rounds=53.0 cpu_ms_per_vacuum=5.341
ndb=3 pbase runs=4 server_cpu_s=25.48 sd=0.31 vacuums=15490 rounds=57.0 cpu_ms_per_vacuum=1.645
ndb=3 pv3 runs=4 server_cpu_s=26.45 sd=1.33 vacuums=15456 rounds=57.0 cpu_ms_per_vacuum=1.711

===== 4 more pairs with 3 workers, pv3 first in each pair =====
load: 1.46 1.66 1.98 | build=pv3 ndb=3 rounds=57 server_cpu_s=25.67 vacuums=15426
load: 0.71 1.31 1.80 | build=pbase ndb=3 rounds=57 server_cpu_s=25.33 vacuums=15222
load: 0.68 1.15 1.67 | build=pv3 ndb=3 rounds=57 server_cpu_s=25.50 vacuums=15436
load: 1.05 1.07 1.56 | build=pbase ndb=3 rounds=57 server_cpu_s=25.90 vacuums=15642
load: 1.45 1.11 1.49 | build=pv3 ndb=3 rounds=57 server_cpu_s=25.81 vacuums=15755
load: 0.83 0.95 1.37 | build=pbase ndb=3 rounds=57 server_cpu_s=26.26 vacuums=15395
load: 0.97 1.01 1.33 | build=pv3 ndb=3 rounds=57 server_cpu_s=26.39 vacuums=15514
load: 0.62 0.95 1.27 | build=pbase ndb=3 rounds=57 server_cpu_s=26.63 vacuums=15988

^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-25 21:41  Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
  parent: Masahiko Sawada <sawada.mshk@gmail.com>
  2 siblings, 2 replies; 43+ messages in thread

From: Bharath Rupireddy @ 2026-09-25 21:41 UTC (permalink / raw)
  To: Masahiko Sawada <sawada.mshk@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; Nikolay Samokhvalov <nik@postgres.ai>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

Hi,

On Thu, Sep 24, 2026 at 4:47 PM Masahiko Sawada <sawada.mshk@gmail.com> wrote:
>
> > > How about doing the check inside the existing wait loop, only when the process is an autovacuum worker, along the lines of the attached WIP? This is simple, the pattern already exists elsewhere in the code, and it looks safe. I also believe the wait loop is not in a performance-critical hot path, and this check is not costly. I checked that it still fixes the reported issues.
>
> > It's not great to sprinkle in worker specific code in the generic parallel
> > handling.  My original thinking was to reject the idea, but the vacuum costing
> > is already used for non-vacuum purposes (and there have been discussions to
> > rename and generalize it) so with that in mind I am less concerned for this
> > particular case.
>
> Same here. After thinking on the Bharath's patch, I now think it might
> make sense. I guess it's likely that when autovacuum wants to wait for
> other workers to finish in other cases like parallel heap vacuum, we
> would want to use WaitForParallelWorkersToFinish() and want it to
> update the cost-based delay params during the wait.
>
> > > What I am less sure about is the fix for the second issue in the attached WIP patch, which needs any leader waiting for its workers to be woken up after the cost limit is rebalanced. I haven't found a better one yet.
>
> > Not sure I see a better solution either, and we are running short of time
> > before 19 RC1.
>
> While it's a good idea to wake the leader up instead of polling, I'm a
> bit concerned that we call SetLatch() on all autovacuum workers
> participating in cost balancing, whereas we need it only in a narrow
> situation: when the worker is participating in cost balancing, using
> parallel vacuum, and waiting for its parallel workers to finish.
> Calling SetLatch() on a process that is not waiting doesn't send a
> signal, butit still leaves the latch set, causing a spurious wakeup
> the next time the process waits for something else.
>
> I think a condition variable fits better here.

Thanks all for the comments.

I understand the cost of setting the latch and waking up every leader
participating in rebalancing could be higher. But we only need to wake
the leaders that are in the wait-for-parallel-workers-to-finish loop.
Those are the ones that need the rebalanced cost limit, so they can
propagate it down to their workers still doing index vacuuming. That
way the workers can adjust to the system needs. The user may have
changed the cost limits to slow things down or speed them up, and
either way it is not good for the workers to miss them.

Now, on the cost of setting the latch. It depends on three things.
First, how many leaders can exist, which depends on
autovacuum_max_workers and autovacuum_worker_slots. Second, how often
the number of leaders sharing the cost limit changes, that is, when
the recount comes out different from the previous count. This happens
when a leader starts or exits, or when the user sets or removes a
per-table cost limit, which moves a leader in or out of the sharing
set the next time that table is vacuumed. Third, how often a leader is
waiting in the loop at that moment.

Having all three happen together looks fairly rare in practice IMHO.
And even if it happens often, the cost is small, since set-latch sends
a signal only to a process that is in a latch wait at that moment. For
a leader not in the loop, it only marks the latch as set and returns,
and that leader's next latch wait just returns once early and
re-checks.

If we still want to set the latch only for the leaders in the loop,
there are a couple of ways, and I don't prefer either. We could check
the leader's wait event info to see whether it is in the loop, or the
leader could set a flag in its WorkerInfo when it enters the loop and
clear it when it leaves. Both are prone to race conditions and add
more code than they save. I would leave it as a follow-up if anyone
still wants to pursue it.

That said, please find the attached v3 patch set. 0001 has only the
fixes, and 0002 has the tests posted upthread along with the injection
points. I think 0001 is ready to go in, unless there are further
comments. I'm fine leaving 0002 here for the record, since I haven't
reviewed it in depth and I don't think it has had a close review yet,
and it needs more dev cycles. Anyone with cycles can pick it up for
HEAD. Given the time left, my suggestion would be to consider 0001 for
PG19 and HEAD.

Thoughts?

--
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com

Attachments:

  [application/x-patch] v3-0001-Refresh-autovacuum-costs-while-waiting-for-parall.patch (7.3K, ../../CALj2ACUBkuGeS8m+o_5OzrTOtf=NhekVJ8PLS5yNTyHYnH0jpg@mail.gmail.com/2-v3-0001-Refresh-autovacuum-costs-while-waiting-for-parall.patch)
  download | inline diff:
From 372d492c439b57c1d79c20c4042a668a5350d99b Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Fri, 25 Sep 2026 18:52:57 +0000
Subject: [PATCH v3 1/2] Refresh autovacuum costs while waiting for parallel
 workers.

Previously, an autovacuum worker running a parallel vacuum
(leader) propagated changes to the cost-based delay parameters to
its parallel workers only from vacuum_delay_point(). Once the
leader had finished its own share of the indexes and was waiting
for the parallel workers in WaitForParallelWorkersToFinish(), it
no longer reached vacuum_delay_point(), so the parallel workers
did not receive any changes until the wait ended, which could
take as long as the largest index.

As a result, a configuration reload during the wait stayed
pending, since SIGHUP woke up the leader but the wait loop only
called CHECK_FOR_INTERRUPTS(), which does not process
ConfigReloadPending. Likewise, a change in the number of
autovacuum workers sharing the cost limit did not reach the
leader, since nothing signaled it, and its parallel workers kept
running at the old share of the limit.

This commit fixes the issue in two parts. First, the wait loop in
WaitForParallelWorkersToFinish() now refreshes the cost-based
delay parameters and propagates them to the parallel workers on
every wakeup, when the process is an autovacuum worker. Second,
autovac_recalculate_workers_for_balance() now sets the latch of
the autovacuum workers sharing the cost limit when the count
changes, so that a leader waiting for its parallel workers wakes
up and picks up the new count. A worker that is not waiting on
its latch is not signaled and picks up the new count on its next
nap as before.

Backpatch to 19, where parallel autovacuum was introduced.

Reported-by: Nikolay Samokhvalov <nik@postgres.ai>
Author: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Reviewed-by: Manu <manuelreyesbravo@gmail.com>
Reviewed-by: Zsolt Parragi <zsolt.parragi@percona.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Reviewed-by: Masahiko Sawada <sawada.mshk@gmail.com>
Discussion: https://postgr.es/m/CAM527d-GL%3DJp2EJXBSnVBGPK-4XEZwWof5Cv8P0hghS_og6oAg%40mail.gmail.com
Backpatch-through: 19
---
 src/backend/access/transam/parallel.c | 11 +++++++
 src/backend/commands/vacuumparallel.c | 42 +++++++++++++++++++++++++++
 src/backend/postmaster/autovacuum.c   | 27 +++++++++++++++++
 src/include/commands/vacuum.h         |  1 +
 4 files changed, 81 insertions(+)

diff --git a/src/backend/access/transam/parallel.c b/src/backend/access/transam/parallel.c
index e1806a9a28a..b51df2a4606 100644
--- a/src/backend/access/transam/parallel.c
+++ b/src/backend/access/transam/parallel.c
@@ -812,6 +812,17 @@ WaitForParallelWorkersToFinish(ParallelContext *pcxt)
 		 */
 		CHECK_FOR_INTERRUPTS();
 
+		/*
+		 * An autovacuum worker running a parallel vacuum (leader) propagates
+		 * changes to the cost-based delay parameters to its parallel workers
+		 * at its own cost delay points, which it no longer reaches while
+		 * waiting here. Do it here instead, on every wakeup, so that a config
+		 * reload or a change in the number of autovacuum workers sharing the
+		 * cost limit reaches the parallel workers before they finish.
+		 */
+		if (AmAutoVacuumWorkerProcess())
+			parallel_vacuum_refresh_cost_params();
+
 		for (i = 0; i < pcxt->nworkers_launched; ++i)
 		{
 			/*
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 4532da60c84..b7de730a87a 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -43,6 +43,7 @@
 #include "executor/instrument.h"
 #include "optimizer/paths.h"
 #include "pgstat.h"
+#include "postmaster/interrupt.h"
 #include "storage/bufmgr.h"
 #include "storage/proc.h"
 #include "tcop/tcopprot.h"
@@ -725,6 +726,47 @@ parallel_vacuum_propagate_shared_delay_params(void)
 	pg_atomic_fetch_add_u32(&pv_shared_cost_params->generation, 1);
 }
 
+/*
+ * Refresh the cost-based vacuum delay parameters of an autovacuum worker
+ * running a parallel vacuum (leader) and propagate them to its parallel
+ * workers.
+ *
+ * The leader normally does this at its own cost delay points, which it no
+ * longer reaches once it is only waiting for the parallel workers to finish.
+ * It calls this from that wait instead, on every wakeup, to pick up a config
+ * reload and a change in the number of autovacuum workers sharing the cost
+ * limit. Both wake the leader, the second through the latch set by
+ * autovac_recalculate_workers_for_balance() when the count is recalculated.
+ */
+void
+parallel_vacuum_refresh_cost_params(void)
+{
+	Assert(AmAutoVacuumWorkerProcess());
+
+	/*
+	 * Quick return if the leader is not sharing the delay parameters with
+	 * parallel workers.
+	 */
+	if (pv_shared_cost_params == NULL)
+		return;
+
+	if (ConfigReloadPending)
+	{
+		ConfigReloadPending = false;
+		ProcessConfigFile(PGC_SIGHUP);
+
+		/* This re-divides the cost limit too */
+		VacuumUpdateCosts();
+	}
+	else
+	{
+		/* The number of workers sharing the cost limit may have changed */
+		AutoVacuumUpdateCostLimit();
+	}
+
+	parallel_vacuum_propagate_shared_delay_params();
+}
+
 /*
  * Compute the number of parallel worker processes to request.  Both index
  * vacuum and index cleanup can be executed with parallel workers.
diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c
index 60ebe828900..11da648e4a2 100644
--- a/src/backend/postmaster/autovacuum.c
+++ b/src/backend/postmaster/autovacuum.c
@@ -1814,8 +1814,35 @@ autovac_recalculate_workers_for_balance(void)
 	}
 
 	if (nworkers_for_balance != orig_nworkers_for_balance)
+	{
 		pg_atomic_write_u32(&AutoVacuumShmem->av_nworkersForBalance,
 							nworkers_for_balance);
+
+		/*
+		 * Wake up the autovacuum workers sharing the cost limit so that they
+		 * pick up the new count. An autovacuum worker that is vacuuming does
+		 * that on its next nap anyway, but one running a parallel vacuum
+		 * (leader) that is only waiting for its parallel workers to finish
+		 * never naps, and nothing else would tell it. The count is written
+		 * above, before the latches are set, so a woken leader always reads
+		 * the new value.
+		 *
+		 * Only the waiting leaders need this, but knowing which ones are
+		 * waiting would need more state. For an autovacuum worker that is not
+		 * in a latch wait, SetLatch() sends no signal and only marks the
+		 * latch set, which costs one early return from its next latch wait.
+		 */
+		dlist_foreach(iter, &AutoVacuumShmem->av_runningWorkers)
+		{
+			WorkerInfo	worker = dlist_container(WorkerInfoData, wi_links, iter.cur);
+
+			if (worker->wi_proc == NULL ||
+				pg_atomic_unlocked_test_flag(&worker->wi_dobalance))
+				continue;
+
+			SetLatch(&worker->wi_proc->procLatch);
+		}
+	}
 }
 
 /*
diff --git a/src/include/commands/vacuum.h b/src/include/commands/vacuum.h
index 6e3c912bf5c..89fa1fcd261 100644
--- a/src/include/commands/vacuum.h
+++ b/src/include/commands/vacuum.h
@@ -432,6 +432,7 @@ extern void parallel_vacuum_cleanup_all_indexes(ParallelVacuumState *pvs,
 												PVWorkerStats *wstats);
 extern void parallel_vacuum_update_shared_delay_params(void);
 extern void parallel_vacuum_propagate_shared_delay_params(void);
+extern void parallel_vacuum_refresh_cost_params(void);
 extern void parallel_vacuum_main(dsm_segment *seg, shm_toc *toc);
 
 /* in commands/analyze.c */
-- 
2.47.3



  [application/x-patch] v3-0002-Add-tests-for-autovacuum-cost-parameter-refresh-w.patch (13.1K, ../../CALj2ACUBkuGeS8m+o_5OzrTOtf=NhekVJ8PLS5yNTyHYnH0jpg@mail.gmail.com/3-v3-0002-Add-tests-for-autovacuum-cost-parameter-refresh-w.patch)
  download | inline diff:
From 6eb82c30ff618c6819c8e706f1d4def69e0d790c Mon Sep 17 00:00:00 2001
From: Nikolay Samokhvalov <nik@postgres.ai>
Date: Fri, 25 Sep 2026 19:01:30 +0000
Subject: [PATCH v3 2/2] Add tests for autovacuum cost parameter refresh while
 waiting.

---
 src/backend/commands/vacuumparallel.c         |  15 ++
 .../t/001_parallel_autovacuum.pl              | 166 ++++++++++++++++++
 .../t/003_cost_reload_while_waiting.pl        | 109 ++++++++++++
 3 files changed, 290 insertions(+)
 create mode 100644 src/test/modules/test_autovacuum/t/003_cost_reload_while_waiting.pl

diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index b7de730a87a..d29031dbd65 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -41,12 +41,14 @@
 #include "commands/progress.h"
 #include "commands/vacuum.h"
 #include "executor/instrument.h"
+#include "miscadmin.h"
 #include "optimizer/paths.h"
 #include "pgstat.h"
 #include "postmaster/interrupt.h"
 #include "storage/bufmgr.h"
 #include "storage/proc.h"
 #include "tcop/tcopprot.h"
+#include "utils/injection_point.h"
 #include "utils/lsyscache.h"
 #include "utils/rel.h"
 
@@ -765,6 +767,8 @@ parallel_vacuum_refresh_cost_params(void)
 	}
 
 	parallel_vacuum_propagate_shared_delay_params();
+
+	INJECTION_POINT("parallel-autovacuum-leader-cost-updated", NULL);
 }
 
 /*
@@ -971,6 +975,9 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 	/* Vacuum the indexes that can be processed by only leader process */
 	parallel_vacuum_process_unsafe_indexes(pvs);
 
+	if (pvs->shared->is_autovacuum)
+		INJECTION_POINT("parallel-autovacuum-leader-before-index", NULL);
+
 	/*
 	 * Join as a parallel worker.  The leader vacuums alone processes all
 	 * parallel-safe indexes in the case where no workers are launched.
@@ -986,6 +993,11 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 		/* Wait for all vacuum workers to finish */
 		WaitForParallelWorkersToFinish(pvs->pcxt);
 
+		if (pvs->shared->is_autovacuum)
+			INJECTION_POINT("parallel-autovacuum-leader-after-worker-wait",
+							ConfigReloadPending ? "reload pending" :
+							"reload processed");
+
 		for (int i = 0; i < pvs->pcxt->nworkers_launched; i++)
 			InstrAccumParallelQuery(&pvs->buffer_usage[i], &pvs->wal_usage[i]);
 	}
@@ -1052,6 +1064,9 @@ parallel_vacuum_process_safe_indexes(ParallelVacuumState *pvs)
 		if (!indstats->parallel_workers_can_process)
 			continue;
 
+		if (IsParallelWorker())
+			INJECTION_POINT("parallel-autovacuum-worker-before-index", NULL);
+
 		/* Do vacuum or cleanup of the index */
 		parallel_vacuum_process_one_index(pvs, pvs->indrels[idx], indstats);
 	}
diff --git a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
index 33c86bbdc94..bd1d574a4cf 100644
--- a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
+++ b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
@@ -252,6 +252,172 @@ $node->safe_psql('postgres',
 	"SELECT injection_points_wakeup('autovacuum-worker-cost-balanced')");
 $node->safe_psql('postgres',
 	"SELECT injection_points_detach('autovacuum-worker-cost-balanced')");
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) = 0 FROM pg_stat_activity
+	WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
+}) or die "second autovacuum worker did not finish";
+
+# Test 4:
+# Check whether a config reload is serviced while the autovacuum leader waits
+# for a parallel worker to finish an index.  Hold the worker after it claims
+# an index, so the leader can process the remaining indexes and enter
+# ParallelFinish.
+my $postgresoid = $node->safe_psql('postgres',
+	"SELECT oid FROM pg_database WHERE datname = 'postgres'");
+my $testautovacid =
+  $node->safe_psql('postgres', "SELECT 'test_autovac'::regclass::oid");
+
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_max_workers = 1;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 700;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_delay = 0;
+	SELECT pg_reload_conf();
+});
+
+prepare_for_next_test($node, 4);
+$log_offset = -s $node->logfile;
+
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_attach('parallel-autovacuum-worker-before-index', 'wait');
+	SELECT injection_points_attach('parallel-autovacuum-leader-before-index', 'wait');
+});
+$node->safe_psql('postgres',
+	'ALTER TABLE test_autovac SET (autovacuum_enabled = true)');
+$node->wait_for_event('autovacuum worker',
+	'parallel-autovacuum-leader-before-index');
+$node->wait_for_event('parallel worker',
+	'parallel-autovacuum-worker-before-index');
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_wakeup('parallel-autovacuum-leader-before-index');
+	SELECT injection_points_detach('parallel-autovacuum-leader-before-index');
+});
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) > 0
+	FROM pg_stat_activity worker
+	JOIN pg_stat_activity leader ON leader.pid = worker.leader_pid
+	WHERE worker.backend_type = 'parallel worker'
+	  AND worker.wait_event = 'parallel-autovacuum-worker-before-index'
+	  AND leader.wait_event = 'ParallelFinish'
+}) or die "autovacuum leader did not enter ParallelFinish";
+
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 800;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_delay = 8;
+	ALTER SYSTEM SET vacuum_cost_page_miss = 11;
+	ALTER SYSTEM SET vacuum_cost_page_dirty = 12;
+	ALTER SYSTEM SET vacuum_cost_page_hit = 13;
+	SELECT pg_reload_conf();
+});
+
+$node->wait_for_log(
+	qr/Autovacuum VacuumUpdateCosts\(db=$postgresoid, rel=$testautovacid, dobalance=yes, cost_limit=800, cost_delay=8 /,
+	$log_offset);
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('parallel-autovacuum-worker-before-index')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('parallel-autovacuum-worker-before-index')");
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=800, cost_delay=8, cost_page_miss=11, cost_page_dirty=12, cost_page_hit=13/,
+	$log_offset);
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+ok(1, "config reload is propagated while the leader waits for workers");
+
+# Test 5:
+# Check the same wait path for an un-signalled cost-limit rebalance.  A second
+# autovacuum worker joins the balance while the first leader and its parallel
+# worker remain held.
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_max_workers = 2;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 600;
+	SELECT pg_reload_conf();
+});
+
+prepare_for_next_test($node, 5);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = false)');
+$node->safe_psql('regress_db2', 'UPDATE filler SET id = id + 1');
+
+$log_offset = -s $node->logfile;
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_attach('parallel-autovacuum-worker-before-index', 'wait');
+	SELECT injection_points_attach('parallel-autovacuum-leader-before-index', 'wait');
+});
+$node->safe_psql('postgres',
+	'ALTER TABLE test_autovac SET (autovacuum_enabled = true)');
+$node->wait_for_event('autovacuum worker',
+	'parallel-autovacuum-leader-before-index');
+$node->wait_for_event('parallel worker',
+	'parallel-autovacuum-worker-before-index');
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_wakeup('parallel-autovacuum-leader-before-index');
+	SELECT injection_points_detach('parallel-autovacuum-leader-before-index');
+});
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) > 0
+	FROM pg_stat_activity worker
+	JOIN pg_stat_activity leader ON leader.pid = worker.leader_pid
+	WHERE worker.backend_type = 'parallel worker'
+	  AND worker.wait_event = 'parallel-autovacuum-worker-before-index'
+	  AND leader.wait_event = 'ParallelFinish'
+}) or die "autovacuum leader did not enter ParallelFinish";
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_attach('autovacuum-worker-cost-balanced', 'wait')"
+);
+
+# Attach before the rebalance, as the rebalance wakeup is the only thing that
+# brings the waiting leader to this point.
+$node->safe_psql('postgres',
+	"SELECT injection_points_attach('parallel-autovacuum-leader-cost-updated', 'notice')"
+);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = true)');
+$node->wait_for_log(
+	qr/VacuumUpdateCosts\(db=$db2oid, rel=$filleroid, dobalance=yes, cost_limit=300,/,
+	$log_offset);
+$node->wait_for_log(
+	qr/notice triggered for injection point parallel-autovacuum-leader-cost-updated/,
+	$log_offset);
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('parallel-autovacuum-leader-cost-updated')");
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('parallel-autovacuum-worker-before-index')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('parallel-autovacuum-worker-before-index')");
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=300,/,
+	$log_offset);
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('autovacuum-worker-cost-balanced')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('autovacuum-worker-cost-balanced')");
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) = 0 FROM pg_stat_activity
+	WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
+}) or die "second autovacuum worker did not finish";
+ok(1, "cost rebalance is propagated while the leader waits for workers");
 
 $node->stop;
 done_testing();
diff --git a/src/test/modules/test_autovacuum/t/003_cost_reload_while_waiting.pl b/src/test/modules/test_autovacuum/t/003_cost_reload_while_waiting.pl
new file mode 100644
index 00000000000..afc3a733670
--- /dev/null
+++ b/src/test/modules/test_autovacuum/t/003_cost_reload_while_waiting.pl
@@ -0,0 +1,109 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+use strict;
+use warnings FATAL => 'all';
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+	plan skip_all => 'Injection points not supported by this build';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+$node->append_conf(
+	'postgresql.conf', qq{
+autovacuum_max_workers = 1
+autovacuum_max_parallel_workers = 1
+autovacuum_naptime = '1s'
+autovacuum_vacuum_cost_delay = '20ms'
+autovacuum_vacuum_cost_limit = 200
+log_min_messages = debug2
+min_parallel_index_scan_size = 0
+});
+$node->start;
+
+if (!$node->check_extension('injection_points'))
+{
+	plan skip_all => 'Extension injection_points not installed';
+}
+
+$node->safe_psql(
+	'postgres', q{
+	CREATE EXTENSION injection_points;
+	CREATE TABLE test_autovac (id int, a int)
+		WITH (autovacuum_enabled = false,
+			autovacuum_parallel_workers = 1,
+			autovacuum_vacuum_threshold = 0,
+			autovacuum_vacuum_scale_factor = 0);
+	INSERT INTO test_autovac
+		SELECT g, g FROM generate_series(1, 100) g;
+	CREATE INDEX test_autovac_id_idx ON test_autovac (id);
+	CREATE INDEX test_autovac_a_idx ON test_autovac (a);
+	UPDATE test_autovac SET a = a + 1;
+	SELECT injection_points_attach(
+		'parallel-autovacuum-worker-before-index', 'wait');
+	SELECT injection_points_attach(
+		'parallel-autovacuum-leader-before-index', 'wait');
+	SELECT injection_points_attach(
+		'parallel-autovacuum-leader-after-worker-wait', 'notice');
+	ALTER TABLE test_autovac SET (autovacuum_enabled = true);
+});
+
+$node->wait_for_event('autovacuum worker',
+	'parallel-autovacuum-leader-before-index');
+$node->wait_for_event('parallel worker',
+	'parallel-autovacuum-worker-before-index');
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_wakeup(
+		'parallel-autovacuum-leader-before-index');
+	SELECT injection_points_detach(
+		'parallel-autovacuum-leader-before-index');
+});
+$node->poll_query_until(
+	'postgres', q{
+	SELECT EXISTS (
+		SELECT 1
+		FROM pg_stat_activity
+		WHERE backend_type = 'autovacuum worker'
+		AND wait_event = 'ParallelFinish')
+}) or die "autovacuum leader did not reach ParallelFinish";
+
+my $log_offset = -s $node->logfile;
+if (!$ENV{NO_RELOAD_CONTROL})
+{
+	$node->safe_psql(
+		'postgres', q{
+		ALTER SYSTEM SET autovacuum_vacuum_cost_delay = 0;
+		SELECT pg_reload_conf();
+	});
+}
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_wakeup(
+		'parallel-autovacuum-worker-before-index');
+	SELECT injection_points_detach(
+		'parallel-autovacuum-worker-before-index');
+});
+
+$node->wait_for_log(
+	qr/parallel-autovacuum-leader-after-worker-wait \(reload (?:pending|processed)\)/,
+	$log_offset);
+
+my $log = slurp_file($node->logfile, $log_offset);
+my ($reload_state) =
+	$log =~ /parallel-autovacuum-leader-after-worker-wait \(reload (pending|processed)\)/;
+is($reload_state, 'processed',
+	'autovacuum leader processes a configuration reload while waiting');
+
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_detach(
+		'parallel-autovacuum-leader-after-worker-wait');
+});
+$node->stop;
+
+done_testing();
-- 
2.47.3



^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-25 22:03  Manu <manuelreyesbravo@gmail.com>
  parent: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
  1 sibling, 0 replies; 43+ messages in thread

From: Manu @ 2026-09-25 22:03 UTC (permalink / raw)
  To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: Masahiko Sawada <sawada.mshk@gmail.com>; Daniel Gustafsson <daniel@yesql.se>; Nikolay Samokhvalov <nik@postgres.ai>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

Hi,

Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com> wrote:
> I'm fine leaving 0002 here for the record, since I haven't
> reviewed it in depth and I don't think it has had a close review yet,

I ran v3 0001+0002 on REL_19_STABLE (407687a0fb1), with assertions and
injection points.  001_parallel_autovacuum.pl and
003_cost_reload_while_waiting.pl passed 10 runs out of 10.

As a control I removed only the SetLatch() loop that 0001 adds to
autovac_recalculate_workers_for_balance().  Then test 5 in 001 times
out waiting for parallel-autovacuum-leader-cost-updated, while 003
still passes.  So 003 covers the reload path and test 5 the rebalance
path, each on its own.

The opposite case, a worker leaving the balance while the leader
waits, is not covered.  I tried a test for it and could not make it
deterministic: with autovacuum_naptime = 1s the launcher starts other
workers, and one of them joining the balance wakes the leader too, so
the test passed even with the launcher's wakeup removed.

Regards,
Manu






^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-26 00:31  Masahiko Sawada <sawada.mshk@gmail.com>
  parent: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
  1 sibling, 1 reply; 43+ messages in thread

From: Masahiko Sawada @ 2026-09-26 00:31 UTC (permalink / raw)
  To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; Nikolay Samokhvalov <nik@postgres.ai>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

On Fri, Sep 25, 2026 at 2:41 PM Bharath Rupireddy
<bharath.rupireddyforpostgres@gmail.com> wrote:
>
> Hi,
>
> On Thu, Sep 24, 2026 at 4:47 PM Masahiko Sawada <sawada.mshk@gmail.com> wrote:
> >
> > > > How about doing the check inside the existing wait loop, only when the process is an autovacuum worker, along the lines of the attached WIP? This is simple, the pattern already exists elsewhere in the code, and it looks safe. I also believe the wait loop is not in a performance-critical hot path, and this check is not costly. I checked that it still fixes the reported issues.
> >
> > > It's not great to sprinkle in worker specific code in the generic parallel
> > > handling.  My original thinking was to reject the idea, but the vacuum costing
> > > is already used for non-vacuum purposes (and there have been discussions to
> > > rename and generalize it) so with that in mind I am less concerned for this
> > > particular case.
> >
> > Same here. After thinking on the Bharath's patch, I now think it might
> > make sense. I guess it's likely that when autovacuum wants to wait for
> > other workers to finish in other cases like parallel heap vacuum, we
> > would want to use WaitForParallelWorkersToFinish() and want it to
> > update the cost-based delay params during the wait.
> >
> > > > What I am less sure about is the fix for the second issue in the attached WIP patch, which needs any leader waiting for its workers to be woken up after the cost limit is rebalanced. I haven't found a better one yet.
> >
> > > Not sure I see a better solution either, and we are running short of time
> > > before 19 RC1.
> >
> > While it's a good idea to wake the leader up instead of polling, I'm a
> > bit concerned that we call SetLatch() on all autovacuum workers
> > participating in cost balancing, whereas we need it only in a narrow
> > situation: when the worker is participating in cost balancing, using
> > parallel vacuum, and waiting for its parallel workers to finish.
> > Calling SetLatch() on a process that is not waiting doesn't send a
> > signal, butit still leaves the latch set, causing a spurious wakeup
> > the next time the process waits for something else.
> >
> > I think a condition variable fits better here.
>
> Thanks all for the comments.
>
> I understand the cost of setting the latch and waking up every leader
> participating in rebalancing could be higher. But we only need to wake
> the leaders that are in the wait-for-parallel-workers-to-finish loop.
> Those are the ones that need the rebalanced cost limit, so they can
> propagate it down to their workers still doing index vacuuming. That
> way the workers can adjust to the system needs. The user may have
> changed the cost limits to slow things down or speed them up, and
> either way it is not good for the workers to miss them.
>
> Now, on the cost of setting the latch. It depends on three things.
> First, how many leaders can exist, which depends on
> autovacuum_max_workers and autovacuum_worker_slots. Second, how often
> the number of leaders sharing the cost limit changes, that is, when
> the recount comes out different from the previous count. This happens
> when a leader starts or exits, or when the user sets or removes a
> per-table cost limit, which moves a leader in or out of the sharing
> set the next time that table is vacuumed. Third, how often a leader is
> waiting in the loop at that moment.
>
> Having all three happen together looks fairly rare in practice IMHO.
> And even if it happens often, the cost is small, since set-latch sends
> a signal only to a process that is in a latch wait at that moment. For
> a leader not in the loop, it only marks the latch as set and returns,
> and that leader's next latch wait just returns once early and
> re-checks.

I think the cost argument should be the other way around. Setting the
latch of a leader that is waiting for its parallel workers is exactly
what we want, and using a cv wouldn't change how often that happens.
What I was concerned about is the case where SetLatch() is wasted,
i.e., when the worker is not waiting in that loop. Since the wakeups
we actually need are not frequent, most SetLatch() calls would be
wasted, but a wasted SetLatch() is cheap as you and Mau mentioned. It
only marks the latch set if the worker is not waiting, and at worst
causes one spurious wakeup if the worker is waiting on something else.
And it happens only when the number of workers sharing the cost limit
changes. So I agree that it's acceptable, and given we're close to
RC1, I'm fine with the SetLatch() approach.

> That said, please find the attached v3 patch set. 0001 has only the
> fixes, and 0002 has the tests posted upthread along with the injection
> points. I think 0001 is ready to go in, unless there are further
> comments. I'm fine leaving 0002 here for the record, since I haven't
> reviewed it in depth and I don't think it has had a close review yet,
> and it needs more dev cycles. Anyone with cycles can pick it up for
> HEAD. Given the time left, my suggestion would be to consider 0001 for
> PG19 and HEAD.
>
> Thoughts?

I'll review the 0002 patch in depth and see if it's reasonable to push
some parts of it along with the fix.

As for the 0001 patch, it looks good to me. I have one minor comment,
though it's a matter of personal preference:

+/*
+ * Refresh the cost-based vacuum delay parameters of an autovacuum worker
+ * running a parallel vacuum (leader) and propagate them to its parallel
+ * workers.
+ *
+ * The leader normally does this at its own cost delay points, which it no
+ * longer reaches once it is only waiting for the parallel workers to finish.
+ * It calls this from that wait instead, on every wakeup, to pick up a config
+ * reload and a change in the number of autovacuum workers sharing the cost
+ * limit. Both wake the leader, the second through the latch set by
+ * autovac_recalculate_workers_for_balance() when the count is recalculated.
+ */
+void
+parallel_vacuum_refresh_cost_params(void)

The second paragraph explains how this function is called. Since what
the function does, as well as its name, is not specific to that wait,
I guess it's better to explain it on the caller side, and the comment
newly added in WaitForParallelWorkersToFinish() already explains a
similar thing. Then the function comment can focus on what the
function does and its side effects, such as reloading the
configuration file. For example:

  /*
   * Reload the configuration file if requested, and refresh the cost-based
   * delay parameters of an autovacuum worker running a parallel vacuum
   * (leader), propagating any change to its parallel workers.
   */

Regards,

-- 
Masahiko Sawada
Amazon Web Services: https://aws.amazon.com






^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-26 04:37  Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
  parent: Masahiko Sawada <sawada.mshk@gmail.com>
  0 siblings, 1 reply; 43+ messages in thread

From: Bharath Rupireddy @ 2026-09-26 04:37 UTC (permalink / raw)
  To: Masahiko Sawada <sawada.mshk@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; Nikolay Samokhvalov <nik@postgres.ai>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

Hi,

On Fri, Sep 25, 2026 at 5:32 PM Masahiko Sawada <sawada.mshk@gmail.com> wrote:
>
> And it happens only when the number of workers sharing the cost limit
> changes. So I agree that it's acceptable, and given we're close to
> RC1, I'm fine with the SetLatch() approach.
>
> As for the 0001 patch, it looks good to me. I have one minor comment,
> though it's a matter of personal preference:

Thanks for reviewing it.

> +void
> +parallel_vacuum_refresh_cost_params(void)
>
> The second paragraph explains how this function is called. Since what
> the function does, as well as its name, is not specific to that wait,
> I guess it's better to explain it on the caller side, and the comment
> newly added in WaitForParallelWorkersToFinish() already explains a
> similar thing. Then the function comment can focus on what the
> function does and its side effects, such as reloading the
> configuration file. For example:
>
>   /*
>    * Reload the configuration file if requested, and refresh the cost-based
>    * delay parameters of an autovacuum worker running a parallel vacuum
>    * (leader), propagating any change to its parallel workers.
>    */

Agreed.

Please find the attached v4 patches. 0002 remains the same.

--
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com

Attachments:

  [application/octet-stream] v4-0001-Refresh-autovacuum-costs-while-waiting-for-parall.patch (6.9K, ../../CALj2ACVnxudCDbi1DaxpkR=jLC4Ff01msAgL_w8W=P2EOeN3Aw@mail.gmail.com/2-v4-0001-Refresh-autovacuum-costs-while-waiting-for-parall.patch)
  download | inline diff:
From 4e6fefb064323566826dcb483b38357607d0972a Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Fri, 25 Sep 2026 18:52:57 +0000
Subject: [PATCH v4 1/2] Refresh autovacuum costs while waiting for parallel
 workers.

Previously, an autovacuum worker running a parallel vacuum
(leader) propagated changes to the cost-based delay parameters to
its parallel workers only from vacuum_delay_point(). Once the
leader had finished its own share of the indexes and was waiting
for the parallel workers in WaitForParallelWorkersToFinish(), it
no longer reached vacuum_delay_point(), so the parallel workers
did not receive any changes until the wait ended, which could
take as long as the largest index.

As a result, a configuration reload during the wait stayed
pending, since SIGHUP woke up the leader but the wait loop only
called CHECK_FOR_INTERRUPTS(), which does not process
ConfigReloadPending. Likewise, a change in the number of
autovacuum workers sharing the cost limit did not reach the
leader, since nothing signaled it, and its parallel workers kept
running at the old share of the limit.

This commit fixes the issue in two parts. First, the wait loop in
WaitForParallelWorkersToFinish() now refreshes the cost-based
delay parameters and propagates them to the parallel workers on
every wakeup, when the process is an autovacuum worker. Second,
autovac_recalculate_workers_for_balance() now sets the latch of
the autovacuum workers sharing the cost limit when the count
changes, so that a leader waiting for its parallel workers wakes
up and picks up the new count. A worker that is not waiting on
its latch is not signaled and picks up the new count on its next
nap as before.

Backpatch to 19, where parallel autovacuum was introduced.

Reported-by: Nikolay Samokhvalov <nik@postgres.ai>
Author: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Reviewed-by: Manu <manuelreyesbravo@gmail.com>
Reviewed-by: Zsolt Parragi <zsolt.parragi@percona.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Reviewed-by: Masahiko Sawada <sawada.mshk@gmail.com>
Discussion: https://postgr.es/m/CAM527d-GL%3DJp2EJXBSnVBGPK-4XEZwWof5Cv8P0hghS_og6oAg%40mail.gmail.com
Backpatch-through: 19
---
 src/backend/access/transam/parallel.c | 11 +++++++++
 src/backend/commands/vacuumparallel.c | 35 +++++++++++++++++++++++++++
 src/backend/postmaster/autovacuum.c   | 27 +++++++++++++++++++++
 src/include/commands/vacuum.h         |  1 +
 4 files changed, 74 insertions(+)

diff --git a/src/backend/access/transam/parallel.c b/src/backend/access/transam/parallel.c
index e1806a9a28a..b51df2a4606 100644
--- a/src/backend/access/transam/parallel.c
+++ b/src/backend/access/transam/parallel.c
@@ -812,6 +812,17 @@ WaitForParallelWorkersToFinish(ParallelContext *pcxt)
 		 */
 		CHECK_FOR_INTERRUPTS();
 
+		/*
+		 * An autovacuum worker running a parallel vacuum (leader) propagates
+		 * changes to the cost-based delay parameters to its parallel workers
+		 * at its own cost delay points, which it no longer reaches while
+		 * waiting here. Do it here instead, on every wakeup, so that a config
+		 * reload or a change in the number of autovacuum workers sharing the
+		 * cost limit reaches the parallel workers before they finish.
+		 */
+		if (AmAutoVacuumWorkerProcess())
+			parallel_vacuum_refresh_cost_params();
+
 		for (i = 0; i < pcxt->nworkers_launched; ++i)
 		{
 			/*
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 4532da60c84..2dd127e9118 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -43,6 +43,7 @@
 #include "executor/instrument.h"
 #include "optimizer/paths.h"
 #include "pgstat.h"
+#include "postmaster/interrupt.h"
 #include "storage/bufmgr.h"
 #include "storage/proc.h"
 #include "tcop/tcopprot.h"
@@ -725,6 +726,40 @@ parallel_vacuum_propagate_shared_delay_params(void)
 	pg_atomic_fetch_add_u32(&pv_shared_cost_params->generation, 1);
 }
 
+/*
+ * Reload the configuration file if requested, and refresh the cost-based
+ * delay parameters of an autovacuum worker running a parallel vacuum
+ * (leader), propagating any change to its parallel workers.
+ */
+void
+parallel_vacuum_refresh_cost_params(void)
+{
+	Assert(AmAutoVacuumWorkerProcess());
+
+	/*
+	 * Quick return if the leader is not sharing the delay parameters with
+	 * parallel workers.
+	 */
+	if (pv_shared_cost_params == NULL)
+		return;
+
+	if (ConfigReloadPending)
+	{
+		ConfigReloadPending = false;
+		ProcessConfigFile(PGC_SIGHUP);
+
+		/* This re-divides the cost limit too */
+		VacuumUpdateCosts();
+	}
+	else
+	{
+		/* The number of workers sharing the cost limit may have changed */
+		AutoVacuumUpdateCostLimit();
+	}
+
+	parallel_vacuum_propagate_shared_delay_params();
+}
+
 /*
  * Compute the number of parallel worker processes to request.  Both index
  * vacuum and index cleanup can be executed with parallel workers.
diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c
index 60ebe828900..11da648e4a2 100644
--- a/src/backend/postmaster/autovacuum.c
+++ b/src/backend/postmaster/autovacuum.c
@@ -1814,8 +1814,35 @@ autovac_recalculate_workers_for_balance(void)
 	}
 
 	if (nworkers_for_balance != orig_nworkers_for_balance)
+	{
 		pg_atomic_write_u32(&AutoVacuumShmem->av_nworkersForBalance,
 							nworkers_for_balance);
+
+		/*
+		 * Wake up the autovacuum workers sharing the cost limit so that they
+		 * pick up the new count. An autovacuum worker that is vacuuming does
+		 * that on its next nap anyway, but one running a parallel vacuum
+		 * (leader) that is only waiting for its parallel workers to finish
+		 * never naps, and nothing else would tell it. The count is written
+		 * above, before the latches are set, so a woken leader always reads
+		 * the new value.
+		 *
+		 * Only the waiting leaders need this, but knowing which ones are
+		 * waiting would need more state. For an autovacuum worker that is not
+		 * in a latch wait, SetLatch() sends no signal and only marks the
+		 * latch set, which costs one early return from its next latch wait.
+		 */
+		dlist_foreach(iter, &AutoVacuumShmem->av_runningWorkers)
+		{
+			WorkerInfo	worker = dlist_container(WorkerInfoData, wi_links, iter.cur);
+
+			if (worker->wi_proc == NULL ||
+				pg_atomic_unlocked_test_flag(&worker->wi_dobalance))
+				continue;
+
+			SetLatch(&worker->wi_proc->procLatch);
+		}
+	}
 }
 
 /*
diff --git a/src/include/commands/vacuum.h b/src/include/commands/vacuum.h
index 6e3c912bf5c..89fa1fcd261 100644
--- a/src/include/commands/vacuum.h
+++ b/src/include/commands/vacuum.h
@@ -432,6 +432,7 @@ extern void parallel_vacuum_cleanup_all_indexes(ParallelVacuumState *pvs,
 												PVWorkerStats *wstats);
 extern void parallel_vacuum_update_shared_delay_params(void);
 extern void parallel_vacuum_propagate_shared_delay_params(void);
+extern void parallel_vacuum_refresh_cost_params(void);
 extern void parallel_vacuum_main(dsm_segment *seg, shm_toc *toc);
 
 /* in commands/analyze.c */
-- 
2.47.3



  [application/octet-stream] v4-0002-Add-tests-for-autovacuum-cost-parameter-refresh-w.patch (13.1K, ../../CALj2ACVnxudCDbi1DaxpkR=jLC4Ff01msAgL_w8W=P2EOeN3Aw@mail.gmail.com/3-v4-0002-Add-tests-for-autovacuum-cost-parameter-refresh-w.patch)
  download | inline diff:
From 10f257503d4276e6b7a4468bdbf32ec91c2750c9 Mon Sep 17 00:00:00 2001
From: Nikolay Samokhvalov <nik@postgres.ai>
Date: Fri, 25 Sep 2026 19:01:30 +0000
Subject: [PATCH v4 2/2] Add tests for autovacuum cost parameter refresh while
 waiting.

---
 src/backend/commands/vacuumparallel.c         |  15 ++
 .../t/001_parallel_autovacuum.pl              | 166 ++++++++++++++++++
 .../t/003_cost_reload_while_waiting.pl        | 109 ++++++++++++
 3 files changed, 290 insertions(+)
 create mode 100644 src/test/modules/test_autovacuum/t/003_cost_reload_while_waiting.pl

diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 2dd127e9118..405f0b5e7b1 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -41,12 +41,14 @@
 #include "commands/progress.h"
 #include "commands/vacuum.h"
 #include "executor/instrument.h"
+#include "miscadmin.h"
 #include "optimizer/paths.h"
 #include "pgstat.h"
 #include "postmaster/interrupt.h"
 #include "storage/bufmgr.h"
 #include "storage/proc.h"
 #include "tcop/tcopprot.h"
+#include "utils/injection_point.h"
 #include "utils/lsyscache.h"
 #include "utils/rel.h"
 
@@ -758,6 +760,8 @@ parallel_vacuum_refresh_cost_params(void)
 	}
 
 	parallel_vacuum_propagate_shared_delay_params();
+
+	INJECTION_POINT("parallel-autovacuum-leader-cost-updated", NULL);
 }
 
 /*
@@ -964,6 +968,9 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 	/* Vacuum the indexes that can be processed by only leader process */
 	parallel_vacuum_process_unsafe_indexes(pvs);
 
+	if (pvs->shared->is_autovacuum)
+		INJECTION_POINT("parallel-autovacuum-leader-before-index", NULL);
+
 	/*
 	 * Join as a parallel worker.  The leader vacuums alone processes all
 	 * parallel-safe indexes in the case where no workers are launched.
@@ -979,6 +986,11 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 		/* Wait for all vacuum workers to finish */
 		WaitForParallelWorkersToFinish(pvs->pcxt);
 
+		if (pvs->shared->is_autovacuum)
+			INJECTION_POINT("parallel-autovacuum-leader-after-worker-wait",
+							ConfigReloadPending ? "reload pending" :
+							"reload processed");
+
 		for (int i = 0; i < pvs->pcxt->nworkers_launched; i++)
 			InstrAccumParallelQuery(&pvs->buffer_usage[i], &pvs->wal_usage[i]);
 	}
@@ -1045,6 +1057,9 @@ parallel_vacuum_process_safe_indexes(ParallelVacuumState *pvs)
 		if (!indstats->parallel_workers_can_process)
 			continue;
 
+		if (IsParallelWorker())
+			INJECTION_POINT("parallel-autovacuum-worker-before-index", NULL);
+
 		/* Do vacuum or cleanup of the index */
 		parallel_vacuum_process_one_index(pvs, pvs->indrels[idx], indstats);
 	}
diff --git a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
index 33c86bbdc94..bd1d574a4cf 100644
--- a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
+++ b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
@@ -252,6 +252,172 @@ $node->safe_psql('postgres',
 	"SELECT injection_points_wakeup('autovacuum-worker-cost-balanced')");
 $node->safe_psql('postgres',
 	"SELECT injection_points_detach('autovacuum-worker-cost-balanced')");
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) = 0 FROM pg_stat_activity
+	WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
+}) or die "second autovacuum worker did not finish";
+
+# Test 4:
+# Check whether a config reload is serviced while the autovacuum leader waits
+# for a parallel worker to finish an index.  Hold the worker after it claims
+# an index, so the leader can process the remaining indexes and enter
+# ParallelFinish.
+my $postgresoid = $node->safe_psql('postgres',
+	"SELECT oid FROM pg_database WHERE datname = 'postgres'");
+my $testautovacid =
+  $node->safe_psql('postgres', "SELECT 'test_autovac'::regclass::oid");
+
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_max_workers = 1;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 700;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_delay = 0;
+	SELECT pg_reload_conf();
+});
+
+prepare_for_next_test($node, 4);
+$log_offset = -s $node->logfile;
+
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_attach('parallel-autovacuum-worker-before-index', 'wait');
+	SELECT injection_points_attach('parallel-autovacuum-leader-before-index', 'wait');
+});
+$node->safe_psql('postgres',
+	'ALTER TABLE test_autovac SET (autovacuum_enabled = true)');
+$node->wait_for_event('autovacuum worker',
+	'parallel-autovacuum-leader-before-index');
+$node->wait_for_event('parallel worker',
+	'parallel-autovacuum-worker-before-index');
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_wakeup('parallel-autovacuum-leader-before-index');
+	SELECT injection_points_detach('parallel-autovacuum-leader-before-index');
+});
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) > 0
+	FROM pg_stat_activity worker
+	JOIN pg_stat_activity leader ON leader.pid = worker.leader_pid
+	WHERE worker.backend_type = 'parallel worker'
+	  AND worker.wait_event = 'parallel-autovacuum-worker-before-index'
+	  AND leader.wait_event = 'ParallelFinish'
+}) or die "autovacuum leader did not enter ParallelFinish";
+
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 800;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_delay = 8;
+	ALTER SYSTEM SET vacuum_cost_page_miss = 11;
+	ALTER SYSTEM SET vacuum_cost_page_dirty = 12;
+	ALTER SYSTEM SET vacuum_cost_page_hit = 13;
+	SELECT pg_reload_conf();
+});
+
+$node->wait_for_log(
+	qr/Autovacuum VacuumUpdateCosts\(db=$postgresoid, rel=$testautovacid, dobalance=yes, cost_limit=800, cost_delay=8 /,
+	$log_offset);
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('parallel-autovacuum-worker-before-index')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('parallel-autovacuum-worker-before-index')");
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=800, cost_delay=8, cost_page_miss=11, cost_page_dirty=12, cost_page_hit=13/,
+	$log_offset);
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+ok(1, "config reload is propagated while the leader waits for workers");
+
+# Test 5:
+# Check the same wait path for an un-signalled cost-limit rebalance.  A second
+# autovacuum worker joins the balance while the first leader and its parallel
+# worker remain held.
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_max_workers = 2;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 600;
+	SELECT pg_reload_conf();
+});
+
+prepare_for_next_test($node, 5);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = false)');
+$node->safe_psql('regress_db2', 'UPDATE filler SET id = id + 1');
+
+$log_offset = -s $node->logfile;
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_attach('parallel-autovacuum-worker-before-index', 'wait');
+	SELECT injection_points_attach('parallel-autovacuum-leader-before-index', 'wait');
+});
+$node->safe_psql('postgres',
+	'ALTER TABLE test_autovac SET (autovacuum_enabled = true)');
+$node->wait_for_event('autovacuum worker',
+	'parallel-autovacuum-leader-before-index');
+$node->wait_for_event('parallel worker',
+	'parallel-autovacuum-worker-before-index');
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_wakeup('parallel-autovacuum-leader-before-index');
+	SELECT injection_points_detach('parallel-autovacuum-leader-before-index');
+});
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) > 0
+	FROM pg_stat_activity worker
+	JOIN pg_stat_activity leader ON leader.pid = worker.leader_pid
+	WHERE worker.backend_type = 'parallel worker'
+	  AND worker.wait_event = 'parallel-autovacuum-worker-before-index'
+	  AND leader.wait_event = 'ParallelFinish'
+}) or die "autovacuum leader did not enter ParallelFinish";
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_attach('autovacuum-worker-cost-balanced', 'wait')"
+);
+
+# Attach before the rebalance, as the rebalance wakeup is the only thing that
+# brings the waiting leader to this point.
+$node->safe_psql('postgres',
+	"SELECT injection_points_attach('parallel-autovacuum-leader-cost-updated', 'notice')"
+);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = true)');
+$node->wait_for_log(
+	qr/VacuumUpdateCosts\(db=$db2oid, rel=$filleroid, dobalance=yes, cost_limit=300,/,
+	$log_offset);
+$node->wait_for_log(
+	qr/notice triggered for injection point parallel-autovacuum-leader-cost-updated/,
+	$log_offset);
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('parallel-autovacuum-leader-cost-updated')");
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('parallel-autovacuum-worker-before-index')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('parallel-autovacuum-worker-before-index')");
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=300,/,
+	$log_offset);
+
+$node->safe_psql('postgres',
+	"SELECT injection_points_wakeup('autovacuum-worker-cost-balanced')");
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('autovacuum-worker-cost-balanced')");
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+$node->poll_query_until(
+	'postgres', q{
+	SELECT count(*) = 0 FROM pg_stat_activity
+	WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
+}) or die "second autovacuum worker did not finish";
+ok(1, "cost rebalance is propagated while the leader waits for workers");
 
 $node->stop;
 done_testing();
diff --git a/src/test/modules/test_autovacuum/t/003_cost_reload_while_waiting.pl b/src/test/modules/test_autovacuum/t/003_cost_reload_while_waiting.pl
new file mode 100644
index 00000000000..afc3a733670
--- /dev/null
+++ b/src/test/modules/test_autovacuum/t/003_cost_reload_while_waiting.pl
@@ -0,0 +1,109 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+
+use strict;
+use warnings FATAL => 'all';
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+if ($ENV{enable_injection_points} ne 'yes')
+{
+	plan skip_all => 'Injection points not supported by this build';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('main');
+$node->init;
+$node->append_conf(
+	'postgresql.conf', qq{
+autovacuum_max_workers = 1
+autovacuum_max_parallel_workers = 1
+autovacuum_naptime = '1s'
+autovacuum_vacuum_cost_delay = '20ms'
+autovacuum_vacuum_cost_limit = 200
+log_min_messages = debug2
+min_parallel_index_scan_size = 0
+});
+$node->start;
+
+if (!$node->check_extension('injection_points'))
+{
+	plan skip_all => 'Extension injection_points not installed';
+}
+
+$node->safe_psql(
+	'postgres', q{
+	CREATE EXTENSION injection_points;
+	CREATE TABLE test_autovac (id int, a int)
+		WITH (autovacuum_enabled = false,
+			autovacuum_parallel_workers = 1,
+			autovacuum_vacuum_threshold = 0,
+			autovacuum_vacuum_scale_factor = 0);
+	INSERT INTO test_autovac
+		SELECT g, g FROM generate_series(1, 100) g;
+	CREATE INDEX test_autovac_id_idx ON test_autovac (id);
+	CREATE INDEX test_autovac_a_idx ON test_autovac (a);
+	UPDATE test_autovac SET a = a + 1;
+	SELECT injection_points_attach(
+		'parallel-autovacuum-worker-before-index', 'wait');
+	SELECT injection_points_attach(
+		'parallel-autovacuum-leader-before-index', 'wait');
+	SELECT injection_points_attach(
+		'parallel-autovacuum-leader-after-worker-wait', 'notice');
+	ALTER TABLE test_autovac SET (autovacuum_enabled = true);
+});
+
+$node->wait_for_event('autovacuum worker',
+	'parallel-autovacuum-leader-before-index');
+$node->wait_for_event('parallel worker',
+	'parallel-autovacuum-worker-before-index');
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_wakeup(
+		'parallel-autovacuum-leader-before-index');
+	SELECT injection_points_detach(
+		'parallel-autovacuum-leader-before-index');
+});
+$node->poll_query_until(
+	'postgres', q{
+	SELECT EXISTS (
+		SELECT 1
+		FROM pg_stat_activity
+		WHERE backend_type = 'autovacuum worker'
+		AND wait_event = 'ParallelFinish')
+}) or die "autovacuum leader did not reach ParallelFinish";
+
+my $log_offset = -s $node->logfile;
+if (!$ENV{NO_RELOAD_CONTROL})
+{
+	$node->safe_psql(
+		'postgres', q{
+		ALTER SYSTEM SET autovacuum_vacuum_cost_delay = 0;
+		SELECT pg_reload_conf();
+	});
+}
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_wakeup(
+		'parallel-autovacuum-worker-before-index');
+	SELECT injection_points_detach(
+		'parallel-autovacuum-worker-before-index');
+});
+
+$node->wait_for_log(
+	qr/parallel-autovacuum-leader-after-worker-wait \(reload (?:pending|processed)\)/,
+	$log_offset);
+
+my $log = slurp_file($node->logfile, $log_offset);
+my ($reload_state) =
+	$log =~ /parallel-autovacuum-leader-after-worker-wait \(reload (pending|processed)\)/;
+is($reload_state, 'processed',
+	'autovacuum leader processes a configuration reload while waiting');
+
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_detach(
+		'parallel-autovacuum-leader-after-worker-wait');
+});
+$node->stop;
+
+done_testing();
-- 
2.47.3



^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-28 13:13  Daniel Gustafsson <daniel@yesql.se>
  parent: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
  0 siblings, 1 reply; 43+ messages in thread

From: Daniel Gustafsson @ 2026-09-28 13:13 UTC (permalink / raw)
  To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: Masahiko Sawada <sawada.mshk@gmail.com>; Nikolay Samokhvalov <nik@postgres.ai>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

> On 26 Sep 2026, at 06:37, Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com> wrote:
> On Fri, Sep 25, 2026 at 5:32 PM Masahiko Sawada <sawada.mshk@gmail.com> wrote:

>> As for the 0001 patch, it looks good to me.

Agreed, +1 pn 0001.

> Please find the attached v4 patches. 0002 remains the same.

Regarding 0002, ISTM that it will be stable but we could also just commit to
master and hold off on REL_19_STABLE for now.  Once proven it can be back-
patched even post GA as it is test-only.  That could give us more time while
reducing the risk of unstable tests this very late in the cycle.  A few small
comments on 0002:
	
	+	if (pvs->shared->is_autovacuum)
	+		INJECTION_POINT("parallel-autovacuum-leader-before-index", NULL);

I think our common coding pattern is to wrap constructions like this (and many
more in the patch) in #ifdef USE_INJECTION_POINTS.

	+$node->poll_query_until(
	+	'postgres', q{
	+	SELECT count(*) = 0 FROM pg_stat_activity
	+	WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
	+}) or die "second autovacuum worker did not finish";

I'm not a fan of tests that die() when the condition fails rather than report a
test failure.


	+$node->wait_for_log(
	+	qr/parallel-autovacuum-leader-after-worker-wait \(reload (?:pending|processed)\)/,
	+	$log_offset);
	+
	+my $log = slurp_file($node->logfile, $log_offset);
	+my ($reload_state) =
	+	$log =~ /parallel-autovacuum-leader-after-worker-wait \(reload (pending|processed)\)/;
	+is($reload_state, 'processed',
	+	'autovacuum leader processes a configuration reload while waiting');

While not a problem with this patch, I wish wait_for_log could either return
the match, or at least the offset of the match, and not just the size of the
file.  If we had that we could avoid reading excessive amounts of log data and
reduce the risk of buggy tests matching on the wrong part of the log.

--
Daniel Gustafsson







^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-30 02:19  Masahiko Sawada <sawada.mshk@gmail.com>
  parent: Daniel Gustafsson <daniel@yesql.se>
  0 siblings, 2 replies; 43+ messages in thread

From: Masahiko Sawada @ 2026-09-30 02:19 UTC (permalink / raw)
  To: Daniel Gustafsson <daniel@yesql.se>; +Cc: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Nikolay Samokhvalov <nik@postgres.ai>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

On Mon, Sep 28, 2026 at 9:13 AM Daniel Gustafsson <daniel@yesql.se> wrote:
>
> > On 26 Sep 2026, at 06:37, Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com> wrote:
> > On Fri, Sep 25, 2026 at 5:32 PM Masahiko Sawada <sawada.mshk@gmail.com> wrote:
>
> >> As for the 0001 patch, it looks good to me.
>
> Agreed, +1 pn 0001.
>
> > Please find the attached v4 patches. 0002 remains the same.
>
> Regarding 0002, ISTM that it will be stable but we could also just commit to
> master and hold off on REL_19_STABLE for now.  Once proven it can be back-
> patched even post GA as it is test-only.  That could give us more time while
> reducing the risk of unstable tests this very late in the cycle.

Agreed.

Daniel, are you planning to push the 0001 patch separately from the
tests? I will be attending Postgres Summit US 2026 this week, so I can
push 0001 early next week instead just in case the buildfarm turns red
and requires monitoring.

> comments on 0002:

Aside from these comments, I think we could simplify the tests by
introducing a boolean variable, say, leader_participates that controls
whether the leader participates in parallel index vacuuming. It would
default to true, but we could toggle it off via an injection point.
When set to false, the leader would vacuum parallel-unsafe indexes but
skip parallel-safe ones.

Regards,

-- 
Masahiko Sawada
Amazon Web Services: https://aws.amazon.com






^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-30 09:00  Zsolt Parragi <zsolt.parragi@percona.com>
  parent: Masahiko Sawada <sawada.mshk@gmail.com>
  1 sibling, 1 reply; 43+ messages in thread

From: Zsolt Parragi @ 2026-09-30 09:00 UTC (permalink / raw)
  To: Masahiko Sawada <sawada.mshk@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Nikolay Samokhvalov <nik@postgres.ai>; pgsql-bugs@lists.postgresql.org

I addressed these in v5. 0001 is unchanged.

> I think we could simplify the tests by
> introducing a boolean variable, say, leader_participates
> ...

I am not convinced that this change made things simpler, but it is
included in 0002.

> While not a problem with this patch, I wish wait_for_log could either return
> the match, or at least the offset of the match, and not just the size of the
> file.  If we had that we could avoid reading excessive amounts of log data and
> reduce the risk of buggy tests matching on the wrong part of the log.

I ended up completely removing that test, it seemed redundant.

Attachments:

  [application/x-patch] v5-0001-Refresh-autovacuum-costs-while-waiting-for-parall.patch (6.9K, ../../CAN4CZFOGQCy+fxtQEXHUqwF6gTLpNSdMtV9CxHC_0yd-f=NhBw@mail.gmail.com/2-v5-0001-Refresh-autovacuum-costs-while-waiting-for-parall.patch)
  download | inline diff:
From 2af0bce739701970478e9c9a6838ec279e27b96c Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Fri, 25 Sep 2026 18:52:57 +0000
Subject: [PATCH v5 1/2] Refresh autovacuum costs while waiting for parallel
 workers.

Previously, an autovacuum worker running a parallel vacuum
(leader) propagated changes to the cost-based delay parameters to
its parallel workers only from vacuum_delay_point(). Once the
leader had finished its own share of the indexes and was waiting
for the parallel workers in WaitForParallelWorkersToFinish(), it
no longer reached vacuum_delay_point(), so the parallel workers
did not receive any changes until the wait ended, which could
take as long as the largest index.

As a result, a configuration reload during the wait stayed
pending, since SIGHUP woke up the leader but the wait loop only
called CHECK_FOR_INTERRUPTS(), which does not process
ConfigReloadPending. Likewise, a change in the number of
autovacuum workers sharing the cost limit did not reach the
leader, since nothing signaled it, and its parallel workers kept
running at the old share of the limit.

This commit fixes the issue in two parts. First, the wait loop in
WaitForParallelWorkersToFinish() now refreshes the cost-based
delay parameters and propagates them to the parallel workers on
every wakeup, when the process is an autovacuum worker. Second,
autovac_recalculate_workers_for_balance() now sets the latch of
the autovacuum workers sharing the cost limit when the count
changes, so that a leader waiting for its parallel workers wakes
up and picks up the new count. A worker that is not waiting on
its latch is not signaled and picks up the new count on its next
nap as before.

Backpatch to 19, where parallel autovacuum was introduced.

Reported-by: Nikolay Samokhvalov <nik@postgres.ai>
Author: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Reviewed-by: Manu <manuelreyesbravo@gmail.com>
Reviewed-by: Zsolt Parragi <zsolt.parragi@percona.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Reviewed-by: Masahiko Sawada <sawada.mshk@gmail.com>
Discussion: https://postgr.es/m/CAM527d-GL%3DJp2EJXBSnVBGPK-4XEZwWof5Cv8P0hghS_og6oAg%40mail.gmail.com
Backpatch-through: 19
---
 src/backend/access/transam/parallel.c | 11 +++++++++
 src/backend/commands/vacuumparallel.c | 35 +++++++++++++++++++++++++++
 src/backend/postmaster/autovacuum.c   | 27 +++++++++++++++++++++
 src/include/commands/vacuum.h         |  1 +
 4 files changed, 74 insertions(+)

diff --git a/src/backend/access/transam/parallel.c b/src/backend/access/transam/parallel.c
index e1806a9a28a..b51df2a4606 100644
--- a/src/backend/access/transam/parallel.c
+++ b/src/backend/access/transam/parallel.c
@@ -812,6 +812,17 @@ WaitForParallelWorkersToFinish(ParallelContext *pcxt)
 		 */
 		CHECK_FOR_INTERRUPTS();
 
+		/*
+		 * An autovacuum worker running a parallel vacuum (leader) propagates
+		 * changes to the cost-based delay parameters to its parallel workers
+		 * at its own cost delay points, which it no longer reaches while
+		 * waiting here. Do it here instead, on every wakeup, so that a config
+		 * reload or a change in the number of autovacuum workers sharing the
+		 * cost limit reaches the parallel workers before they finish.
+		 */
+		if (AmAutoVacuumWorkerProcess())
+			parallel_vacuum_refresh_cost_params();
+
 		for (i = 0; i < pcxt->nworkers_launched; ++i)
 		{
 			/*
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index d4572861000..da0ec3cc8ee 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -43,6 +43,7 @@
 #include "executor/instrument.h"
 #include "optimizer/paths.h"
 #include "pgstat.h"
+#include "postmaster/interrupt.h"
 #include "storage/bufmgr.h"
 #include "storage/proc.h"
 #include "tcop/tcopprot.h"
@@ -729,6 +730,40 @@ parallel_vacuum_propagate_shared_delay_params(void)
 	pg_atomic_fetch_add_u32(&pv_shared_cost_params->generation, 1);
 }
 
+/*
+ * Reload the configuration file if requested, and refresh the cost-based
+ * delay parameters of an autovacuum worker running a parallel vacuum
+ * (leader), propagating any change to its parallel workers.
+ */
+void
+parallel_vacuum_refresh_cost_params(void)
+{
+	Assert(AmAutoVacuumWorkerProcess());
+
+	/*
+	 * Quick return if the leader is not sharing the delay parameters with
+	 * parallel workers.
+	 */
+	if (pv_shared_cost_params == NULL)
+		return;
+
+	if (ConfigReloadPending)
+	{
+		ConfigReloadPending = false;
+		ProcessConfigFile(PGC_SIGHUP);
+
+		/* This re-divides the cost limit too */
+		VacuumUpdateCosts();
+	}
+	else
+	{
+		/* The number of workers sharing the cost limit may have changed */
+		AutoVacuumUpdateCostLimit();
+	}
+
+	parallel_vacuum_propagate_shared_delay_params();
+}
+
 /*
  * Compute the number of parallel worker processes to request.  Both index
  * vacuum and index cleanup can be executed with parallel workers.
diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c
index 60ebe828900..11da648e4a2 100644
--- a/src/backend/postmaster/autovacuum.c
+++ b/src/backend/postmaster/autovacuum.c
@@ -1814,8 +1814,35 @@ autovac_recalculate_workers_for_balance(void)
 	}
 
 	if (nworkers_for_balance != orig_nworkers_for_balance)
+	{
 		pg_atomic_write_u32(&AutoVacuumShmem->av_nworkersForBalance,
 							nworkers_for_balance);
+
+		/*
+		 * Wake up the autovacuum workers sharing the cost limit so that they
+		 * pick up the new count. An autovacuum worker that is vacuuming does
+		 * that on its next nap anyway, but one running a parallel vacuum
+		 * (leader) that is only waiting for its parallel workers to finish
+		 * never naps, and nothing else would tell it. The count is written
+		 * above, before the latches are set, so a woken leader always reads
+		 * the new value.
+		 *
+		 * Only the waiting leaders need this, but knowing which ones are
+		 * waiting would need more state. For an autovacuum worker that is not
+		 * in a latch wait, SetLatch() sends no signal and only marks the
+		 * latch set, which costs one early return from its next latch wait.
+		 */
+		dlist_foreach(iter, &AutoVacuumShmem->av_runningWorkers)
+		{
+			WorkerInfo	worker = dlist_container(WorkerInfoData, wi_links, iter.cur);
+
+			if (worker->wi_proc == NULL ||
+				pg_atomic_unlocked_test_flag(&worker->wi_dobalance))
+				continue;
+
+			SetLatch(&worker->wi_proc->procLatch);
+		}
+	}
 }
 
 /*
diff --git a/src/include/commands/vacuum.h b/src/include/commands/vacuum.h
index 6e3c912bf5c..89fa1fcd261 100644
--- a/src/include/commands/vacuum.h
+++ b/src/include/commands/vacuum.h
@@ -432,6 +432,7 @@ extern void parallel_vacuum_cleanup_all_indexes(ParallelVacuumState *pvs,
 												PVWorkerStats *wstats);
 extern void parallel_vacuum_update_shared_delay_params(void);
 extern void parallel_vacuum_propagate_shared_delay_params(void);
+extern void parallel_vacuum_refresh_cost_params(void);
 extern void parallel_vacuum_main(dsm_segment *seg, shm_toc *toc);
 
 /* in commands/analyze.c */
-- 
2.55.0



  [application/x-patch] v5-0002-Add-tests-for-cost-parameter-refresh-in-parallel-.patch (9.5K, ../../CAN4CZFOGQCy+fxtQEXHUqwF6gTLpNSdMtV9CxHC_0yd-f=NhBw@mail.gmail.com/3-v5-0002-Add-tests-for-cost-parameter-refresh-in-parallel-.patch)
  download | inline diff:
From 8834f65ae1ab05c10ba992223691c8076cf5e64c Mon Sep 17 00:00:00 2001
From: Nikolay Samokhvalov <nik@postgres.ai>
Date: Fri, 25 Sep 2026 19:01:30 +0000
Subject: [PATCH v5 2/2] Add tests for cost parameter refresh in parallel
 autovacuum wait.

The previous commit made an autovacuum worker running a parallel
vacuum (leader) refresh the cost-based delay parameters while it
waits in WaitForParallelWorkersToFinish(), both after a
configuration reload and after a change in the number of
autovacuum workers sharing the cost limit.  Add tests for both
cases.

To get the leader into that wait while its parallel worker still
has work left, add an injection point that keeps the leader out of
parallel-safe index vacuuming, leaving those indexes to the
parallel workers, and one that holds a parallel worker before each
index.  A third injection point fires once the leader has
refreshed the parameters, so the rebalance test knows when it is
safe to release the parallel worker.

Author: Nikolay Samokhvalov <nik@postgres.ai>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Reviewed-by: Masahiko Sawada <sawada.mshk@gmail.com>
Discussion: https://postgr.es/m/CAM527d-GL%3DJp2EJXBSnVBGPK-4XEZwWof5Cv8P0hghS_og6oAg%40mail.gmail.com
---
 src/backend/commands/vacuumparallel.c         |  23 ++-
 .../t/001_parallel_autovacuum.pl              | 152 ++++++++++++++++++
 2 files changed, 174 insertions(+), 1 deletion(-)

diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index da0ec3cc8ee..affb363cd5f 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -47,6 +47,7 @@
 #include "storage/bufmgr.h"
 #include "storage/proc.h"
 #include "tcop/tcopprot.h"
+#include "utils/injection_point.h"
 #include "utils/lsyscache.h"
 #include "utils/rel.h"
 
@@ -762,6 +763,8 @@ parallel_vacuum_refresh_cost_params(void)
 	}
 
 	parallel_vacuum_propagate_shared_delay_params();
+
+	INJECTION_POINT("parallel-autovacuum-leader-cost-updated", NULL);
 }
 
 /*
@@ -852,6 +855,7 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 {
 	int			nworkers;
 	PVIndVacStatus new_status;
+	bool		leader_participates = true;
 
 	Assert(!IsParallelWorker());
 
@@ -965,6 +969,17 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 							pvs->pcxt->nworkers_launched, nworkers)));
 	}
 
+#ifdef USE_INJECTION_POINTS
+
+	/*
+	 * Used by tests to leave all parallel-safe indexes to the parallel
+	 * workers, so that the leader waits for them to finish.
+	 */
+	if (nworkers > 0 && pvs->pcxt->nworkers_launched > 0 &&
+		IS_INJECTION_POINT_ATTACHED("parallel-vacuum-leader-skip-safe-indexes"))
+		leader_participates = false;
+#endif
+
 	/* Vacuum the indexes that can be processed by only leader process */
 	parallel_vacuum_process_unsafe_indexes(pvs);
 
@@ -972,7 +987,8 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 	 * Join as a parallel worker.  The leader vacuums alone processes all
 	 * parallel-safe indexes in the case where no workers are launched.
 	 */
-	parallel_vacuum_process_safe_indexes(pvs);
+	if (leader_participates)
+		parallel_vacuum_process_safe_indexes(pvs);
 
 	/*
 	 * Next, accumulate buffer and WAL usage.  (This must wait for the workers
@@ -1049,6 +1065,11 @@ parallel_vacuum_process_safe_indexes(ParallelVacuumState *pvs)
 		if (!indstats->parallel_workers_can_process)
 			continue;
 
+#ifdef USE_INJECTION_POINTS
+		if (IsParallelWorker())
+			INJECTION_POINT("parallel-vacuum-worker-before-index", NULL);
+#endif
+
 		/* Do vacuum or cleanup of the index */
 		parallel_vacuum_process_one_index(pvs, pvs->indrels[idx], indstats);
 	}
diff --git a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
index 33c86bbdc94..e49544eb0ca 100644
--- a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
+++ b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
@@ -252,6 +252,158 @@ $node->safe_psql('postgres',
 	"SELECT injection_points_wakeup('autovacuum-worker-cost-balanced')");
 $node->safe_psql('postgres',
 	"SELECT injection_points_detach('autovacuum-worker-cost-balanced')");
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+ok( $node->poll_query_until(
+		'postgres', q{
+		SELECT count(*) = 0 FROM pg_stat_activity
+		WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
+	}),
+	'second autovacuum worker finished');
+
+# Start autovacuum on test_autovac with its parallel worker held before its
+# first index, and the leader leaving all indexes to the parallel worker.
+# Returns once the leader waits for the parallel worker to finish.
+sub start_leader_waiting
+{
+	my ($node) = @_;
+
+	$node->safe_psql(
+		'postgres', q{
+		SELECT injection_points_attach('parallel-vacuum-worker-before-index', 'wait');
+		SELECT injection_points_attach('parallel-vacuum-leader-skip-safe-indexes', 'notice');
+		ALTER TABLE test_autovac SET (autovacuum_enabled = true);
+	});
+	$node->wait_for_event('parallel worker',
+		'parallel-vacuum-worker-before-index');
+	ok( $node->poll_query_until(
+			'postgres', q{
+			SELECT count(*) > 0 FROM pg_stat_activity
+			WHERE backend_type = 'autovacuum worker'
+			  AND wait_event = 'ParallelFinish'
+		}),
+		'autovacuum leader waits for its parallel worker');
+}
+
+# Release the parallel worker held by start_leader_waiting().
+sub release_parallel_worker
+{
+	my ($node) = @_;
+
+	$node->safe_psql(
+		'postgres', q{
+		SELECT injection_points_detach('parallel-vacuum-leader-skip-safe-indexes');
+		SELECT injection_points_detach('parallel-vacuum-worker-before-index');
+		SELECT injection_points_wakeup('parallel-vacuum-worker-before-index');
+	});
+}
+
+# Test 4:
+# Check whether a config reload is serviced while the autovacuum leader waits
+# for its parallel worker.
+my $postgresoid = $node->safe_psql('postgres',
+	"SELECT oid FROM pg_database WHERE datname = 'postgres'");
+my $testautovacid =
+  $node->safe_psql('postgres', "SELECT 'test_autovac'::regclass::oid");
+
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_max_workers = 1;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 700;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_delay = 0;
+	SELECT pg_reload_conf();
+});
+
+prepare_for_next_test($node, 4);
+$log_offset = -s $node->logfile;
+
+start_leader_waiting($node);
+
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 800;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_delay = 8;
+	ALTER SYSTEM SET vacuum_cost_page_miss = 11;
+	ALTER SYSTEM SET vacuum_cost_page_dirty = 12;
+	ALTER SYSTEM SET vacuum_cost_page_hit = 13;
+	SELECT pg_reload_conf();
+});
+
+# The leader must process the reload before the parallel worker is released.
+$node->wait_for_log(
+	qr/Autovacuum VacuumUpdateCosts\(db=$postgresoid, rel=$testautovacid, dobalance=yes, cost_limit=800, cost_delay=8 /,
+	$log_offset);
+
+release_parallel_worker($node);
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=800, cost_delay=8, cost_page_miss=11, cost_page_dirty=12, cost_page_hit=13/,
+	$log_offset);
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+ok(1, "config reload is propagated while the leader waits for workers");
+
+# Test 5:
+# Check the same wait path for a cost limit rebalance, which is not signaled
+# by a config reload.  A second autovacuum worker joins the balance while the
+# leader waits for its parallel worker.
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_max_workers = 2;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 600;
+	SELECT pg_reload_conf();
+});
+
+prepare_for_next_test($node, 5);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = false)');
+$node->safe_psql('regress_db2', 'UPDATE filler SET id = id + 1');
+
+$log_offset = -s $node->logfile;
+
+start_leader_waiting($node);
+
+# Hold the second worker, so that the balance stays at 2.  Attach the notice
+# before the rebalance, as the rebalance wakeup is the only thing that brings
+# the waiting leader to this point.
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_attach('autovacuum-worker-cost-balanced', 'wait');
+	SELECT injection_points_attach('parallel-autovacuum-leader-cost-updated', 'notice');
+});
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = true)');
+$node->wait_for_log(
+	qr/VacuumUpdateCosts\(db=$db2oid, rel=$filleroid, dobalance=yes, cost_limit=300,/,
+	$log_offset);
+$node->wait_for_log(
+	qr/notice triggered for injection point parallel-autovacuum-leader-cost-updated/,
+	$log_offset);
+$node->safe_psql('postgres',
+	"SELECT injection_points_detach('parallel-autovacuum-leader-cost-updated')"
+);
+
+release_parallel_worker($node);
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=300,/,
+	$log_offset);
+
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_wakeup('autovacuum-worker-cost-balanced');
+	SELECT injection_points_detach('autovacuum-worker-cost-balanced');
+});
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+ok( $node->poll_query_until(
+		'postgres', q{
+		SELECT count(*) = 0 FROM pg_stat_activity
+		WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
+	}),
+	'second autovacuum worker finished');
+ok(1, "cost rebalance is propagated while the leader waits for workers");
 
 $node->stop;
 done_testing();
-- 
2.55.0



^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-09-30 18:24  Daniel Gustafsson <daniel@yesql.se>
  parent: Masahiko Sawada <sawada.mshk@gmail.com>
  1 sibling, 1 reply; 43+ messages in thread

From: Daniel Gustafsson @ 2026-09-30 18:24 UTC (permalink / raw)
  To: Masahiko Sawada <sawada.mshk@gmail.com>; +Cc: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Nikolay Samokhvalov <nik@postgres.ai>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

> On 30 Sep 2026, at 04:19, Masahiko Sawada <sawada.mshk@gmail.com> wrote:

> Daniel, are you planning to push the 0001 patch separately from the
> tests? I will be attending Postgres Summit US 2026 this week, so I can
> push 0001 early next week instead just in case the buildfarm turns red
> and requires monitoring.

I hadn't really thought about that part yet, but if you are at a conference I
can get 0001 in this week to get it building in the BF well ahead of the next
release.  When you are back we can see about the 0002 part.

--
Daniel Gustafsson







^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-10-01 03:35  Masahiko Sawada <sawada.mshk@gmail.com>
  parent: Daniel Gustafsson <daniel@yesql.se>
  0 siblings, 1 reply; 43+ messages in thread

From: Masahiko Sawada @ 2026-10-01 03:35 UTC (permalink / raw)
  To: Daniel Gustafsson <daniel@yesql.se>; +Cc: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Nikolay Samokhvalov <nik@postgres.ai>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

On Wed, Sep 30, 2026 at 11:25 AM Daniel Gustafsson <daniel@yesql.se> wrote:
>
> > On 30 Sep 2026, at 04:19, Masahiko Sawada <sawada.mshk@gmail.com> wrote:
>
> > Daniel, are you planning to push the 0001 patch separately from the
> > tests? I will be attending Postgres Summit US 2026 this week, so I can
> > push 0001 early next week instead just in case the buildfarm turns red
> > and requires monitoring.
>
> I hadn't really thought about that part yet, but if you are at a conference I
> can get 0001 in this week to get it building in the BF well ahead of the next
> release.  When you are back we can see about the 0002 part.

That's very helpful for me. I'll monitor buildfarm animals as much as I can.

Regards,

-- 
Masahiko Sawada
Amazon Web Services: https://aws.amazon.com






^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-10-01 04:26  Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
  parent: Masahiko Sawada <sawada.mshk@gmail.com>
  0 siblings, 1 reply; 43+ messages in thread

From: Bharath Rupireddy @ 2026-10-01 04:26 UTC (permalink / raw)
  To: Masahiko Sawada <sawada.mshk@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; Nikolay Samokhvalov <nik@postgres.ai>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

Hi,

On Wed, Sep 30, 2026 at 8:36 PM Masahiko Sawada <sawada.mshk@gmail.com> wrote:
>
> On Wed, Sep 30, 2026 at 11:25 AM Daniel Gustafsson <daniel@yesql.se> wrote:
> >
> > > On 30 Sep 2026, at 04:19, Masahiko Sawada <sawada.mshk@gmail.com> wrote:
> >
> > > Daniel, are you planning to push the 0001 patch separately from the
> > > tests? I will be attending Postgres Summit US 2026 this week, so I can
> > > push 0001 early next week instead just in case the buildfarm turns red
> > > and requires monitoring.
> >
> > I hadn't really thought about that part yet, but if you are at a conference I
> > can get 0001 in this week to get it building in the BF well ahead of the next
> > release.  When you are back we can see about the 0002 part.
>
> That's very helpful for me. I'll monitor buildfarm animals as much as I can.

Thanks Daniel, and thanks Sawada-san. I will also keep an eye on the BF animals.

--
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com





^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-10-01 05:45  Masahiko Sawada <sawada.mshk@gmail.com>
  parent: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
  0 siblings, 1 reply; 43+ messages in thread

From: Masahiko Sawada @ 2026-10-01 05:45 UTC (permalink / raw)
  To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; Nikolay Samokhvalov <nik@postgres.ai>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

On Wed, Sep 30, 2026 at 9:26 PM Bharath Rupireddy
<bharath.rupireddyforpostgres@gmail.com> wrote:
>
> Hi,
>
> On Wed, Sep 30, 2026 at 8:36 PM Masahiko Sawada <sawada.mshk@gmail.com> wrote:
> >
> > On Wed, Sep 30, 2026 at 11:25 AM Daniel Gustafsson <daniel@yesql.se> wrote:
> > >
> > > > On 30 Sep 2026, at 04:19, Masahiko Sawada <sawada.mshk@gmail.com> wrote:
> > >
> > > > Daniel, are you planning to push the 0001 patch separately from the
> > > > tests? I will be attending Postgres Summit US 2026 this week, so I can
> > > > push 0001 early next week instead just in case the buildfarm turns red
> > > > and requires monitoring.
> > >
> > > I hadn't really thought about that part yet, but if you are at a conference I
> > > can get 0001 in this week to get it building in the BF well ahead of the next
> > > release.  When you are back we can see about the 0002 part.
> >
> > That's very helpful for me. I'll monitor buildfarm animals as much as I can.
>
> Thanks Daniel, and thanks Sawada-san. I will also keep an eye on the BF animals.

After more thought, I need to push another fix for the buildfarm
failure anyway, so I can take this patch too. Sorry for going back and
forth on this. I'll push the attached 0001 patch  tomorrow, barring
any objections. I tested the patch several times and confirmed it
passes every time. There's still room for improvement, though, so I'll
come back to the test patch next week and push it separately.

Regards,

-- 
Masahiko Sawada
Amazon Web Services: https://aws.amazon.com

Attachments:

  [text/x-patch] v6-0001-Refresh-autovacuum-costs-while-waiting-for-parall.patch (6.1K, ../../CAD21AoCJwm+88y5K0YW1hSBBhx5c8jLVjXjr__28Qxdh4QR74Q@mail.gmail.com/2-v6-0001-Refresh-autovacuum-costs-while-waiting-for-parall.patch)
  download | inline diff:
From a7b21c3cd9ba898c7da72ece98cbc3ddfffc46c2 Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Fri, 25 Sep 2026 18:52:57 +0000
Subject: [PATCH v6] Refresh autovacuum costs while waiting for parallel
 workers.

Commit 1ff3180ca01 made a parallel autovacuum leader propagate changes
to the cost-based delay parameters to its workers only from
vacuum_delay_point(), which it no longer reaches once it is waiting in
WaitForParallelWorkersToFinish(). While waiting, the leader neither
processed a config reload nor noticed a change in the number of
autovacuum workers sharing the cost limit, so its workers kept running
with stale parameters until they finished.

Fix by refreshing and propagating the parameters on every wakeup of
that wait loop, and by setting the latches of the balanced autovacuum
workers when their count changes so that a waiting leader wakes up.

A regression test will be added in a separate commit.

Backpatch to v19, where parallel autovacuum was introduced.

Reported-by: Nikolay Samokhvalov <nik@postgres.ai>
Author: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Reviewed-by: Manu <manuelreyesbravo@gmail.com>
Reviewed-by: Zsolt Parragi <zsolt.parragi@percona.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Reviewed-by: Masahiko Sawada <sawada.mshk@gmail.com>
Discussion: https://postgr.es/m/CAM527d-GL=Jp2EJXBSnVBGPK-4XEZwWof5Cv8P0hghS_og6oAg@mail.gmail.com
Backpatch-through: 19
---
 src/backend/access/transam/parallel.c | 11 +++++++++
 src/backend/commands/vacuumparallel.c | 35 +++++++++++++++++++++++++++
 src/backend/postmaster/autovacuum.c   | 25 +++++++++++++++++++
 src/include/commands/vacuum.h         |  1 +
 4 files changed, 72 insertions(+)

diff --git a/src/backend/access/transam/parallel.c b/src/backend/access/transam/parallel.c
index e1806a9a28a..b51df2a4606 100644
--- a/src/backend/access/transam/parallel.c
+++ b/src/backend/access/transam/parallel.c
@@ -812,6 +812,17 @@ WaitForParallelWorkersToFinish(ParallelContext *pcxt)
 		 */
 		CHECK_FOR_INTERRUPTS();
 
+		/*
+		 * An autovacuum worker running a parallel vacuum (leader) propagates
+		 * changes to the cost-based delay parameters to its parallel workers
+		 * at its own cost delay points, which it no longer reaches while
+		 * waiting here. Do it here instead, on every wakeup, so that a config
+		 * reload or a change in the number of autovacuum workers sharing the
+		 * cost limit reaches the parallel workers before they finish.
+		 */
+		if (AmAutoVacuumWorkerProcess())
+			parallel_vacuum_refresh_cost_params();
+
 		for (i = 0; i < pcxt->nworkers_launched; ++i)
 		{
 			/*
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index b8ceb5a7d29..b07124d1467 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -43,6 +43,7 @@
 #include "executor/instrument.h"
 #include "optimizer/paths.h"
 #include "pgstat.h"
+#include "postmaster/interrupt.h"
 #include "storage/bufmgr.h"
 #include "storage/proc.h"
 #include "tcop/tcopprot.h"
@@ -734,6 +735,40 @@ parallel_vacuum_propagate_shared_delay_params(void)
 	pg_atomic_fetch_add_u32(&pv_shared_cost_params->generation, 1);
 }
 
+/*
+ * Reload the configuration file if requested, and refresh the cost-based
+ * delay parameters of an autovacuum worker running a parallel vacuum
+ * (leader), propagating any change to its parallel workers.
+ */
+void
+parallel_vacuum_refresh_cost_params(void)
+{
+	Assert(AmAutoVacuumWorkerProcess());
+
+	/*
+	 * Quick return if the leader is not sharing the delay parameters with
+	 * parallel workers.
+	 */
+	if (pv_shared_cost_params == NULL)
+		return;
+
+	if (ConfigReloadPending)
+	{
+		ConfigReloadPending = false;
+		ProcessConfigFile(PGC_SIGHUP);
+
+		/* This re-divides the cost limit too */
+		VacuumUpdateCosts();
+	}
+	else
+	{
+		/* The number of workers sharing the cost limit may have changed */
+		AutoVacuumUpdateCostLimit();
+	}
+
+	parallel_vacuum_propagate_shared_delay_params();
+}
+
 /*
  * Compute the number of parallel worker processes to request.  Both index
  * vacuum and index cleanup can be executed with parallel workers.
diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c
index 60ebe828900..0ec01c5da6f 100644
--- a/src/backend/postmaster/autovacuum.c
+++ b/src/backend/postmaster/autovacuum.c
@@ -1814,8 +1814,33 @@ autovac_recalculate_workers_for_balance(void)
 	}
 
 	if (nworkers_for_balance != orig_nworkers_for_balance)
+	{
 		pg_atomic_write_u32(&AutoVacuumShmem->av_nworkersForBalance,
 							nworkers_for_balance);
+
+		/*
+		 * Wake up the autovacuum workers sharing the cost limit so that they
+		 * pick up the new count. An autovacuum worker that is vacuuming does
+		 * that on its next nap anyway, but one running a parallel vacuum
+		 * (leader) that is only waiting for its parallel workers to finish
+		 * never naps, and nothing else would tell it.
+		 *
+		 * Only the waiting leaders need this, but knowing which ones are
+		 * waiting would need more state. For an autovacuum worker that is not
+		 * in a latch wait, SetLatch() sends no signal and only marks the
+		 * latch set, which costs one early return from its next latch wait.
+		 */
+		dlist_foreach(iter, &AutoVacuumShmem->av_runningWorkers)
+		{
+			WorkerInfo	worker = dlist_container(WorkerInfoData, wi_links, iter.cur);
+
+			if (worker->wi_proc == NULL ||
+				pg_atomic_unlocked_test_flag(&worker->wi_dobalance))
+				continue;
+
+			SetLatch(&worker->wi_proc->procLatch);
+		}
+	}
 }
 
 /*
diff --git a/src/include/commands/vacuum.h b/src/include/commands/vacuum.h
index 6e3c912bf5c..89fa1fcd261 100644
--- a/src/include/commands/vacuum.h
+++ b/src/include/commands/vacuum.h
@@ -432,6 +432,7 @@ extern void parallel_vacuum_cleanup_all_indexes(ParallelVacuumState *pvs,
 												PVWorkerStats *wstats);
 extern void parallel_vacuum_update_shared_delay_params(void);
 extern void parallel_vacuum_propagate_shared_delay_params(void);
+extern void parallel_vacuum_refresh_cost_params(void);
 extern void parallel_vacuum_main(dsm_segment *seg, shm_toc *toc);
 
 /* in commands/analyze.c */
-- 
2.55.0



^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-10-01 07:11  Daniel Gustafsson <daniel@yesql.se>
  parent: Masahiko Sawada <sawada.mshk@gmail.com>
  0 siblings, 1 reply; 43+ messages in thread

From: Daniel Gustafsson @ 2026-10-01 07:11 UTC (permalink / raw)
  To: Masahiko Sawada <sawada.mshk@gmail.com>; +Cc: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Nikolay Samokhvalov <nik@postgres.ai>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

> On 1 Oct 2026, at 07:45, Masahiko Sawada <sawada.mshk@gmail.com> wrote:
> On Wed, Sep 30, 2026 at 9:26 PM Bharath Rupireddy
> <bharath.rupireddyforpostgres@gmail.com> wrote:
>> 
>> On Wed, Sep 30, 2026 at 8:36 PM Masahiko Sawada <sawada.mshk@gmail.com> wrote:

>>> That's very helpful for me. I'll monitor buildfarm animals as much as I can.
>> 
>> Thanks Daniel, and thanks Sawada-san. I will also keep an eye on the BF animals.
> 
> After more thought, I need to push another fix for the buildfarm
> failure anyway, so I can take this patch too. Sorry for going back and
> forth on this. I'll push the attached 0001 patch  tomorrow, barring
> any objections.

No worries at all, whatever works best for you.  Let me know if I can be of
assistance in any way while you're travelling.

> I'll come back to the test patch next week and push it separately.

Sounds good.

--
Daniel Gustafsson







^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-10-01 18:31  Masahiko Sawada <sawada.mshk@gmail.com>
  parent: Daniel Gustafsson <daniel@yesql.se>
  0 siblings, 0 replies; 43+ messages in thread

From: Masahiko Sawada @ 2026-10-01 18:31 UTC (permalink / raw)
  To: Daniel Gustafsson <daniel@yesql.se>; +Cc: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Nikolay Samokhvalov <nik@postgres.ai>; Zsolt Parragi <zsolt.parragi@percona.com>; pgsql-bugs@lists.postgresql.org

On Thu, Oct 1, 2026 at 12:11 AM Daniel Gustafsson <daniel@yesql.se> wrote:
>
> > On 1 Oct 2026, at 07:45, Masahiko Sawada <sawada.mshk@gmail.com> wrote:
> > On Wed, Sep 30, 2026 at 9:26 PM Bharath Rupireddy
> > <bharath.rupireddyforpostgres@gmail.com> wrote:
> >>
> >> On Wed, Sep 30, 2026 at 8:36 PM Masahiko Sawada <sawada.mshk@gmail.com> wrote:
>
> >>> That's very helpful for me. I'll monitor buildfarm animals as much as I can.
> >>
> >> Thanks Daniel, and thanks Sawada-san. I will also keep an eye on the BF animals.
> >
> > After more thought, I need to push another fix for the buildfarm
> > failure anyway, so I can take this patch too. Sorry for going back and
> > forth on this. I'll push the attached 0001 patch  tomorrow, barring
> > any objections.
>
> No worries at all, whatever works best for you.  Let me know if I can be of
> assistance in any way while you're travelling.

Pushed. Thank you for all your support.

Regards,

-- 
Masahiko Sawada
Amazon Web Services: https://aws.amazon.com






^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-10-06 20:14  Masahiko Sawada <sawada.mshk@gmail.com>
  parent: Zsolt Parragi <zsolt.parragi@percona.com>
  0 siblings, 1 reply; 43+ messages in thread

From: Masahiko Sawada @ 2026-10-06 20:14 UTC (permalink / raw)
  To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Nikolay Samokhvalov <nik@postgres.ai>; pgsql-bugs@lists.postgresql.org

On Wed, Sep 30, 2026 at 2:00 AM Zsolt Parragi <zsolt.parragi@percona.com> wrote:
>
> I addressed these in v5. 0001 is unchanged.
>
> > I think we could simplify the tests by
> > introducing a boolean variable, say, leader_participates
> > ...
>
> I am not convinced that this change made things simpler, but it is
> included in 0002.

I think it did; the v5 adds 3 injection points while the v4 added 4
injection points.

>
> > While not a problem with this patch, I wish wait_for_log could either return
> > the match, or at least the offset of the match, and not just the size of the
> > file.  If we had that we could avoid reading excessive amounts of log data and
> > reduce the risk of buggy tests matching on the wrong part of the log.
>
> I ended up completely removing that test, it seemed redundant.

Thinking about these regression tests, I think we can simplify them
further by using a single injection point. Instead of having the
workers process the indexes while the leader waits for them, we can
stop the worker at the beginning of parallel_vacuum_main() and have
the leader process all the indexes and wait in
WaitForParallelWorkersToFinish(). That way, we can create a situation
where the leader waits in WaitForParallelWorkersToFinish() with just
one injection point. This injection point can also be reused in the
other tests being discussed[1].

Also, instead of adding a 'parallel-autovacuum-leader-cost-updated'
injection point, the leader writes a DEBUG2 log after propagating the
shared delay params.

The attached patch implements this idea and can be applied on top of
the v5-0002 patch. It also adds more comments to these tests for
better readability.

Regards,

[1] https://www.postgresql.org/message-id/CALj2ACWK6PkOk5MTfibtJyTYpb%3DAq_8%3DYaexnPKhOj11bOCROQ%40mail...

--
Masahiko Sawada
Amazon Web Services: https://aws.amazon.com

Attachments:

  [text/x-patch] nocfbot_v5_masahiko.patch (10.3K, ../../CAD21AoDWYaz9h3C8erxFzRGjAjpet=_SfK2UBgYtZ--d4GBwSw@mail.gmail.com/2-nocfbot_v5_masahiko.patch)
  download | inline diff:
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 70364c0292e..8e33e9d8cbb 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -734,6 +734,12 @@ parallel_vacuum_propagate_shared_delay_params(void)
 	 * know that they should re-read shared cost params.
 	 */
 	pg_atomic_fetch_add_u32(&pv_shared_cost_params->generation, 1);
+
+	elog(DEBUG2,
+		 "parallel autovacuum leader propagated cost params: cost_limit=%d, cost_delay=%g, cost_page_miss=%d, cost_page_dirty=%d, cost_page_hit=%d, track_cost_delay_timing=%s",
+		 vacuum_cost_limit, vacuum_cost_delay,
+		 VacuumCostPageMiss, VacuumCostPageDirty, VacuumCostPageHit,
+		 track_cost_delay_timing ? "yes" : "no");
 }
 
 /*
@@ -768,8 +774,6 @@ parallel_vacuum_refresh_cost_params(void)
 	}
 
 	parallel_vacuum_propagate_shared_delay_params();
-
-	INJECTION_POINT("parallel-autovacuum-leader-cost-updated", NULL);
 }
 
 /*
@@ -860,7 +864,6 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 {
 	int			nworkers;
 	PVIndVacStatus new_status;
-	bool		leader_participates = true;
 
 	Assert(!IsParallelWorker());
 
@@ -974,17 +977,6 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 							pvs->pcxt->nworkers_launched, nworkers)));
 	}
 
-#ifdef USE_INJECTION_POINTS
-
-	/*
-	 * Used by tests to leave all parallel-safe indexes to the parallel
-	 * workers, so that the leader waits for them to finish.
-	 */
-	if (nworkers > 0 && pvs->pcxt->nworkers_launched > 0 &&
-		IS_INJECTION_POINT_ATTACHED("parallel-vacuum-leader-skip-safe-indexes"))
-		leader_participates = false;
-#endif
-
 	/* Vacuum the indexes that can be processed by only leader process */
 	parallel_vacuum_process_unsafe_indexes(pvs);
 
@@ -992,8 +984,7 @@ parallel_vacuum_process_all_indexes(ParallelVacuumState *pvs, int num_index_scan
 	 * Join as a parallel worker.  The leader vacuums alone processes all
 	 * parallel-safe indexes in the case where no workers are launched.
 	 */
-	if (leader_participates)
-		parallel_vacuum_process_safe_indexes(pvs);
+	parallel_vacuum_process_safe_indexes(pvs);
 
 	/*
 	 * Next, accumulate buffer and WAL usage.  (This must wait for the workers
@@ -1070,11 +1061,6 @@ parallel_vacuum_process_safe_indexes(ParallelVacuumState *pvs)
 		if (!indstats->parallel_workers_can_process)
 			continue;
 
-#ifdef USE_INJECTION_POINTS
-		if (IsParallelWorker())
-			INJECTION_POINT("parallel-vacuum-worker-before-index", NULL);
-#endif
-
 		/* Do vacuum or cleanup of the index */
 		parallel_vacuum_process_one_index(pvs, pvs->indrels[idx], indstats);
 	}
@@ -1310,6 +1296,8 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
 
 	elog(DEBUG1, "starting parallel vacuum worker");
 
+	INJECTION_POINT("parallel-vacuum-worker-start", NULL);
+
 	shared = (PVShared *) shm_toc_lookup(toc, PARALLEL_VACUUM_KEY_SHARED, false);
 
 	/* Set debug_query_string for individual workers */
diff --git a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
index 39960f6be74..a16df331fc6 100644
--- a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
+++ b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
@@ -263,50 +263,9 @@ ok( $node->poll_query_until(
 	}),
 	'second autovacuum worker finished');
 
-# Start autovacuum on test_autovac with its parallel worker held before its
-# first index, and the leader leaving all indexes to the parallel worker.
-# Returns once the leader waits for the parallel worker to finish.
-sub start_leader_waiting
-{
-	my ($node) = @_;
-
-	$node->safe_psql(
-		'postgres', q{
-		SELECT injection_points_attach('parallel-vacuum-worker-before-index', 'wait');
-		SELECT injection_points_attach('parallel-vacuum-leader-skip-safe-indexes', 'notice');
-		ALTER TABLE test_autovac SET (autovacuum_enabled = true);
-	});
-	$node->wait_for_event('parallel worker',
-		'parallel-vacuum-worker-before-index');
-	ok( $node->poll_query_until(
-			'postgres', q{
-			SELECT count(*) > 0 FROM pg_stat_activity
-			WHERE backend_type = 'autovacuum worker'
-			  AND wait_event = 'ParallelFinish'
-		}),
-		'autovacuum leader waits for its parallel worker');
-}
-
-# Release the parallel worker held by start_leader_waiting().
-sub release_parallel_worker
-{
-	my ($node) = @_;
-
-	$node->safe_psql(
-		'postgres', q{
-		SELECT injection_points_detach('parallel-vacuum-leader-skip-safe-indexes');
-		SELECT injection_points_detach('parallel-vacuum-worker-before-index');
-		SELECT injection_points_wakeup('parallel-vacuum-worker-before-index');
-	});
-}
-
 # Test 4:
 # Check whether a config reload is serviced while the autovacuum leader waits
 # for its parallel worker.
-my $postgresoid = $node->safe_psql('postgres',
-	"SELECT oid FROM pg_database WHERE datname = 'postgres'");
-my $testautovacid =
-  $node->safe_psql('postgres', "SELECT 'test_autovac'::regclass::oid");
 
 $node->safe_psql(
 	'postgres', qq{
@@ -319,8 +278,19 @@ $node->safe_psql(
 prepare_for_next_test($node, 4);
 $log_offset = -s $node->logfile;
 
-start_leader_waiting($node);
+# Let an autovacuum worker process test_autovac with its parallel worker held
+# before the parallel worker reads the cost-based delay parameters.  The
+# leader then processes all indexes by itself and waits for the parallel
+# worker to finish.  Change the parameters only once the leader is waiting,
+# or it would pick up the change before reaching the code path under test.
+$node->safe_psql(
+	'postgres', q{
+SELECT injection_points_attach('parallel-vacuum-worker-start', 'wait');
+ALTER TABLE test_autovac SET (autovacuum_enabled = true);
+});
+$node->wait_for_event('autovacuum worker', 'ParallelFinish');
 
+# Update cost-based delay parameters.
 $node->safe_psql(
 	'postgres', qq{
 	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 800;
@@ -331,15 +301,26 @@ $node->safe_psql(
 	SELECT pg_reload_conf();
 });
 
-# The leader must process the reload before the parallel worker is released.
+# The parallel worker reads the parameters only once, when it starts, since
+# the leader has already processed all indexes.  So the leader must have
+# propagated the new parameters before the parallel worker is released.
 $node->wait_for_log(
-	qr/Autovacuum VacuumUpdateCosts\(db=$postgresoid, rel=$testautovacid, dobalance=yes, cost_limit=800, cost_delay=8 /,
+	qr/parallel autovacuum leader propagated cost params: cost_limit=800,/,
 	$log_offset);
 
-release_parallel_worker($node);
+# Release the parallel worker.  It reads the cost-based delay parameters the
+# leader has propagated as soon as it resumes.
+$node->safe_psql(
+	'postgres',
+	q{
+SELECT injection_points_detach('parallel-vacuum-worker-start');
+SELECT injection_points_wakeup('parallel-vacuum-worker-start');
+});
 $node->wait_for_log(
 	qr/parallel autovacuum worker updated cost params: cost_limit=800, cost_delay=8, cost_page_miss=11, cost_page_dirty=12, cost_page_hit=13/,
 	$log_offset);
+
+# Wait for the autovacuum on test_autovac to finish.
 $node->wait_for_log(
 	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
 	$log_offset);
@@ -363,47 +344,63 @@ $node->safe_psql('regress_db2', 'UPDATE filler SET id = id + 1');
 
 $log_offset = -s $node->logfile;
 
-start_leader_waiting($node);
+# As in Test 4, hold the parallel worker and wait for the leader to process
+# all indexes and wait for the parallel worker to finish.
+$node->safe_psql(
+	'postgres', q{
+SELECT injection_points_attach('parallel-vacuum-worker-start', 'wait');
+ALTER TABLE test_autovac SET (autovacuum_enabled = true);
+});
+$node->wait_for_event('autovacuum worker', 'ParallelFinish');
 
-# Hold the second worker, so that the balance stays at 2.  Attach the notice
-# before the rebalance, as the rebalance wakeup is the only thing that brings
-# the waiting leader to this point.
+# Hold the second worker, so that the number of autovacuum workers sharing
+# the cost limit stays at 2 until the parallel worker has read the
+# parameters.  Once the second worker finishes, the leader would propagate
+# the original cost limit again.
 $node->safe_psql(
 	'postgres', q{
-	SELECT injection_points_attach('autovacuum-worker-cost-balanced', 'wait');
-	SELECT injection_points_attach('parallel-autovacuum-leader-cost-updated', 'notice');
+   SELECT injection_points_attach('autovacuum-worker-cost-balanced', 'wait');
 });
 $node->safe_psql('regress_db2',
 	'ALTER TABLE filler SET (autovacuum_enabled = true)');
+
+# Wait for the second worker to update its cost parameters.  It has
+# recalculated the number of workers sharing the cost limit, now 2, and woken
+# up the leader.
 $node->wait_for_log(
 	qr/VacuumUpdateCosts\(db=$db2oid, rel=$filleroid, dobalance=yes, cost_limit=300,/,
 	$log_offset);
+
+# Likewise, the leader must propagate the rebalanced cost limit before the
+# parallel worker is released.
 $node->wait_for_log(
-	qr/notice triggered for injection point parallel-autovacuum-leader-cost-updated/,
+	qr/parallel autovacuum leader propagated cost params: cost_limit=300,/,
 	$log_offset);
-$node->safe_psql('postgres',
-	"SELECT injection_points_detach('parallel-autovacuum-leader-cost-updated')"
-);
 
-release_parallel_worker($node);
+# Release the parallel worker.  It reads the cost-based delay parameters the
+# leader has propagated as soon as it resumes.
+$node->safe_psql(
+	'postgres',
+	q{
+SELECT injection_points_detach('parallel-vacuum-worker-start');
+SELECT injection_points_wakeup('parallel-vacuum-worker-start');
+});
 $node->wait_for_log(
 	qr/parallel autovacuum worker updated cost params: cost_limit=300,/,
 	$log_offset);
 
+# Release the second worker.
 $node->safe_psql(
-	'postgres', q{
-	SELECT injection_points_wakeup('autovacuum-worker-cost-balanced');
-	SELECT injection_points_detach('autovacuum-worker-cost-balanced');
+	'postgres',
+	q{
+SELECT injection_points_detach('autovacuum-worker-cost-balanced');
+SELECT injection_points_wakeup('autovacuum-worker-cost-balanced');
 });
+
+# Wait for the autovacuum on test_autovac to finish.
 $node->wait_for_log(
 	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
 	$log_offset);
-ok( $node->poll_query_until(
-		'postgres', q{
-		SELECT count(*) = 0 FROM pg_stat_activity
-		WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
-	}),
-	'second autovacuum worker finished');
 ok(1, "cost rebalance is propagated while the leader waits for workers");
 
 $node->stop;


^ permalink  raw  reply  [nested|flat] 43+ messages in thread

* Re: autovacuum: automatically propagate updated parameters
@ 2026-10-06 21:24  Zsolt Parragi <zsolt.parragi@percona.com>
  parent: Masahiko Sawada <sawada.mshk@gmail.com>
  0 siblings, 0 replies; 43+ messages in thread

From: Zsolt Parragi @ 2026-10-06 21:24 UTC (permalink / raw)
  To: Masahiko Sawada <sawada.mshk@gmail.com>; +Cc: Daniel Gustafsson <daniel@yesql.se>; Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Nikolay Samokhvalov <nik@postgres.ai>; pgsql-bugs@lists.postgresql.org

Hello!

> I think it did; the v5 adds 3 injection points while the v4 added 4
> injection points.

The overall patch size wasn't that different, the main reduction came
from removing the redundant last testcase, but your changes definitely
move it toward a simpler direction.

The changes look good to me. I squashed the previous 0002 and this
diff together for easier review / CI, fixed one whitespace issue, and
adjusted the commit message.

Attachments:

  [application/octet-stream] v6-0001-Add-tests-for-cost-parameter-refresh-in-parallel-.patch (8.7K, ../../CAN4CZFNziQAVBzSpRjeLvG7Wx7zqeokzw8_HQAYxg2knrEGhSw@mail.gmail.com/2-v6-0001-Add-tests-for-cost-parameter-refresh-in-parallel-.patch)
  download | inline diff:
From 3c24e1e581e5ad8339f99d2afa8cabdbbb9b9c85 Mon Sep 17 00:00:00 2001
From: Nikolay Samokhvalov <nik@postgres.ai>
Date: Fri, 25 Sep 2026 19:01:30 +0000
Subject: [PATCH v6] Add tests for cost parameter refresh in parallel
 autovacuum wait.

Commit 42e96cf2fe0 made an autovacuum worker running a parallel
vacuum (leader) refresh the cost-based delay parameters while it
waits in WaitForParallelWorkersToFinish(), both after a
configuration reload and after a change in the number of autovacuum
workers sharing the cost limit. Add tests for both cases.

To get the leader into that wait, add an injection point at the
start of parallel_vacuum_main() that holds the parallel worker
before it reads the cost-based delay parameters. The leader then
processes all indexes by itself and waits for the parallel worker
to finish. Also, make the leader log the parameters it propagates at
DEBUG2, so the tests can tell when it is safe to release the
parallel worker.

Author: Nikolay Samokhvalov <nik@postgres.ai>
Co-authored-by: Masahiko Sawada <sawada.mshk@gmail.com>
Reviewed-by: Zsolt Parragi <zsolt.parragi@percona.com>
Reviewed-by: Daniel Gustafsson <daniel@yesql.se>
Discussion: https://postgr.es/m/CAM527d-GL%3DJp2EJXBSnVBGPK-4XEZwWof5Cv8P0hghS_og6oAg%40mail.gmail.com
---
 src/backend/commands/vacuumparallel.c         |   9 ++
 .../t/001_parallel_autovacuum.pl              | 149 ++++++++++++++++++
 2 files changed, 158 insertions(+)

diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 4e432c50d37..8e33e9d8cbb 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -47,6 +47,7 @@
 #include "storage/bufmgr.h"
 #include "storage/proc.h"
 #include "tcop/tcopprot.h"
+#include "utils/injection_point.h"
 #include "utils/lsyscache.h"
 #include "utils/rel.h"
 
@@ -733,6 +734,12 @@ parallel_vacuum_propagate_shared_delay_params(void)
 	 * know that they should re-read shared cost params.
 	 */
 	pg_atomic_fetch_add_u32(&pv_shared_cost_params->generation, 1);
+
+	elog(DEBUG2,
+		 "parallel autovacuum leader propagated cost params: cost_limit=%d, cost_delay=%g, cost_page_miss=%d, cost_page_dirty=%d, cost_page_hit=%d, track_cost_delay_timing=%s",
+		 vacuum_cost_limit, vacuum_cost_delay,
+		 VacuumCostPageMiss, VacuumCostPageDirty, VacuumCostPageHit,
+		 track_cost_delay_timing ? "yes" : "no");
 }
 
 /*
@@ -1289,6 +1296,8 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
 
 	elog(DEBUG1, "starting parallel vacuum worker");
 
+	INJECTION_POINT("parallel-vacuum-worker-start", NULL);
+
 	shared = (PVShared *) shm_toc_lookup(toc, PARALLEL_VACUUM_KEY_SHARED, false);
 
 	/* Set debug_query_string for individual workers */
diff --git a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
index 115f92a49e4..93126bb6682 100644
--- a/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
+++ b/src/test/modules/test_autovacuum/t/001_parallel_autovacuum.pl
@@ -253,6 +253,155 @@ $node->safe_psql('postgres',
 	"SELECT injection_points_wakeup('autovacuum-worker-cost-balanced')");
 $node->safe_psql('postgres',
 	"SELECT injection_points_detach('autovacuum-worker-cost-balanced')");
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+ok( $node->poll_query_until(
+		'postgres', q{
+		SELECT count(*) = 0 FROM pg_stat_activity
+		WHERE backend_type = 'autovacuum worker' AND datname = 'regress_db2'
+	}),
+	'second autovacuum worker finished');
+
+# Test 4:
+# Check whether a config reload is serviced while the autovacuum leader waits
+# for its parallel worker.
+
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_max_workers = 1;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 700;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_delay = 0;
+	SELECT pg_reload_conf();
+});
+
+prepare_for_next_test($node, 4);
+$log_offset = -s $node->logfile;
+
+# Let an autovacuum worker process test_autovac with its parallel worker held
+# before the parallel worker reads the cost-based delay parameters.  The
+# leader then processes all indexes by itself and waits for the parallel
+# worker to finish.  Change the parameters only once the leader is waiting,
+# or it would pick up the change before reaching the code path under test.
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_attach('parallel-vacuum-worker-start', 'wait');
+	ALTER TABLE test_autovac SET (autovacuum_enabled = true);
+});
+$node->wait_for_event('autovacuum worker', 'ParallelFinish');
+
+# Update cost-based delay parameters.
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 800;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_delay = 8;
+	ALTER SYSTEM SET vacuum_cost_page_miss = 11;
+	ALTER SYSTEM SET vacuum_cost_page_dirty = 12;
+	ALTER SYSTEM SET vacuum_cost_page_hit = 13;
+	SELECT pg_reload_conf();
+});
+
+# The parallel worker reads the parameters only once, when it starts, since
+# the leader has already processed all indexes.  So the leader must have
+# propagated the new parameters before the parallel worker is released.
+$node->wait_for_log(
+	qr/parallel autovacuum leader propagated cost params: cost_limit=800,/,
+	$log_offset);
+
+# Release the parallel worker.  It reads the cost-based delay parameters the
+# leader has propagated as soon as it resumes.
+$node->safe_psql(
+	'postgres',
+	q{
+	SELECT injection_points_detach('parallel-vacuum-worker-start');
+	SELECT injection_points_wakeup('parallel-vacuum-worker-start');
+});
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=800, cost_delay=8, cost_page_miss=11, cost_page_dirty=12, cost_page_hit=13/,
+	$log_offset);
+
+# Wait for the autovacuum on test_autovac to finish.
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+ok(1, "config reload is propagated while the leader waits for workers");
+
+# Test 5:
+# Check the same wait path for a cost limit rebalance, which is not signaled
+# by a config reload.  A second autovacuum worker joins the balance while the
+# leader waits for its parallel worker.
+$node->safe_psql(
+	'postgres', qq{
+	ALTER SYSTEM SET autovacuum_max_workers = 2;
+	ALTER SYSTEM SET autovacuum_vacuum_cost_limit = 600;
+	SELECT pg_reload_conf();
+});
+
+prepare_for_next_test($node, 5);
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = false)');
+$node->safe_psql('regress_db2', 'UPDATE filler SET id = id + 1');
+
+$log_offset = -s $node->logfile;
+
+# As in Test 4, hold the parallel worker and wait for the leader to process
+# all indexes and wait for the parallel worker to finish.
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_attach('parallel-vacuum-worker-start', 'wait');
+	ALTER TABLE test_autovac SET (autovacuum_enabled = true);
+});
+$node->wait_for_event('autovacuum worker', 'ParallelFinish');
+
+# Hold the second worker, so that the number of autovacuum workers sharing
+# the cost limit stays at 2 until the parallel worker has read the
+# parameters.  Once the second worker finishes, the leader would propagate
+# the original cost limit again.
+$node->safe_psql(
+	'postgres', q{
+	SELECT injection_points_attach('autovacuum-worker-cost-balanced', 'wait');
+});
+$node->safe_psql('regress_db2',
+	'ALTER TABLE filler SET (autovacuum_enabled = true)');
+
+# Wait for the second worker to update its cost parameters.  It has
+# recalculated the number of workers sharing the cost limit, now 2, and woken
+# up the leader.
+$node->wait_for_log(
+	qr/VacuumUpdateCosts\(db=$db2oid, rel=$filleroid, dobalance=yes, cost_limit=300,/,
+	$log_offset);
+
+# Likewise, the leader must propagate the rebalanced cost limit before the
+# parallel worker is released.
+$node->wait_for_log(
+	qr/parallel autovacuum leader propagated cost params: cost_limit=300,/,
+	$log_offset);
+
+# Release the parallel worker.  It reads the cost-based delay parameters the
+# leader has propagated as soon as it resumes.
+$node->safe_psql(
+	'postgres',
+	q{
+	SELECT injection_points_detach('parallel-vacuum-worker-start');
+	SELECT injection_points_wakeup('parallel-vacuum-worker-start');
+});
+$node->wait_for_log(
+	qr/parallel autovacuum worker updated cost params: cost_limit=300,/,
+	$log_offset);
+
+# Release the second worker.
+$node->safe_psql(
+	'postgres',
+	q{
+	SELECT injection_points_detach('autovacuum-worker-cost-balanced');
+	SELECT injection_points_wakeup('autovacuum-worker-cost-balanced');
+});
+
+# Wait for the autovacuum on test_autovac to finish.
+$node->wait_for_log(
+	qr/automatic vacuum of table "postgres\.public\.test_autovac"/,
+	$log_offset);
+ok(1, "cost rebalance is propagated while the leader waits for workers");
 
 $node->stop;
 done_testing();
-- 
2.55.0



^ permalink  raw  reply  [nested|flat] 43+ messages in thread


end of thread, other threads:[~2026-10-06 21:24 UTC | newest]

Thread overview: 43+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2026-07-24 08:33 autovacuum: automatically propagate updated parameters Zsolt Parragi <zsolt.parragi@percona.com>
2026-08-24 12:31 ` Daniel Gustafsson <daniel@yesql.se>
2026-08-24 19:42   ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-24 22:06     ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-08-24 22:16       ` Daniel Gustafsson <daniel@yesql.se>
2026-08-24 22:26         ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-25 11:16           ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-08-27 00:07             ` Masahiko Sawada <sawada.mshk@gmail.com>
2026-08-27 08:37               ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-08-27 11:57                 ` Daniel Gustafsson <daniel@yesql.se>
2026-08-27 21:15                   ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-27 22:02                     ` Daniel Gustafsson <daniel@yesql.se>
2026-08-27 22:43                       ` Masahiko Sawada <sawada.mshk@gmail.com>
2026-08-27 22:52                         ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-08-28 07:13                           ` Daniel Gustafsson <daniel@yesql.se>
2026-08-28 22:31                             ` Daniel Gustafsson <daniel@yesql.se>
2026-09-21 18:28                         ` Nikolay Samokhvalov <nik@postgres.ai>
2026-09-24 00:38                           ` Masahiko Sawada <sawada.mshk@gmail.com>
2026-09-24 04:44                             ` Nikolay Samokhvalov <nik@postgres.ai>
2026-09-24 06:34                             ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-24 09:04                               ` Daniel Gustafsson <daniel@yesql.se>
2026-09-24 20:41                                 ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-09-24 21:35                                   ` Manu <manuelreyesbravo@gmail.com>
2026-09-24 23:46                                 ` Masahiko Sawada <sawada.mshk@gmail.com>
2026-09-25 00:50                                   ` Manu <manuelreyesbravo@gmail.com>
2026-09-25 07:45                                   ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-09-25 08:41                                     ` Daniel Gustafsson <daniel@yesql.se>
2026-09-25 19:12                                       ` Manu <manuelreyesbravo@gmail.com>
2026-09-25 21:41                                   ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-25 22:03                                     ` Manu <manuelreyesbravo@gmail.com>
2026-09-26 00:31                                     ` Masahiko Sawada <sawada.mshk@gmail.com>
2026-09-26 04:37                                       ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-28 13:13                                         ` Daniel Gustafsson <daniel@yesql.se>
2026-09-30 02:19                                           ` Masahiko Sawada <sawada.mshk@gmail.com>
2026-09-30 09:00                                             ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-10-06 20:14                                               ` Masahiko Sawada <sawada.mshk@gmail.com>
2026-10-06 21:24                                                 ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-09-30 18:24                                             ` Daniel Gustafsson <daniel@yesql.se>
2026-10-01 03:35                                               ` Masahiko Sawada <sawada.mshk@gmail.com>
2026-10-01 04:26                                                 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-10-01 05:45                                                   ` Masahiko Sawada <sawada.mshk@gmail.com>
2026-10-01 07:11                                                     ` Daniel Gustafsson <daniel@yesql.se>
2026-10-01 18:31                                                       ` Masahiko Sawada <sawada.mshk@gmail.com>

This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox