agora inbox for pgsql-hackers@postgresql.org
help / color / mirror / Atom feedReport index currently being vacuumed in pg_stat_progress_vacuum
32+ messages / 8 participants
[nested] [flat]
* Report index currently being vacuumed in pg_stat_progress_vacuum
@ 2026-05-04 02:00 Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-04 04:53 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
2026-05-04 09:52 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Antonin Houska <ah@cybertec.at>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
0 siblings, 3 replies; 32+ messages in thread
From: Bharath Rupireddy @ 2026-05-04 02:00 UTC (permalink / raw)
To: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>
Hi,
When VACUUM is in the "vacuuming indexes" or "cleaning up indexes" phase,
there is currently no easy way to tell which specific index is being
processed. The progress report view shows indexes_total and
indexes_processed counters, but not which index is actively being worked on.
This makes it difficult to debug slow or stuck autovacuum workers on tables
with multiple indexes of different types (btree, GIN, GiST, BRIN, HNSW,
etc.), since one cannot determine which index type or which specific index
is causing the delay.
Please find the attached patch adds a new column current_index_relid to
pg_stat_progress_vacuum that reports the OID of the index currently being
vacuumed or cleaned up. The column is reported for both the "vacuuming
indexes" phase and the "cleaning up indexes" phase.
When indexes are being vacuumed in parallel, each parallel worker emits its
own row in pg_stat_progress_vacuum with current_index_relid set to the
index it is currently processing, and leader_pid pointing to the leader
process.
Appreciate any feedback. Thank you!
[1] Example output:
pid | datname | relid | table_name | phase | started_by |
current_index_relid | index_name | leader_pid
------+----------+-------+------------+-------------------+------------+---------------------+---------------+------------
1420 | postgres | 16395 | vac_test | vacuuming indexes | autovacuum |
16398 | vac_test_idx1 |
1421 | postgres | 16395 | vac_test | vacuuming indexes | |
16399 | vac_test_idx2 | 1420
1423 | postgres | 16395 | vac_test | vacuuming indexes | |
16400 | vac_test_idx3 | 1420
(3 rows)
pid | datname | relid | table_name | phase | started_by |
current_index_relid | index_name | leader_pid
------+----------+-------+------------+-------------------+------------+---------------------+---------------+------------
1346 | postgres | 16395 | vac_test | vacuuming indexes | manual |
16398 | vac_test_idx1 |
(1 row)
[2]
SELECT v.pid, v.datname, v.relid, c.relname AS table_name,
v.phase, v.started_by, v.current_index_relid,
COALESCE(ic.relname, '') AS index_name, v.leader_pid
FROM pg_stat_progress_vacuum v
JOIN pg_class c
ON c.oid = v.relid
LEFT JOIN pg_class ic
ON ic.oid = v.current_index_relid
WHERE v.relid = $tbl_oid
ORDER BY
v.leader_pid,
v.pid;
--
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com
Attachments:
[application/x-patch] v1-0001-Report-index-currently-being-vacuumed-in-pg_stat_.patch (10.5K, ../../CALj2ACUgwSchK6jQ2CdKLBWUADTOE_zKdTff2Zg3E6hOuXKv-w@mail.gmail.com/3-v1-0001-Report-index-currently-being-vacuumed-in-pg_stat_.patch)
download | inline diff:
From 556d8c4a8db1b1ee23f19631e8c94ac702a26042 Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Fri, 17 Apr 2026 22:11:20 +0000
Subject: [PATCH v1] Report index currently being vacuumed in
pg_stat_progress_vacuum
When VACUUM is in the "vacuuming indexes" or "cleaning up indexes"
phase, there is no way to tell which specific index is being
processed. The view shows indexes_total and indexes_processed
counters, but not which index is actively being worked on.
This makes it difficult to debug slow or stuck autovacuum workers
on tables with multiple indexes (btree, GIN, GiST, BRIN, HNSW,
etc.), since one cannot determine which index type or which specific
index is causing the delay.
This commit adds a new column current_index_relid to
pg_stat_progress_vacuum that reports the OID of the index currently
being vacuumed or cleaned up. The value is set before each index
vacuum/cleanup begins and reset to 0 when all indexes have been
processed.
The column is reported for both the "vacuuming indexes" phase and
the "cleaning up indexes" phase.
When indexes are vacuumed in parallel, each parallel worker emits
its own row in pg_stat_progress_vacuum with current_index_relid set
to the index it is processing, and leader_pid set to the PID of the
leader. For the leader process itself (or non-parallel vacuum),
leader_pid is NULL.
---
doc/src/sgml/monitoring.sgml | 26 ++++++++++++++++++++++++++
src/backend/access/heap/vacuumlazy.c | 16 ++++++++++++++++
src/backend/access/transam/parallel.c | 9 +++++++++
src/backend/catalog/system_views.sql | 4 +++-
src/backend/commands/vacuumparallel.c | 21 +++++++++++++++++++++
src/include/access/parallel.h | 1 +
src/include/commands/progress.h | 2 ++
src/test/regress/expected/rules.out | 4 +++-
8 files changed, 81 insertions(+), 2 deletions(-)
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index 08d5b824552..0a7223608e6 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -7577,6 +7577,32 @@ FROM pg_stat_get_backend_idset() AS backendid;
</itemizedlist>
</para></entry>
</row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>current_index_relid</structfield> <type>oid</type>
+ </para>
+ <para>
+ If <command>VACUUM</command> is currently processing an index, this
+ column shows the OID of the index being vacuumed. The value is set
+ when the phase is <literal>vacuuming indexes</literal> or
+ <literal>cleaning up indexes</literal> and is reset to 0 when all
+ indexes have been processed. During parallel index vacuum, each
+ parallel worker row shows the index that particular worker is
+ processing.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>leader_pid</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Process ID of the leader process if this process is a parallel
+ vacuum worker; <literal>NULL</literal> if this process is the
+ leader itself or if the vacuum is not running in parallel.
+ </para></entry>
+ </row>
</tbody>
</tgroup>
</table>
diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index 39395aed0d5..ea040036251 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -3025,6 +3025,10 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3046,6 +3050,10 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /* Reset the current index relid to avoid reporting a stale value */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) InvalidOid);
+
return istat;
}
@@ -3076,6 +3084,10 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3095,6 +3107,10 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /* Reset the current index relid to avoid reporting a stale value */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) InvalidOid);
+
return istat;
}
diff --git a/src/backend/access/transam/parallel.c b/src/backend/access/transam/parallel.c
index 89e9d224eec..23a2e5eecf0 100644
--- a/src/backend/access/transam/parallel.c
+++ b/src/backend/access/transam/parallel.c
@@ -131,6 +131,15 @@ static dlist_head pcxt_list = DLIST_STATIC_INIT(pcxt_list);
/* Backend-local copy of data from FixedParallelState. */
static pid_t ParallelLeaderPid;
+/*
+ * Return the PID of the parallel group leader.
+ */
+pid_t
+GetParallelLeaderPid(void)
+{
+ return ParallelLeaderPid;
+}
+
/*
* List of internal parallel worker entry points. We need this for
* reasons explained in LookupParallelWorkerFunction(), below.
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 73a1c1c4670..22e9294bacc 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -1342,7 +1342,9 @@ CREATE VIEW pg_stat_progress_vacuum AS
CASE S.param13 WHEN 1 THEN 'manual'
WHEN 2 THEN 'autovacuum'
WHEN 3 THEN 'autovacuum_wraparound'
- ELSE NULL END AS started_by
+ ELSE NULL END AS started_by,
+ CAST(S.param14 AS oid) AS current_index_relid,
+ NULLIF(S.param15, 0) AS leader_pid
FROM pg_stat_get_progress_info('VACUUM') AS S
LEFT JOIN pg_database D ON S.datid = D.oid;
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 979c2be4abd..613b379a3ca 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -36,6 +36,7 @@
#include "postgres.h"
#include "access/amapi.h"
+#include "access/parallel.h"
#include "access/table.h"
#include "access/xact.h"
#include "commands/progress.h"
@@ -1076,6 +1077,11 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
IndexBulkDeleteResult *istat = NULL;
IndexBulkDeleteResult *istat_res;
IndexVacuumInfo ivinfo;
+ const int progress_index[] = {
+ PROGRESS_VACUUM_PHASE,
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID
+ };
+ int64 progress_val[2];
/*
* Update the pointer to the corresponding bulk-deletion result if someone
@@ -1112,6 +1118,13 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
RelationGetRelationName(indrel));
}
+ /* Report which index we're currently processing and the current phase */
+ progress_val[0] = (indstats->status == PARALLEL_INDVAC_STATUS_NEED_BULKDELETE)
+ ? PROGRESS_VACUUM_PHASE_VACUUM_INDEX
+ : PROGRESS_VACUUM_PHASE_INDEX_CLEANUP;
+ progress_val[1] = (int64) RelationGetRelid(indrel);
+ pgstat_progress_update_multi_param(2, progress_index, progress_val);
+
/*
* Copy the index bulk-deletion result returned from ambulkdelete and
* amvacuumcleanup to the DSM segment if it's the first cycle because they
@@ -1307,6 +1320,11 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Prepare to track buffer usage during parallel execution */
InstrStartParallelQuery();
+ /* Register this worker for vacuum progress reporting */
+ pgstat_progress_start_command(PROGRESS_COMMAND_VACUUM, shared->relid);
+ pgstat_progress_update_param(PROGRESS_VACUUM_LEADER_PID,
+ (int64) GetParallelLeaderPid());
+
/* Process indexes to perform vacuum/cleanup */
parallel_vacuum_process_safe_indexes(&pvs);
@@ -1326,6 +1344,9 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Pop the error context stack */
error_context_stack = errcallback.previous;
+ /* Unregister this worker from vacuum progress reporting */
+ pgstat_progress_end_command();
+
vac_close_indexes(nindexes, indrels, RowExclusiveLock);
table_close(rel, ShareUpdateExclusiveLock);
FreeAccessStrategy(pvs.bstrategy);
diff --git a/src/include/access/parallel.h b/src/include/access/parallel.h
index 60f857675e0..d329a2424bd 100644
--- a/src/include/access/parallel.h
+++ b/src/include/access/parallel.h
@@ -77,6 +77,7 @@ extern void ProcessParallelMessages(void);
extern void AtEOXact_Parallel(bool isCommit);
extern void AtEOSubXact_Parallel(bool isCommit, SubTransactionId mySubId);
extern void ParallelWorkerReportLastRecEnd(XLogRecPtr last_xlog_end);
+extern pid_t GetParallelLeaderPid(void);
extern void ParallelWorkerMain(Datum main_arg);
diff --git a/src/include/commands/progress.h b/src/include/commands/progress.h
index 2a12920c75f..f9f3e187718 100644
--- a/src/include/commands/progress.h
+++ b/src/include/commands/progress.h
@@ -31,6 +31,8 @@
#define PROGRESS_VACUUM_DELAY_TIME 10
#define PROGRESS_VACUUM_MODE 11
#define PROGRESS_VACUUM_STARTED_BY 12
+#define PROGRESS_VACUUM_CURRENT_INDEX_RELID 13
+#define PROGRESS_VACUUM_LEADER_PID 14
/* Phases of vacuum (as advertised via PROGRESS_VACUUM_PHASE) */
#define PROGRESS_VACUUM_PHASE_SCAN_HEAP 1
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index a65a5bf0c4f..00c0af6c24c 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -2202,7 +2202,9 @@ pg_stat_progress_vacuum| SELECT s.pid,
WHEN 2 THEN 'autovacuum'::text
WHEN 3 THEN 'autovacuum_wraparound'::text
ELSE NULL::text
- END AS started_by
+ END AS started_by,
+ (s.param14)::oid AS current_index_relid,
+ NULLIF(s.param15, 0) AS leader_pid
FROM (pg_stat_get_progress_info('VACUUM'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
LEFT JOIN pg_database d ON ((s.datid = d.oid)));
pg_stat_recovery| SELECT promote_triggered,
--
2.47.3
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
@ 2026-05-04 04:53 ` SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
2 siblings, 0 replies; 32+ messages in thread
From: SATYANARAYANA NARLAPURAM @ 2026-05-04 04:53 UTC (permalink / raw)
To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>
Hi,
On Sun, May 3, 2026 at 7:01 PM Bharath Rupireddy <
bharath.rupireddyforpostgres@gmail.com> wrote:
> Hi,
>
> When VACUUM is in the "vacuuming indexes" or "cleaning up indexes" phase,
> there is currently no easy way to tell which specific index is being
> processed. The progress report view shows indexes_total and
> indexes_processed counters, but not which index is actively being worked on.
>
> This makes it difficult to debug slow or stuck autovacuum workers on
> tables with multiple indexes of different types (btree, GIN, GiST, BRIN,
> HNSW, etc.), since one cannot determine which index type or which specific
> index is causing the delay.
>
> Please find the attached patch adds a new column current_index_relid to
> pg_stat_progress_vacuum that reports the OID of the index currently being
> vacuumed or cleaned up. The column is reported for both the "vacuuming
> indexes" phase and the "cleaning up indexes" phase.
>
> When indexes are being vacuumed in parallel, each parallel worker emits
> its own row in pg_stat_progress_vacuum with current_index_relid set to the
> index it is currently processing, and leader_pid pointing to the leader
> process.
>
> Appreciate any feedback. Thank you!
>
> [1] Example output:
>
> pid | datname | relid | table_name | phase | started_by |
> current_index_relid | index_name | leader_pid
>
> ------+----------+-------+------------+-------------------+------------+---------------------+---------------+------------
> 1420 | postgres | 16395 | vac_test | vacuuming indexes | autovacuum |
> 16398 | vac_test_idx1 |
> 1421 | postgres | 16395 | vac_test | vacuuming indexes | |
> 16399 | vac_test_idx2 | 1420
> 1423 | postgres | 16395 | vac_test | vacuuming indexes | |
> 16400 | vac_test_idx3 | 1420
> (3 rows)
>
> pid | datname | relid | table_name | phase | started_by |
> current_index_relid | index_name | leader_pid
>
> ------+----------+-------+------------+-------------------+------------+---------------------+---------------+------------
> 1346 | postgres | 16395 | vac_test | vacuuming indexes | manual |
> 16398 | vac_test_idx1 |
> (1 row)
>
> [2]
> SELECT v.pid, v.datname, v.relid, c.relname AS table_name,
> v.phase, v.started_by, v.current_index_relid,
> COALESCE(ic.relname, '') AS index_name, v.leader_pid
> FROM pg_stat_progress_vacuum v
> JOIN pg_class c
> ON c.oid = v.relid
> LEFT JOIN pg_class ic
> ON ic.oid = v.current_index_relid
> WHERE v.relid = $tbl_oid
> ORDER BY
> v.leader_pid,
> v.pid;
>
Bharath, thanks for the patch! A few comments:
(1) Do we need a global API? Can we add a leader_pid field in PVShared?
+pid_t
+GetParallelLeaderPid(void)
+{
+ return ParallelLeaderPid;
+}
(2): Looks like current_index_relid is not cleared when we leave the index
phases.As a result, once any index has been processed,
pg_stat_progress_vacuum.current_index_relid keeps reporting that relid
through vacuuming heap, truncating heap, cleaning up indexes.
This will be confusing to the user. Something like below:
1795819|vacuuming heap|0/0|16392|t1_pkey|LEADER
(3) leader_pid type should be integer type similar to pg_Stat_activity?
Thanks,
Satya
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
@ 2026-05-04 09:52 ` Antonin Houska <ah@cybertec.at>
2 siblings, 0 replies; 32+ messages in thread
From: Antonin Houska @ 2026-05-04 09:52 UTC (permalink / raw)
To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>
Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com> wrote:
> When VACUUM is in the "vacuuming indexes" or "cleaning up indexes" phase, there is currently no easy way to tell which specific index is
> being processed. The progress report view shows indexes_total and indexes_processed counters, but not which index is actively being worked
> on.
>
> This makes it difficult to debug slow or stuck autovacuum workers on tables with multiple indexes of different types (btree, GIN, GiST, BRIN,
> HNSW, etc.), since one cannot determine which index type or which specific index is causing the delay.
>
> Please find the attached patch adds a new column current_index_relid to pg_stat_progress_vacuum that reports the OID of the index
> currently being vacuumed or cleaned up. The column is reported for both the "vacuuming indexes" phase and the "cleaning up indexes"
> phase.
>
> When indexes are being vacuumed in parallel, each parallel worker emits its own row in pg_stat_progress_vacuum with current_index_relid
> set to the index it is currently processing, and leader_pid pointing to the leader process.
>
> Appreciate any feedback. Thank you!
This problem seems to be similar to what I noticed when workign on the REPACK
command: progress reporting of index build needs to be disabled if the build
is part of REPACK, otherwise the index build can overwrite the counters of
REPACK (whether the overwriting actually happens or not is another question).
The solution I suggest is to allow progress tracking of a "sub-command" - see
the attached patch. Wouldn't that also resolve your problem? (My plan is to
incorporate this in the series of REPACK enhancements soon.)
--
Antonin Houska
Web: https://www.cybertec-postgresql.com
Attachments:
[text/x-diff] nocfbot.Allow-progress-tracking-of-sub-commands.patch (14.3K, ../../30939.1777888333@localhost/2-nocfbot.Allow-progress-tracking-of-sub-commands.patch)
download | inline diff:
From 132ca28afd8d47bea0d40eb1db782709e549aa07 Mon Sep 17 00:00:00 2001
From: Antonin Houska <ah@cybertec.at>
Date: Mon, 4 May 2026 11:32:21 +0200
Subject: [PATCH] Allow progress tracking of sub-commands.
Some commands that support progress reporting run sub-commands, which also
report their progress. The typical case is that REPACK builds indexes. Instead
of disabling the progress tracking of the sub-commands, we can allow both the
"parent" command and the sub-command to report their progress at the same
time.
---
contrib/file_fdw/file_fdw.c | 44 +++++++++++
src/backend/commands/indexcmds.c | 10 ++-
src/backend/commands/repack.c | 14 +++-
src/backend/utils/activity/backend_progress.c | 76 ++++++++++++++++---
src/backend/utils/activity/backend_status.c | 1 +
src/backend/utils/adt/pgstatfuncs.c | 9 ++-
src/include/utils/backend_status.h | 7 ++
7 files changed, 142 insertions(+), 19 deletions(-)
diff --git a/contrib/file_fdw/file_fdw.c b/contrib/file_fdw/file_fdw.c
index 33a37d832ce..a60fb226320 100644
--- a/contrib/file_fdw/file_fdw.c
+++ b/contrib/file_fdw/file_fdw.c
@@ -36,6 +36,7 @@
#include "optimizer/pathnode.h"
#include "optimizer/planmain.h"
#include "optimizer/restrictinfo.h"
+#include "utils/backend_status.h"
#include "utils/acl.h"
#include "utils/memutils.h"
#include "utils/rel.h"
@@ -119,6 +120,16 @@ typedef struct FileFdwExecutionState
CopyFromState cstate; /* COPY execution state */
} FileFdwExecutionState;
+/*
+ * Since progress tracking of multiple COPY commands is not supported, the
+ * first file_fdw node of the plan needs to set pgstat_track_activities to
+ * false during startup, and the last active node needs to restore the
+ * original value during shutdown.
+ */
+static bool save_pgstat_track_activities = false;
+static int fdw_nodes = 0;
+static int active_fdw_nodes = 0;
+
/*
* SQL functions
*/
@@ -616,6 +627,12 @@ fileGetForeignPlan(PlannerInfo *root,
{
Index scan_relid = baserel->relid;
+ /*
+ * This seems to be the appropriate place to count file_fdw nodes in the
+ * plan.
+ */
+ fdw_nodes++;
+
/*
* We have no native ability to evaluate restriction clauses, so we just
* put all the scan_clauses into the plan node's qual list for the
@@ -695,6 +712,18 @@ fileBeginForeignScan(ForeignScanState *node, int eflags)
/* Add any options from the plan (currently only convert_selectively) */
options = list_concat(options, plan->fdw_private);
+ /*
+ * Save the value of pgstat_track_activities if this is the first file_fdw
+ * node of a plan containing multiple file_fdw nodes, and disable the
+ * progress tracking. The monitoring infrastructure currently does not
+ * support monitoring of multiple COPY commands.
+ */
+ if (fdw_nodes > 1 && active_fdw_nodes++ == 0)
+ {
+ save_pgstat_track_activities = pgstat_track_activities;
+ pgstat_track_activities = false;
+ }
+
/*
* Create CopyState from FDW options. We always acquire all columns, so
* as to match the expected ScanTupleSlot signature.
@@ -861,6 +890,21 @@ fileEndForeignScan(ForeignScanState *node)
festate->cstate->num_errors));
EndCopyFrom(festate->cstate);
+
+
+ /*
+ * Restore the value of pgstat_track_activities if this is the last
+ * file_fdw node of a plan containing multiple file_fdw nodes, and enable
+ * progress tracking if we disabled it earlier.
+ */
+ if (active_fdw_nodes > 0)
+ {
+ if (--active_fdw_nodes == 0)
+ {
+ pgstat_track_activities = save_pgstat_track_activities;
+ fdw_nodes = 0;
+ }
+ }
}
/*
diff --git a/src/backend/commands/indexcmds.c b/src/backend/commands/indexcmds.c
index 9ab74c8df0a..7280e64d118 100644
--- a/src/backend/commands/indexcmds.c
+++ b/src/backend/commands/indexcmds.c
@@ -3918,6 +3918,12 @@ ReindexRelationConcurrently(const ReindexStmt *stmt, Oid relationOid, const Rein
* more detailed comments.
*/
+ /*
+ * XXX Is there a reason not to start progress reporting here? If it's ok,
+ * then INDEX_CREATE_SUPPRESS_PROGRESS below is probably not needed.
+ */
+ pgstat_progress_start_command(PROGRESS_COMMAND_CREATE_INDEX, relationOid);
+
foreach(lc, indexIds)
{
char *concurrentName;
@@ -3966,8 +3972,6 @@ ReindexRelationConcurrently(const ReindexStmt *stmt, Oid relationOid, const Rein
if (indexRel->rd_rel->relpersistence == RELPERSISTENCE_TEMP)
elog(ERROR, "cannot reindex a temporary table concurrently");
- pgstat_progress_start_command(PROGRESS_COMMAND_CREATE_INDEX, idx->tableId);
-
progress_vals[0] = PROGRESS_CREATEIDX_COMMAND_REINDEX_CONCURRENTLY;
progress_vals[1] = 0; /* initializing */
progress_vals[2] = idx->indexId;
@@ -4144,7 +4148,6 @@ ReindexRelationConcurrently(const ReindexStmt *stmt, Oid relationOid, const Rein
* Update progress for the index to build, with the correct parent
* table involved.
*/
- pgstat_progress_start_command(PROGRESS_COMMAND_CREATE_INDEX, newidx->tableId);
progress_vals[0] = PROGRESS_CREATEIDX_COMMAND_REINDEX_CONCURRENTLY;
progress_vals[1] = PROGRESS_CREATEIDX_PHASE_BUILD;
progress_vals[2] = newidx->indexId;
@@ -4208,7 +4211,6 @@ ReindexRelationConcurrently(const ReindexStmt *stmt, Oid relationOid, const Rein
* Update progress for the index to build, with the correct parent
* table involved.
*/
- pgstat_progress_start_command(PROGRESS_COMMAND_CREATE_INDEX, newidx->tableId);
progress_vals[0] = PROGRESS_CREATEIDX_COMMAND_REINDEX_CONCURRENTLY;
progress_vals[1] = PROGRESS_CREATEIDX_PHASE_VALIDATE_IDXSCAN;
progress_vals[2] = newidx->indexId;
diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 9d162957bc3..f360f37a9da 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -1937,6 +1937,7 @@ finish_heap_swap(Oid OIDOldHeap, Oid OIDNewHeap,
pgstat_progress_update_param(PROGRESS_REPACK_PHASE,
PROGRESS_REPACK_PHASE_REBUILD_INDEX);
+ reindex_params.options |= REINDEXOPT_REPORT_PROGRESS;
reindex_relation(NULL, OIDOldHeap, reindex_flags, &reindex_params);
}
@@ -3204,9 +3205,20 @@ build_new_indexes(Relation NewHeap, Relation OldHeap, List *OldIndexes)
"repacknew",
get_rel_namespace(ind->rd_index->indrelid),
false);
- newindex = index_create_copy(NewHeap, INDEX_CREATE_SUPPRESS_PROGRESS,
+
+ /*
+ * We build the index on the new heap, but after the swap phase it'll
+ * become an index on the old heap. It makes more sense to report the
+ * progress this way. (The reporting API expects that both command and
+ * subcommand deal with the same target.)
+ */
+ pgstat_progress_start_command(PROGRESS_COMMAND_CREATE_INDEX,
+ RelationGetRelid(OldHeap));
+ newindex = index_create_copy(NewHeap, 0,
oldindex, ind->rd_rel->reltablespace,
newName);
+ pgstat_progress_end_command();
+
copy_index_constraints(ind, newindex, RelationGetRelid(NewHeap));
result = lappend_oid(result, newindex);
diff --git a/src/backend/utils/activity/backend_progress.c b/src/backend/utils/activity/backend_progress.c
index b0359771de5..97965a6973c 100644
--- a/src/backend/utils/activity/backend_progress.c
+++ b/src/backend/utils/activity/backend_progress.c
@@ -22,6 +22,10 @@
*
* Set st_progress_command (and st_progress_command_target) in own backend
* entry. Also, zero-initialize st_progress_param array.
+ *
+ * If command has already been started, start a sub-command. Only parameters
+ * of the sub-command are updated until pgstat_progress_end_command() is
+ * called. (Target relation must be the same for both commands.)
*-----------
*/
void
@@ -33,9 +37,30 @@ pgstat_progress_start_command(ProgressCommandType cmdtype, Oid relid)
return;
PGSTAT_BEGIN_WRITE_ACTIVITY(beentry);
- beentry->st_progress_command = cmdtype;
- beentry->st_progress_command_target = relid;
- MemSet(&beentry->st_progress_param, 0, sizeof(beentry->st_progress_param));
+ /* Sub-command should not be started w/o parent command. */
+ if (beentry->st_progress_command == PROGRESS_COMMAND_INVALID)
+ {
+ Assert(beentry->st_progress_command2 == PROGRESS_COMMAND_INVALID);
+
+ beentry->st_progress_command = cmdtype;
+ beentry->st_progress_command_target = relid;
+ MemSet(&beentry->st_progress_param, 0,
+ sizeof(beentry->st_progress_param));
+ }
+ else if (beentry->st_progress_command2 == PROGRESS_COMMAND_INVALID)
+ {
+ Assert(beentry->st_progress_command != PROGRESS_COMMAND_INVALID);
+ Assert(beentry->st_progress_command_target == relid);
+
+ beentry->st_progress_command2 = cmdtype;
+ MemSet(&beentry->st_progress_param2, 0,
+ sizeof(beentry->st_progress_param2));
+ }
+ else
+ {
+ /* Only one level of nesting is supported. */
+ Assert(false);
+ }
PGSTAT_END_WRITE_ACTIVITY(beentry);
}
@@ -49,14 +74,20 @@ void
pgstat_progress_update_param(int index, int64 val)
{
volatile PgBackendStatus *beentry = MyBEEntry;
+ volatile int64 *params;
Assert(index >= 0 && index < PGSTAT_NUM_PROGRESS_PARAM);
+ Assert(beentry->st_progress_command != PROGRESS_COMMAND_INVALID ||
+ beentry->st_progress_command2 != PROGRESS_COMMAND_INVALID);
if (!beentry || !pgstat_track_activities)
return;
+ params = (beentry->st_progress_command2 == PROGRESS_COMMAND_INVALID) ?
+ beentry->st_progress_param : beentry->st_progress_param2;
+
PGSTAT_BEGIN_WRITE_ACTIVITY(beentry);
- beentry->st_progress_param[index] = val;
+ params[index] = val;
PGSTAT_END_WRITE_ACTIVITY(beentry);
}
@@ -70,14 +101,20 @@ void
pgstat_progress_incr_param(int index, int64 incr)
{
volatile PgBackendStatus *beentry = MyBEEntry;
+ volatile int64 *params;
Assert(index >= 0 && index < PGSTAT_NUM_PROGRESS_PARAM);
+ Assert(beentry->st_progress_command != PROGRESS_COMMAND_INVALID ||
+ beentry->st_progress_command2 != PROGRESS_COMMAND_INVALID);
if (!beentry || !pgstat_track_activities)
return;
+ params = (beentry->st_progress_command2 == PROGRESS_COMMAND_INVALID) ?
+ beentry->st_progress_param : beentry->st_progress_param2;
+
PGSTAT_BEGIN_WRITE_ACTIVITY(beentry);
- beentry->st_progress_param[index] += incr;
+ params[index] += incr;
PGSTAT_END_WRITE_ACTIVITY(beentry);
}
@@ -124,17 +161,24 @@ pgstat_progress_update_multi_param(int nparam, const int *index,
{
volatile PgBackendStatus *beentry = MyBEEntry;
int i;
+ volatile int64 *params;
if (!beentry || !pgstat_track_activities || nparam == 0)
return;
+ Assert(beentry->st_progress_command != PROGRESS_COMMAND_INVALID ||
+ beentry->st_progress_command2 != PROGRESS_COMMAND_INVALID);
+
+ params = (beentry->st_progress_command2 == PROGRESS_COMMAND_INVALID) ?
+ beentry->st_progress_param : beentry->st_progress_param2;
+
PGSTAT_BEGIN_WRITE_ACTIVITY(beentry);
for (i = 0; i < nparam; ++i)
{
Assert(index[i] >= 0 && index[i] < PGSTAT_NUM_PROGRESS_PARAM);
- beentry->st_progress_param[index[i]] = val[i];
+ params[index[i]] = val[i];
}
PGSTAT_END_WRITE_ACTIVITY(beentry);
@@ -144,7 +188,7 @@ pgstat_progress_update_multi_param(int nparam, const int *index,
* pgstat_progress_end_command() -
*
* Reset st_progress_command (and st_progress_command_target) in own backend
- * entry. This signals the end of the command.
+ * entry. This signals the end of the command (or a sub-command).
*-----------
*/
void
@@ -155,11 +199,19 @@ pgstat_progress_end_command(void)
if (!beentry || !pgstat_track_activities)
return;
- if (beentry->st_progress_command == PROGRESS_COMMAND_INVALID)
- return;
-
PGSTAT_BEGIN_WRITE_ACTIVITY(beentry);
- beentry->st_progress_command = PROGRESS_COMMAND_INVALID;
- beentry->st_progress_command_target = InvalidOid;
+
+ if (beentry->st_progress_command2 != PROGRESS_COMMAND_INVALID)
+ {
+ Assert(beentry->st_progress_command != PROGRESS_COMMAND_INVALID);
+
+ beentry->st_progress_command2 = PROGRESS_COMMAND_INVALID;
+ }
+ else
+ {
+ beentry->st_progress_command = PROGRESS_COMMAND_INVALID;
+ beentry->st_progress_command_target = InvalidOid;
+ }
+
PGSTAT_END_WRITE_ACTIVITY(beentry);
}
diff --git a/src/backend/utils/activity/backend_status.c b/src/backend/utils/activity/backend_status.c
index d685fc5cd87..15675992415 100644
--- a/src/backend/utils/activity/backend_status.c
+++ b/src/backend/utils/activity/backend_status.c
@@ -284,6 +284,7 @@ pgstat_bestart_initial(void)
lbeentry.st_state = STATE_STARTING;
lbeentry.st_progress_command = PROGRESS_COMMAND_INVALID;
+ lbeentry.st_progress_command2 = PROGRESS_COMMAND_INVALID;
lbeentry.st_progress_command_target = InvalidOid;
lbeentry.st_query_id = INT64CONST(0);
lbeentry.st_plan_id = INT64CONST(0);
diff --git a/src/backend/utils/adt/pgstatfuncs.c b/src/backend/utils/adt/pgstatfuncs.c
index 7a9dfa9ba3b..c2257d900af 100644
--- a/src/backend/utils/adt/pgstatfuncs.c
+++ b/src/backend/utils/adt/pgstatfuncs.c
@@ -314,6 +314,7 @@ pg_stat_get_progress_info(PG_FUNCTION_ARGS)
Datum values[PG_STAT_GET_PROGRESS_COLS] = {0};
bool nulls[PG_STAT_GET_PROGRESS_COLS] = {0};
int i;
+ volatile int64 *params;
local_beentry = pgstat_get_local_beentry_by_index(curr_backend);
beentry = &local_beentry->backendStatus;
@@ -322,7 +323,11 @@ pg_stat_get_progress_info(PG_FUNCTION_ARGS)
* Report values for only those backends which are running the given
* command.
*/
- if (beentry->st_progress_command != cmdtype)
+ if (beentry->st_progress_command == cmdtype)
+ params = beentry->st_progress_param;
+ else if (beentry->st_progress_command2 == cmdtype)
+ params = beentry->st_progress_param2;
+ else
continue;
/* Value available to all callers */
@@ -334,7 +339,7 @@ pg_stat_get_progress_info(PG_FUNCTION_ARGS)
{
values[2] = ObjectIdGetDatum(beentry->st_progress_command_target);
for (i = 0; i < PGSTAT_NUM_PROGRESS_PARAM; i++)
- values[i + 3] = Int64GetDatum(beentry->st_progress_param[i]);
+ values[i + 3] = Int64GetDatum(params[i]);
}
else
{
diff --git a/src/include/utils/backend_status.h b/src/include/utils/backend_status.h
index a334e096e4a..f528f7abeec 100644
--- a/src/include/utils/backend_status.h
+++ b/src/include/utils/backend_status.h
@@ -169,6 +169,13 @@ typedef struct PgBackendStatus
Oid st_progress_command_target;
int64 st_progress_param[PGSTAT_NUM_PROGRESS_PARAM];
+ /*
+ * Some commands have a sub-command, e.g. REPACK (re)builds indexes. The
+ * subcommands are supposed to have the same target.
+ */
+ ProgressCommandType st_progress_command2;
+ int64 st_progress_param2[PGSTAT_NUM_PROGRESS_PARAM];
+
/* query identifier, optionally computed using post_parse_analyze_hook */
int64 st_query_id;
--
2.47.3
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
@ 2026-05-05 17:54 ` Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2 siblings, 1 reply; 32+ messages in thread
From: Sami Imseih @ 2026-05-05 17:54 UTC (permalink / raw)
To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>
Hi,
> Appreciate any feedback. Thank you!
I think it is valuable to show the index being processed. There is
really no other easy way to get this information except for pstack,
etc. I am +1 for the idea.
However, I am not sure that having a separate row for every parallel
worker is the right approach. The pg_stat_progress_* views are designed
to show progress per row. Each row represents one command with
meaningful progress counters (heap_blks_scanned, indexes_total,
indexes_processed, etc.). A parallel worker row would only show
current_index_relid and leader_pid with no actual progress information
of its own. That is status, not progress, and it does not fit the
view. Also, many columns would remain empty or redundant with the
leader's row.
Instead, could we aggregate the parallel worker information into the
leader's row. For example, an array of worker PIDs in one column and an
array of index relids in another?
--
Sami Imseih
Amazon Web Services (AWS)
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
@ 2026-06-29 15:31 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
0 siblings, 1 reply; 32+ messages in thread
From: Bharath Rupireddy @ 2026-06-29 15:31 UTC (permalink / raw)
To: Sami Imseih <samimseih@gmail.com>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>
Hi,
On Tue, May 5, 2026 at 10:54 AM Sami Imseih <samimseih@gmail.com> wrote:
>
> I think it is valuable to show the index being processed. There is
> really no other easy way to get this information except for pstack,
> etc. I am +1 for the idea.
Thanks for reviewing this!
> However, I am not sure that having a separate row for every parallel
> worker is the right approach. The pg_stat_progress_* views are designed
> to show progress per row. Each row represents one command with
> meaningful progress counters (heap_blks_scanned, indexes_total,
> indexes_processed, etc.). A parallel worker row would only show
> current_index_relid and leader_pid with no actual progress information
> of its own. That is status, not progress, and it does not fit the
> view. Also, many columns would remain empty or redundant with the
> leader's row.
>
> Instead, could we aggregate the parallel worker information into the
> leader's row. For example, an array of worker PIDs in one column and an
> array of index relids in another?
Thanks for the review. I read f1889729 and it looks like the
preference was to keep one command = one row, with workers feeding the
leader's row rather than showing up as separate rows. I want to stay
with that approach.
I considered having the leader set up a DSA that workers write into,
shown as an extra column on the leader's row. But that means new
shared memory whose handle has to be stored somewhere other backends
can find it, attached by the reader, and freed safely while a
monitoring query might still be reading it - that's a lot of work for
a small amount of per-worker data.
The simpler approach is to have workers report the index they're
currently on into their own st_progress_param[] slots, which already
exist per backend, along with their leader's pid. The view then groups
the worker entries under the matching leader's pid and shows only the
leader rows, so no new shared memory is needed. The current index oid
column has to be added to that array anyway for the non-parallel case,
which needs it just as much, and the parallel workers already have
their own slots to report into, so there's nothing extra to set up.
One thing to note is that pg_stat_get_progress_info('VACUUM') itself
would still return the worker rows, and the grouping happens in the
pg_stat_progress_vacuum view instead. I prefer keeping it in the view
rather than the function. The function is shared by all the progress
commands and only deals with raw params, while the leader pid grouping
is specific to VACUUM. This way the shared function stays unchanged
and the other progress views are not affected.
While I'm here, in the "vacuuming indexes" phase I also want to report
the total index pages to scan and the pages scanned so far, for the
index currently being vacuumed. On a large index this phase can run
for a long time with no way to tell whether it's making progress.
Does this direction sound reasonable, or do you see a reason to prefer
a different approach?
--
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
@ 2026-07-16 20:49 ` Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
0 siblings, 1 reply; 32+ messages in thread
From: Sami Imseih @ 2026-07-16 20:49 UTC (permalink / raw)
To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>
Hi,
> The simpler approach is to have workers report the index they're
> currently on into their own st_progress_param[] slots, which already
> exist per backend, along with their leader's pid. The view then groups
> the worker entries under the matching leader's pid and shows only the
> leader rows, so no new shared memory is needed.
Right, that is what I am thinking also.
> One thing to note is that pg_stat_get_progress_info('VACUUM') itself
> would still return the worker rows, and the grouping happens in the
> pg_stat_progress_vacuum view instead.
Yes, that would be the best way. Do the aggregation on the SQL level
using array_agg.
> While I'm here, in the "vacuuming indexes" phase I also want to report
> the total index pages to scan and the pages scanned so far, for the
> index currently being vacuumed. On a large index this phase can run
> for a long time with no way to tell whether it's making progress.
This was brought up when the "index progress" columns were being worked on, and
knowing the total was not possible for all index types [1]
> Does this direction sound reasonable, or do you see a reason to prefer
> a different approach?
Yes, I think the SQL level aggregation is sane.
[1] https://www.postgresql.org/message-id/CAH2-Wz%3D3JGBty%3D3tXoBoEYYwQNd7fXJuN9oPcnBAj3JYroBv3w%40mail...
--
Sami Imseih
Amazon Web Services (AWS)
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
@ 2026-07-17 19:24 ` Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
0 siblings, 1 reply; 32+ messages in thread
From: Sami Imseih @ 2026-07-17 19:24 UTC (permalink / raw)
To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>
Hi,
See the attached v2.
It aggregates on the leader (if parallel) at the SQL level with two new
columns: index_vacuum_pids and index_vacuum_oids. These are pid and oid
arrays respectively and are order-aligned. The leader is listed first when
it is itself processing an index; otherwise one of the workers is,
obviously, first.
For a serial vacuum, each array is a single value. The array fields will be
NULL if no index is currently being processed.
Using array typed columns is new for the pg_stat_progress_* views, none of
them expose arrays today, or any other pg_stat_* views. It's a bit unusual here,
but I think it's the natural fit.
Now, with this, parallel workers now register their own progress entry
so they show up as additional rows in pg_stat_get_progress_info('VACUUM').
The progress_vacuum view filters them out, but any code that reads the function
directly will now see an extra (mostly zeros) row per worker.
pg_stat_get_progress_info()
is not meant to be used directly, hence it's not documented for direct
use, so I think this is OK.
I also noticed in v1 that we weren't resetting the index after the vacuum
completed, and the index was being set in the wrong place. I fixed that. I
also removed the unnecessary typecast for InvalidOid.
Lastly, with this approach we no longer need the external function to
retrieve the leader pid.
What do you think?
--
Sami
Attachments:
[application/octet-stream] v2-0001-Report-indexes-currently-being-vacuumed-in-pg_sta.patch (10.1K, ../../CAA5RZ0sw8ZuAYt+bJ_VvJvcHVKe+=1tstC7oqAqUye9_3Ta6gg@mail.gmail.com/2-v2-0001-Report-indexes-currently-being-vacuumed-in-pg_sta.patch)
download | inline diff:
From aedf35bf0807dd174bf6e71a22e2a8c7a3790df7 Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Fri, 17 Apr 2026 22:11:20 +0000
Subject: [PATCH v2 1/1] Report indexes currently being vacuumed in
pg_stat_progress_vacuum
During the "vacuuming indexes" and "cleaning up indexes" phases,
pg_stat_progress_vacuum reported how many indexes were done but not
which ones were in progress, making it hard to tell which index is
slow on a table with many indexes.
Each process (the leader and any parallel workers) now reports the
OID of the index it is currently processing. pg_stat_progress_vacuum
exposes this on the vacuum's (leader) row as two position-aligned arrays,
index_vacuum_pids and indexes_being_vacuumed, so that
unnest(index_vacuum_pids, indexes_being_vacuumed) yields the
(pid, index) pairs.
---
doc/src/sgml/monitoring.sgml | 20 ++++++++++++++++++++
src/backend/access/heap/vacuumlazy.c | 16 ++++++++++++++++
src/backend/catalog/system_views.sql | 18 ++++++++++++++++--
src/backend/commands/vacuumparallel.c | 22 ++++++++++++++++++++++
src/include/commands/progress.h | 1 +
src/test/regress/expected/rules.out | 15 ++++++++++++---
6 files changed, 87 insertions(+), 5 deletions(-)
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index d1a20d001e9..f4d2fad2bce 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -7737,6 +7737,26 @@ FROM pg_stat_get_backend_idset() AS backendid;
</itemizedlist>
</para></entry>
</row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>index_vacuum_pids</structfield> <type>bigint[]</type>
+ </para>
+ <para>
+ A list of the process IDs, the leader process and any parallel workers,
+ currently vacuuming or cleaning up an index.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>index_vacuum_oids</structfield> <type>oid[]</type>
+ </para>
+ <para>
+ A list of the OIDs of the indexes being vacuumed or cleaned up by the
+ corresponding process in <structfield>index_vacuum_pids</structfield>.
+ </para></entry>
+ </row>
</tbody>
</tgroup>
</table>
diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index 39395aed0d5..3831215ba6a 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -3025,6 +3025,10 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3046,6 +3050,10 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /* Reset the current index relid */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ InvalidOid);
+
return istat;
}
@@ -3076,6 +3084,10 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3095,6 +3107,10 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /* Reset the current index relid */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ InvalidOid);
+
return istat;
}
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 090281a03dd..1a4f90cc6e7 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -1353,9 +1353,23 @@ CREATE VIEW pg_stat_progress_vacuum AS
CASE S.param13 WHEN 1 THEN 'manual'
WHEN 2 THEN 'autovacuum'
WHEN 3 THEN 'autovacuum_wraparound'
- ELSE NULL END AS started_by
+ ELSE NULL END AS started_by,
+ I.index_vacuum_pids AS index_vacuum_pids,
+ I.index_vacuum_oids AS index_vacuum_oids
FROM pg_stat_get_progress_info('VACUUM') AS S
- LEFT JOIN pg_database D ON S.datid = D.oid;
+ LEFT JOIN pg_database D ON S.datid = D.oid
+ LEFT JOIN pg_stat_activity A ON S.pid = A.pid,
+ LATERAL (
+ -- Aggregate the indexes being processed by this vacuum's leader
+ -- and its parallel workers (if any) into the single leader row.
+ SELECT array_agg(W.pid ORDER BY W.pid <> S.pid, W.pid) AS index_vacuum_pids,
+ array_agg(CAST(W.param14 AS oid) ORDER BY W.pid <> S.pid, W.pid) AS index_vacuum_oids
+ FROM pg_stat_get_progress_info('VACUUM') AS W
+ LEFT JOIN pg_stat_activity WA ON W.pid = WA.pid
+ WHERE COALESCE(WA.leader_pid, W.pid) = S.pid
+ AND W.param14 <> 0
+ ) I
+ WHERE A.leader_pid IS NULL;
CREATE VIEW pg_stat_progress_repack AS
SELECT
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 41cefcfde54..da7d02f3c9f 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -1076,6 +1076,11 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
IndexBulkDeleteResult *istat = NULL;
IndexBulkDeleteResult *istat_res;
IndexVacuumInfo ivinfo;
+ const int progress_index[] = {
+ PROGRESS_VACUUM_PHASE,
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID
+ };
+ int64 progress_val[2];
/*
* Update the pointer to the corresponding bulk-deletion result if someone
@@ -1097,6 +1102,13 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
pvs->indname = pstrdup(RelationGetRelationName(indrel));
pvs->status = indstats->status;
+ /* Report which index we're currently processing and the current phase */
+ progress_val[0] = (indstats->status == PARALLEL_INDVAC_STATUS_NEED_BULKDELETE)
+ ? PROGRESS_VACUUM_PHASE_VACUUM_INDEX
+ : PROGRESS_VACUUM_PHASE_INDEX_CLEANUP;
+ progress_val[1] = RelationGetRelid(indrel);
+ pgstat_progress_update_multi_param(2, progress_index, progress_val);
+
switch (indstats->status)
{
case PARALLEL_INDVAC_STATUS_NEED_BULKDELETE:
@@ -1144,6 +1156,10 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
pfree(pvs->indname);
pvs->indname = NULL;
+ /* Reset the current index relid */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ InvalidOid);
+
/*
* Call the parallel variant of pgstat_progress_incr_param so workers can
* report progress of index vacuum to the leader.
@@ -1307,6 +1323,9 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Prepare to track buffer usage during parallel execution */
InstrStartParallelQuery();
+ /* Register this worker for vacuum progress reporting */
+ pgstat_progress_start_command(PROGRESS_COMMAND_VACUUM, shared->relid);
+
/* Process indexes to perform vacuum/cleanup */
parallel_vacuum_process_safe_indexes(&pvs);
@@ -1326,6 +1345,9 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Pop the error context stack */
error_context_stack = errcallback.previous;
+ /* Unregister this worker from vacuum progress reporting */
+ pgstat_progress_end_command();
+
vac_close_indexes(nindexes, indrels, RowExclusiveLock);
table_close(rel, ShareUpdateExclusiveLock);
FreeAccessStrategy(pvs.bstrategy);
diff --git a/src/include/commands/progress.h b/src/include/commands/progress.h
index 2a12920c75f..bee311955b6 100644
--- a/src/include/commands/progress.h
+++ b/src/include/commands/progress.h
@@ -31,6 +31,7 @@
#define PROGRESS_VACUUM_DELAY_TIME 10
#define PROGRESS_VACUUM_MODE 11
#define PROGRESS_VACUUM_STARTED_BY 12
+#define PROGRESS_VACUUM_CURRENT_INDEX_RELID 13
/* Phases of vacuum (as advertised via PROGRESS_VACUUM_PHASE) */
#define PROGRESS_VACUUM_PHASE_SCAN_HEAP 1
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 6a3341356da..78a5dff0b82 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -2213,9 +2213,18 @@ pg_stat_progress_vacuum| SELECT s.pid,
WHEN 2 THEN 'autovacuum'::text
WHEN 3 THEN 'autovacuum_wraparound'::text
ELSE NULL::text
- END AS started_by
- FROM (pg_stat_get_progress_info('VACUUM'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
- LEFT JOIN pg_database d ON ((s.datid = d.oid)));
+ END AS started_by,
+ i.index_vacuum_pids,
+ i.index_vacuum_oids
+ FROM ((pg_stat_get_progress_info('VACUUM'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
+ LEFT JOIN pg_database d ON ((s.datid = d.oid)))
+ LEFT JOIN pg_stat_activity a ON ((s.pid = a.pid))),
+ LATERAL ( SELECT array_agg(w.pid ORDER BY (w.pid <> s.pid), w.pid) AS index_vacuum_pids,
+ array_agg((w.param14)::oid ORDER BY (w.pid <> s.pid), w.pid) AS index_vacuum_oids
+ FROM (pg_stat_get_progress_info('VACUUM'::text) w(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
+ LEFT JOIN pg_stat_activity wa ON ((w.pid = wa.pid)))
+ WHERE ((COALESCE(wa.leader_pid, w.pid) = s.pid) AND (w.param14 <> 0))) i
+ WHERE (a.leader_pid IS NULL);
pg_stat_recovery| SELECT promote_triggered,
last_replayed_read_lsn,
last_replayed_end_lsn,
--
2.50.1 (Apple Git-155)
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
@ 2026-08-04 22:20 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-12 21:45 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
0 siblings, 2 replies; 32+ messages in thread
From: Bharath Rupireddy @ 2026-08-04 22:20 UTC (permalink / raw)
To: Sami Imseih <samimseih@gmail.com>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
Hi,
On Fri, Jul 17, 2026 at 12:24 PM Sami Imseih <samimseih@gmail.com> wrote:
>
> See the attached v2.
>
> It aggregates on the leader (if parallel) at the SQL level with two new
> columns: index_vacuum_pids and index_vacuum_oids. These are pid and oid
> arrays respectively and are order-aligned. The leader is listed first when
> it is itself processing an index; otherwise one of the workers is,
> obviously, first.
> For a serial vacuum, each array is a single value. The array fields will be
> NULL if no index is currently being processed.
>
> Using array typed columns is new for the pg_stat_progress_* views, none of
> them expose arrays today, or any other pg_stat_* views. It's a bit unusual here,
> but I think it's the natural fit.
Thanks for the v2 patch, Sami!
I spent some more time thinking about using arrays here, and about the
one-row-per-command policy. I still think emitting the index OIDs and
worker PIDs as position-aligned arrays (like the existing
pg_stats.most_common_vals/most_common_freqs columns) is the simple
solution. I appreciate any thoughts or other ways here.
Please find attached the v3 patch. It ensures the current index is
reset after each index (so vacuuming heap and truncating heap show
NULL arrays with no stale relid), fixes the docs for the type of
index_vacuum_pids, adds a note in the docs about the arrays being
position-aligned, and rewords the commit message a bit.
Here's how the sample output looks.
Index vacuum:
pid | phase | index_vacuum_pids | index_vacuum_oids
------+-------------------+-------------------+-------------------
4955 | vacuuming indexes | {4955} | {16478}
(1 row)
Parallel index vacuum:
pid | phase | index_vacuum_pids | index_vacuum_oids
------+-------------------+-----------------------+---------------------------
5765 | vacuuming indexes | {5765,5768,5769,5770} | {16478,16479,16480,16481}
(1 row)
--
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com
Attachments:
[application/octet-stream] v3-0001-Report-indexes-being-vacuumed-in-pg_stat_progress_vacuum.patch (11.2K, ../../CALj2ACXhLVJh0OeN5W61_UvhyBxH3k98ffXhnRPYr8NsNbMZSw@mail.gmail.com/2-v3-0001-Report-indexes-being-vacuumed-in-pg_stat_progress_vacuum.patch)
download | inline diff:
From 76b329d7353e814765ab7183779e64d75cba5b7c Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Tue, 4 Aug 2026 20:24:51 +0000
Subject: [PATCH v3] Report indexes being vacuumed in pg_stat_progress_vacuum
Previously, during the "vacuuming indexes" and "cleaning up
indexes" phases, pg_stat_progress_vacuum reported how many
indexes had been processed but not which index was being
processed. On a table with many indexes, especially indexes of
different types, this made it hard to tell which index a slow or
stuck vacuum was working on.
This commit adds two columns to pg_stat_progress_vacuum,
index_vacuum_pids and index_vacuum_oids. Each process, the leader
and any parallel workers, reports the OID of the index it is
currently processing into its own progress array slot. The view
gathers these onto the leader's row as two arrays that are
aligned by position, so that unnest(index_vacuum_pids,
index_vacuum_oids) gives the (pid, index) pairs. For a serial
vacuum each array holds a single value, and both are null when no
index is being processed.
Because pg_stat_progress_vacuum reports one row per command, the
per-worker information is aggregated onto the leader's row rather
than shown as separate rows. This needs no new shared memory,
since the per-process progress array slots already exist.
Author: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Reviewed-by: Sami Imseih <samimseih@gmail.com>
Reviewed-by: Satyanarayana Narlapuram <satyanarlapuram@gmail.com>
Discussion: https://postgr.es/m/CALj2ACUgwSchK6jQ2CdKLBWUADTOE_zKdTff2Zg3E6hOuXKv-w@mail.gmail.com
---
doc/src/sgml/monitoring.sgml | 23 +++++++++++++++++++++++
src/backend/access/heap/vacuumlazy.c | 16 ++++++++++++++++
src/backend/catalog/system_views.sql | 21 +++++++++++++++++++--
src/backend/commands/vacuumparallel.c | 22 ++++++++++++++++++++++
src/include/commands/progress.h | 1 +
src/test/regress/expected/rules.out | 15 ++++++++++++---
6 files changed, 93 insertions(+), 5 deletions(-)
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index 099e9b6f4e9..0913a16ea9a 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -7892,6 +7892,29 @@ FROM pg_stat_get_backend_idset() AS backendid;
</itemizedlist>
</para></entry>
</row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>index_vacuum_pids</structfield> <type>integer[]</type>
+ </para>
+ <para>
+ The process IDs currently vacuuming or cleaning up an index (the
+ leader process and any parallel workers). Null when no index is
+ being processed.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>index_vacuum_oids</structfield> <type>oid[]</type>
+ </para>
+ <para>
+ The OIDs of the indexes being vacuumed or cleaned up, aligned by
+ position with <structfield>index_vacuum_pids</structfield>: the index
+ at each position is being processed by the process ID at the same
+ position. Null when no index is being processed.
+ </para></entry>
+ </row>
</tbody>
</tgroup>
</table>
diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index 39395aed0d5..3831215ba6a 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -3025,6 +3025,10 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3046,6 +3050,10 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /* Reset the current index relid */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ InvalidOid);
+
return istat;
}
@@ -3076,6 +3084,10 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3095,6 +3107,10 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /* Reset the current index relid */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ InvalidOid);
+
return istat;
}
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 090281a03dd..8f93ec58e47 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -1353,9 +1353,26 @@ CREATE VIEW pg_stat_progress_vacuum AS
CASE S.param13 WHEN 1 THEN 'manual'
WHEN 2 THEN 'autovacuum'
WHEN 3 THEN 'autovacuum_wraparound'
- ELSE NULL END AS started_by
+ ELSE NULL END AS started_by,
+ I.index_vacuum_pids AS index_vacuum_pids,
+ I.index_vacuum_oids AS index_vacuum_oids
FROM pg_stat_get_progress_info('VACUUM') AS S
- LEFT JOIN pg_database D ON S.datid = D.oid;
+ LEFT JOIN pg_database D ON S.datid = D.oid
+ LEFT JOIN pg_stat_activity A ON S.pid = A.pid,
+ LATERAL (
+ -- Aggregate the indexes being processed by this vacuum's leader
+ -- and its parallel workers (if any) into the single leader row.
+ -- Both arrays use the same ORDER BY so that they stay aligned by
+ -- position; the leader is listed first when it is itself
+ -- processing an index.
+ SELECT array_agg(W.pid ORDER BY W.pid <> S.pid, W.pid) AS index_vacuum_pids,
+ array_agg(CAST(W.param14 AS oid) ORDER BY W.pid <> S.pid, W.pid) AS index_vacuum_oids
+ FROM pg_stat_get_progress_info('VACUUM') AS W
+ LEFT JOIN pg_stat_activity WA ON W.pid = WA.pid
+ WHERE COALESCE(WA.leader_pid, W.pid) = S.pid
+ AND W.param14 <> 0
+ ) I
+ WHERE A.leader_pid IS NULL;
CREATE VIEW pg_stat_progress_repack AS
SELECT
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 767d162e578..01b11bf5803 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -1076,6 +1076,11 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
IndexBulkDeleteResult *istat = NULL;
IndexBulkDeleteResult *istat_res;
IndexVacuumInfo ivinfo;
+ const int progress_index[] = {
+ PROGRESS_VACUUM_PHASE,
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID
+ };
+ int64 progress_val[2];
/*
* Update the pointer to the corresponding bulk-deletion result if someone
@@ -1097,6 +1102,13 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
pvs->indname = pstrdup(RelationGetRelationName(indrel));
pvs->status = indstats->status;
+ /* Report which index we're currently processing and the current phase */
+ progress_val[0] = (indstats->status == PARALLEL_INDVAC_STATUS_NEED_BULKDELETE)
+ ? PROGRESS_VACUUM_PHASE_VACUUM_INDEX
+ : PROGRESS_VACUUM_PHASE_INDEX_CLEANUP;
+ progress_val[1] = RelationGetRelid(indrel);
+ pgstat_progress_update_multi_param(2, progress_index, progress_val);
+
switch (indstats->status)
{
case PARALLEL_INDVAC_STATUS_NEED_BULKDELETE:
@@ -1144,6 +1156,10 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
pfree(pvs->indname);
pvs->indname = NULL;
+ /* Reset the current index relid */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ InvalidOid);
+
/*
* Call the parallel variant of pgstat_progress_incr_param so workers can
* report progress of index vacuum to the leader.
@@ -1315,6 +1331,9 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Prepare to track buffer usage during parallel execution */
InstrStartParallelQuery();
+ /* Register this worker for vacuum progress reporting */
+ pgstat_progress_start_command(PROGRESS_COMMAND_VACUUM, shared->relid);
+
/* Process indexes to perform vacuum/cleanup */
parallel_vacuum_process_safe_indexes(&pvs);
@@ -1334,6 +1353,9 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Pop the error context stack */
error_context_stack = errcallback.previous;
+ /* Unregister this worker from vacuum progress reporting */
+ pgstat_progress_end_command();
+
vac_close_indexes(nindexes, indrels, RowExclusiveLock);
table_close(rel, ShareUpdateExclusiveLock);
FreeAccessStrategy(pvs.bstrategy);
diff --git a/src/include/commands/progress.h b/src/include/commands/progress.h
index 2a12920c75f..bee311955b6 100644
--- a/src/include/commands/progress.h
+++ b/src/include/commands/progress.h
@@ -31,6 +31,7 @@
#define PROGRESS_VACUUM_DELAY_TIME 10
#define PROGRESS_VACUUM_MODE 11
#define PROGRESS_VACUUM_STARTED_BY 12
+#define PROGRESS_VACUUM_CURRENT_INDEX_RELID 13
/* Phases of vacuum (as advertised via PROGRESS_VACUUM_PHASE) */
#define PROGRESS_VACUUM_PHASE_SCAN_HEAP 1
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 6a3341356da..78a5dff0b82 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -2213,9 +2213,18 @@ pg_stat_progress_vacuum| SELECT s.pid,
WHEN 2 THEN 'autovacuum'::text
WHEN 3 THEN 'autovacuum_wraparound'::text
ELSE NULL::text
- END AS started_by
- FROM (pg_stat_get_progress_info('VACUUM'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
- LEFT JOIN pg_database d ON ((s.datid = d.oid)));
+ END AS started_by,
+ i.index_vacuum_pids,
+ i.index_vacuum_oids
+ FROM ((pg_stat_get_progress_info('VACUUM'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
+ LEFT JOIN pg_database d ON ((s.datid = d.oid)))
+ LEFT JOIN pg_stat_activity a ON ((s.pid = a.pid))),
+ LATERAL ( SELECT array_agg(w.pid ORDER BY (w.pid <> s.pid), w.pid) AS index_vacuum_pids,
+ array_agg((w.param14)::oid ORDER BY (w.pid <> s.pid), w.pid) AS index_vacuum_oids
+ FROM (pg_stat_get_progress_info('VACUUM'::text) w(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
+ LEFT JOIN pg_stat_activity wa ON ((w.pid = wa.pid)))
+ WHERE ((COALESCE(wa.leader_pid, w.pid) = s.pid) AND (w.param14 <> 0))) i
+ WHERE (a.leader_pid IS NULL);
pg_stat_recovery| SELECT promote_triggered,
last_replayed_read_lsn,
last_replayed_end_lsn,
--
2.47.3
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
@ 2026-08-12 21:45 ` Sami Imseih <samimseih@gmail.com>
2026-08-13 23:30 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
1 sibling, 1 reply; 32+ messages in thread
From: Sami Imseih @ 2026-08-12 21:45 UTC (permalink / raw)
To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
> Please find attached the v3 patch. It ensures the current index is
> reset after each index (so vacuuming heap and truncating heap show
> NULL arrays with no stale relid), fixes the docs for the type of
> index_vacuum_pids, adds a note in the docs about the arrays being
> position-aligned, and rewords the commit message a bit.
Thanks for the updates in v3.
It turns out, to my surprise, that leader_pid can be NULL if the user
querying pg_stat_progress_vacuum does not have proper privileges, either
pg_read_all_stats or membership in the role running the vacuum.
Here is the case. "foo" is created with
```
CREATE ROLE foo LOGIN;
GRANT CONNECT ON DATABASE postgres TO foo;
```
A superuser aggregates correctly on leader_pid, because the workers
emit leader_pid in pg_stat_activity.
```
pid | phase | index_vacuum_pids | index_vacuum_oids
-------+-------------------+---------------------+---------------------
23492 | vacuuming indexes | {23492,23570,23571} | {16389,16390,16391}
(1 row)
pid | leader_pid | backend_type | state | query
-------+------------+-----------------+--------+------------------------------
23492 | | client backend | active | VACUUM (PARALLEL 4) vac_demo;
23570 | 23492 | parallel worker | active | VACUUM (PARALLEL 4) vac_demo;
23571 | 23492 | parallel worker | active | VACUUM (PARALLEL 4) vac_demo;
(3 rows)
```
But "foo" cannot, because leader_pid is NULL for this user, so the
aggregation falls apart and each worker emits its own row in
pg_stat_progress_vacuum.
pid | phase | index_vacuum_pids | index_vacuum_oids
-------+-------+-------------------+-------------------
23492 | | |
23570 | | |
23571 | | |
(3 rows)
I think for this patch we should drop the reliance on pg_stat_activity
and have pg_stat_get_progress_info() emit leader_pid directly and
unconditionally. The aggregation in the view then works regardless of
the caller's
privileges. It is the same lockGroupLeader value pg_stat_activity
already computes.
```
pg_stat_get_progress_info(PG_FUNCTION_ARGS)
{
-#define PG_STAT_GET_PROGRESS_COLS PGSTAT_NUM_PROGRESS_PARAM + 3
+#define PG_STAT_GET_PROGRESS_COLS PGSTAT_NUM_PROGRESS_PARAM + 4
int num_backends = pgstat_fetch_stat_numbackends();
int curr_backend;
char *cmd = text_to_cstring(PG_GETARG_TEXT_PP(0));
@@ -373,6 +373,7 @@ pg_stat_get_progress_info(PG_FUNCTION_ARGS)
{
LocalPgBackendStatus *local_beentry;
PgBackendStatus *beentry;
+ PGPROC *proc;
Datum values[PG_STAT_GET_PROGRESS_COLS] = {0};
bool nulls[PG_STAT_GET_PROGRESS_COLS] = {0};
int i;
@@ -391,6 +392,23 @@ pg_stat_get_progress_info(PG_FUNCTION_ARGS)
values[0] = Int32GetDatum(beentry->st_procpid);
values[1] = ObjectIdGetDatum(beentry->st_databaseid);
+ proc = BackendPidGetProc(beentry->st_procpid);
+ if (proc != NULL && proc->lockGroupLeader != NULL &&
+ proc->lockGroupLeader->pid != beentry->st_procpid)
+ values[PGSTAT_NUM_PROGRESS_PARAM + 3] =
+ Int32GetDatum(proc->lockGroupLeader->pid);
+ else
+ values[PGSTAT_NUM_PROGRESS_PARAM + 3] =
Int32GetDatum(0);
+
```
That leaves a more interesting question in my mind, which is why
pg_stat_activity puts leader_pid behind permissions at all. It should be
treated just like pid.
There is probably a larger discussion around what should and should not
be permission controlled in pg_stat_activity, and I could not find a
consistent rule. For example, we do not permission control application_name,
which is user controlled free text, yet we do permission control
query_id, which
is not permission controlled elsewhere such as pg_stat_statements. We probably
need a separate thread to clearly lay out the principles for this.
As far as this patch goes, I don't think it should be blocked and it should
continue to emit the leader_pid, but with the idea I shared above
instead of joining with pg_stat_activity.
thoughts?
--
Sami Imseih
Amazon Web Services (AWS)
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-12 21:45 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
@ 2026-08-13 23:30 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 00:47 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
0 siblings, 1 reply; 32+ messages in thread
From: Bharath Rupireddy @ 2026-08-13 23:30 UTC (permalink / raw)
To: Sami Imseih <samimseih@gmail.com>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>; Michael Paquier <michael@paquier.xyz>
Hi,
On Wed, Aug 12, 2026 at 2:45 PM Sami Imseih <samimseih@gmail.com> wrote:
>
> Thanks for the updates in v3.
Thanks for taking a look at it.
> It turns out, to my surprise, that leader_pid can be NULL if the user
> querying pg_stat_progress_vacuum does not have proper privileges, either
> pg_read_all_stats or membership in the role running the vacuum.
Nice catch!
> I think for this patch we should drop the reliance on pg_stat_activity
> and have pg_stat_get_progress_info() emit leader_pid directly and
> unconditionally. The aggregation in the view then works regardless of
> the caller's privileges.
That's one option. There's another option that I originally proposed
upthread, which is to track the leader_pid directly in the progress
report.
> That leaves a more interesting question in my mind, which is why
> pg_stat_activity puts leader_pid behind permissions at all. It should be
> treated just like pid.
Yes, I looked at the commit (b025f32e0) and the discussion. I think
one of the main reasons was to not let unprivileged users take
ProcArrayLock and scan over the entire PGPROC array via
BackendPidGetProc().
To summarize, we have three options:
1/ Make pg_stat_get_progress_info() report the leader_pid like
pg_stat_get_activity() does.
2/ Make pg_stat_get_activity() itself report the leader_pid just like the pid.
3/ Track the leader_pid via a new progress report param (like the v1
did upthread).
(1) and (2) will let unprivileged users take ProcArrayLock and scan
the entire PGPROC array. (3) although it eats up a new slot in the
progress report, gives the leader pid almost for free. I prefer (3)
for its simplicity and without any additional risks.
Adding Michael Paquier to the thread for any thoughts on this.
> There is probably a larger discussion around what should and should not
> be permission controlled in pg_stat_activity, and I could not find a
> consistent rule. For example, we do not permission control application_name,
> which is user controlled free text, yet we do permission control
> query_id, which
> is not permission controlled elsewhere such as pg_stat_statements. We probably
> need a separate thread to clearly lay out the principles for this.
The rule here seems simple. The pid or leader_pid by itself is not
something that requires permission controls, it is what the users will
do to get it that matters. I think this applies to all other params as
well.
Thoughts?
--
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-12 21:45 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-13 23:30 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
@ 2026-08-14 00:47 ` Sami Imseih <samimseih@gmail.com>
2026-08-14 03:33 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
0 siblings, 1 reply; 32+ messages in thread
From: Sami Imseih @ 2026-08-14 00:47 UTC (permalink / raw)
To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>; Michael Paquier <michael@paquier.xyz>
>
> Hi,
3/ Track the leader_pid via a new progress report param (like the v1
> did upthread).
Yes, this seems like the best way. Have the workers report their leader.
> (1) and (2) will let unprivileged users take ProcArrayLock and scan
> the entire PGPROC array. (3) although it eats up a new slot in the
> progress report, gives the leader pid almost for free. I prefer (3)
> for its simplicity and without any additional risks.
>
> Adding Michael Paquier to the thread for any thoughts on this.
>
> > There is probably a larger discussion around what should and should not
> > be permission controlled in pg_stat_activity, and I could not find a
> > consistent rule. For example, we do not permission control
> application_name,
> > which is user controlled free text, yet we do permission control
> > query_id, which
> > is not permission controlled elsewhere such as pg_stat_statements. We
> probably
> > need a separate thread to clearly lay out the principles for this.
>
> The rule here seems simple. The pid or leader_pid by itself is not
> something that requires permission controls, it is what the users will
> do to get it that matters. I think this applies to all other params as
> well.
>
> Thoughts?
I agree. At least the leader_pid should not be permission controlled and we
should
be able to perform the aggregation as we do in v3- at the sql level. Other
fields like
relid, phase, etc. sit behind permission controls and should remain that
way. If there
is different opinion for those fields, that is a separate discussion.
WDYT?
--
Sami
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-12 21:45 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-13 23:30 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 00:47 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
@ 2026-08-14 03:33 ` Michael Paquier <michael@paquier.xyz>
0 siblings, 0 replies; 32+ messages in thread
From: Michael Paquier @ 2026-08-14 03:33 UTC (permalink / raw)
To: Sami Imseih <samimseih@gmail.com>; +Cc: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
On Thu, Aug 13, 2026 at 07:47:06PM -0500, Sami Imseih wrote:
> I agree. At least the leader_pid should not be permission controlled and we
> should
> be able to perform the aggregation as we do in v3- at the sql level. Other
> fields like
> relid, phase, etc. sit behind permission controls and should remain that
> way. If there
> is different opinion for those fields, that is a separate discussion.
>
> WDYT?
That's debatable perhaps, but the leader PID is in the same kind of
category as the wait events: no information derived from a PGPROC
entry should be viewable except for a role with pg_read_all_stats
privileges or if a role is a member of the role whose information is
queried.
--
Michael
Attachments:
[application/pgp-signature] signature.asc (832B, ../../an6Mbz9RUlAxKslz@paquier.xyz/2-signature.asc)
download
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
@ 2026-08-14 03:48 ` Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-19 18:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
1 sibling, 2 replies; 32+ messages in thread
From: Michael Paquier @ 2026-08-14 03:48 UTC (permalink / raw)
To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: Sami Imseih <samimseih@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
On Tue, Aug 04, 2026 at 03:20:00PM -0700, Bharath Rupireddy wrote:
> I spent some more time thinking about using arrays here, and about the
> one-row-per-command policy. I still think emitting the index OIDs and
> worker PIDs as position-aligned arrays (like the existing
> pg_stats.most_common_vals/most_common_freqs columns) is the simple
> solution. I appreciate any thoughts or other ways here.
>
> Please find attached the v3 patch. It ensures the current index is
> reset after each index (so vacuuming heap and truncating heap show
> NULL arrays with no stale relid), fixes the docs for the type of
> index_vacuum_pids, adds a note in the docs about the arrays being
> position-aligned, and rewords the commit message a bit.
+ ELSE NULL END AS started_by,
+ I.index_vacuum_pids AS index_vacuum_pids,
+ I.index_vacuum_oids AS index_vacuum_oids
FROM pg_stat_get_progress_info('VACUUM') AS S
- LEFT JOIN pg_database D ON S.datid = D.oid;
+ LEFT JOIN pg_database D ON S.datid = D.oid
+ LEFT JOIN pg_stat_activity A ON S.pid = A.pid,
+ LATERAL (
Exposing the information of an index a worker is processing is a good
idea, but I think that this choice lacks a long-term vision. I think
that we should expose one row for each worker rather than an array of
PIDs and index OIDs in the row of a leader. The main issue for me is
the granularity of the information provided, where it would actually
make sense to provide more information for each worker. Choosing how
an index clean works in vacuum for parallel workers is an
implementation choice, where we could think about approaches like:
- Distribute the workload of one index across N workers (for a 1TB
index, spawn N workers each sharing 1/N TB of data to clean)
- Have each worker do one index.
- Or more strategies, etc.
My point is not the strategy or the design we choose, which could vary
depending on an index AM. It's that for any design, any strategy or
any index AM, at the end it is going to be way more important for the
end-user how *each* individual worker behaves. One thing could be for
example reusing heap_blks_total and heap_blks_scanned for indexes, so
as it is possible how much each worker has done (let's perhaps rename
them). Being able to map a leader with its worker is an information
already provided by pg_stat_activity, adding this information in the
progress view seems unnecessary for me to add here as a JOIN is
already able to solve that anyway. If extra SQL knowledge is
necessary, that's more a documentation problem to me, adding more
fields for data that's already available is just more information
bloat.
--
Michael
Attachments:
[application/pgp-signature] signature.asc (832B, ../../an6P_WQNCM4P843q@paquier.xyz/2-signature.asc)
download
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
@ 2026-08-14 21:26 ` Sami Imseih <samimseih@gmail.com>
2026-08-14 23:35 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-20 02:37 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
1 sibling, 2 replies; 32+ messages in thread
From: Sami Imseih @ 2026-08-14 21:26 UTC (permalink / raw)
To: Michael Paquier <michael@paquier.xyz>; +Cc: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
Thanks for the feedback, Michael!
The view already reports indexes_total and indexes_processed, how many
indexes are done. These columns add which ones are in progress, and by
which PID.
> Exposing the information of an index a worker is processing is a good
> idea, but I think that this choice lacks a long-term vision. I think
> that we should expose one row for each worker rather than an array of
> PIDs and index OIDs in the row of a leader.
I put the PID and OID in the leader's row as two position-aligned arrays
precisely to keep the shape of the view, one row per command. The docs
describe pg_stat_progress_vacuum as "one row for each backend (including
autovacuum worker processes) that is currently vacuuming". I read that
as one row for the backend that launched the operation. A parallel
worker isn't running a command, it's a consequence of the configuration,
so it belongs in the leader's row, not a row of its own.
> The main issue for me is the granularity of the information provided,
> where it would actually make sense to provide more information for
> each worker. Choosing how an index clean works in vacuum for parallel
> workers is an implementation choice, where we could think about
> approaches like:
> - Distribute the workload of one index across N workers (for a 1TB
> index, spawn N workers each sharing 1/N TB of data to clean)
> - Have each worker do one index.
> - Or more strategies, etc.
The two arrays already roll up what we care about, which PID is on which
index, into the leader's row, and the approach we pick only changes
whether an OID repeats.
Today, index vacuuming is one PID per OID. An index is claimed in full by
one process and never scanned by two at once, so the OIDs are distinct.
```
leader_pid | vacuum_pids | vacuum_oids
------------+---------------+---------------------
100 | {100,200,300} | {10000,10001,10002}
```
More than one PID on a single OID, whether that's distributing one index
across workers as you describe, or parallel heap vacuum [1] which is
still being discussed, is handled by repeating the OID. The phase says
what the OID is, here a table's relid rather than an index.
```
leader_pid | phase | vacuum_pids | vacuum_oids
------------+----------------+-------------------+-----------------------
100 | vacuuming heap | {100,200,300,400} | {9000,9000,9000,9000}
```
On naming, we should probablt call the columns vacuum_pids and vacuum_oids
rather than index_vacuum_*, since the phase says what they are for.
So this doesn't hamstring us, and it doesn't need a row per worker.
> One thing could be for example reusing heap_blks_total and
> heap_blks_scanned for indexes, so as it is possible how much each
> worker has done (let's perhaps rename them).
Right, there's a need here. For some indexes we know how far we need to
scan, like btree, and for some we can't, like GIN. This was discussed
before for index scan progress [2].
Also, reusing heap_blks_* isn't the right interface for it. They hold
how far the heap got, and even though they stop advancing once we're
vacuuming indexes, a user still wants to know at any given moment just
how far the table has been vacuumed. Double purposing the same field to
count index blocks would overwrite it, which makes monitoring this field
impractical.
If we do expose per-worker progress, a separate view is the better home.
Most columns here are the leader's or shared, so a per-worker row would
be mostly empty anyway.
> Being able to map a leader with its worker is an information already
> provided by pg_stat_activity, adding this information in the progress
> view seems unnecessary for me to add here as a JOIN is already able to
> solve that anyway.
This isn't really about mapping a leader to its workers. A JOIN with
pg_stat_activity relates the PIDs, but it can't tell you which index
each worker is on, and that's what these two columns add.
[1] https://www.postgresql.org/message-id/flat/CAD21AoAEfCNv-GgaDheDJ%2Bs-p_Lv1H24AiJeNoPGCmZNSwL1YA%40m...
[2] https://www.postgresql.org/message-id/CAH2-Wz=3JGBty=3tXoBoEYYwQNd7fXJuN9oPcnBAj3JYroBv3w@mail.gmail...
--
Sami Imseih
Amazon Web Services (AWS)
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
@ 2026-08-14 23:35 ` Sami Imseih <samimseih@gmail.com>
1 sibling, 0 replies; 32+ messages in thread
From: Sami Imseih @ 2026-08-14 23:35 UTC (permalink / raw)
To: Michael Paquier <michael@paquier.xyz>; +Cc: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
> If we do expose per-worker progress, a separate view is the better home.
> Most columns here are the leader's or shared, so a per-worker row would
> be mostly empty anyway.
A different idea other than aggregating the pids and oid's into a list
as is currently
being proposed, would be to have a "pg_stat_progress_vacuum_worker" view,
which initially will be 4 columns:
"pid"
"leader_pid"
"phase"
"oid"
and it will have a row for every worker ( or leader ) and the "oid" they are
processing, which could be a index ( or a heap if we get to that point of
parallel heap vacuum ).
My hesitation is this is not really a "progress" view, since it's not showing
progress related data. Just thought I'll put this out there for discussion
as well.
--
Sami
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
@ 2026-08-20 02:37 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-20 23:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
1 sibling, 1 reply; 32+ messages in thread
From: Bharath Rupireddy @ 2026-08-20 02:37 UTC (permalink / raw)
To: Sami Imseih <samimseih@gmail.com>; +Cc: Michael Paquier <michael@paquier.xyz>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
Hi,
Thanks Sami for the thoughts. I addressed most of these upthread [1],
responding to the remaining ones here.
[1] https://www.postgresql.org/message-id/CALj2ACX6gyBmQfaqbCsycDmaSPbq1%3DiPJw1OqUT%2BaLqaKjW8dQ%40mail...
On Fri, Aug 14, 2026 at 2:26 PM Sami Imseih <samimseih@gmail.com> wrote:
>
> The two arrays already roll up what we care about, which PID is on which
> index, into the leader's row, and the approach we pick only changes
> whether an OID repeats.
>
> Today, index vacuuming is one PID per OID. An index is claimed in full by
> one process and never scanned by two at once, so the OIDs are distinct.
>
> ```
> leader_pid | vacuum_pids | vacuum_oids
> ------------+---------------+---------------------
> 100 | {100,200,300} | {10000,10001,10002}
> ```
Right, and this holds only for BTree. Some other index AM might
implement intra-parallel index vacuum (vacuuming one index with
multiple workers), in which case OIDs would repeat.
> More than one PID on a single OID, whether that's distributing one index
> across workers as you describe, or parallel heap vacuum [1] which is
> still being discussed, is handled by repeating the OID. The phase says
> what the OID is, here a table's relid rather than an index.
>
> ```
> leader_pid | phase | vacuum_pids | vacuum_oids
> ------------+----------------+-------------------+-----------------------
> 100 | vacuuming heap | {100,200,300,400} | {9000,9000,9000,9000}
> ```
>
> [1] https://www.postgresql.org/message-id/flat/CAD21AoAEfCNv-GgaDheDJ%2Bs-p_Lv1H24AiJeNoPGCmZNSwL1YA%40m...
I haven't thought about parallel heap vacuum in depth yet and will do
that as part of that thread. Quick thoughts. One row per worker lets
us report heap_blks_* per worker naturally. Alternatively, we could
report a single overall value in the leader's heap_blks_*, the way
parallel CREATE INDEX does today, tracking total scan progress across
all workers rather than any one worker's share.
> > One thing could be for example reusing heap_blks_total and
> > heap_blks_scanned for indexes, so as it is possible how much each
> > worker has done (let's perhaps rename them).
>
> Right, there's a need here. For some indexes we know how far we need to
> scan, like btree, and for some we can't, like GIN. This was discussed
> before for index scan progress [2].
>
> [2] https://www.postgresql.org/message-id/CAH2-Wz=3JGBty=3tXoBoEYYwQNd7fXJuN9oPcnBAj3JYroBv3w@mail.gmail...
Correct, but today BTree scan progress is supported via create index
progress reporting (see my response upthread around
PROGRESS_SCAN_BLOCKS_TOTAL/DONE).
> > If we do expose per-worker progress, a separate view is the better home.
> > Most columns here are the leader's or shared, so a per-worker row would
> > be mostly empty anyway.
>
> A different idea other than aggregating the pids and oid's into a list
> as is currently
> being proposed, would be to have a "pg_stat_progress_vacuum_worker" view,
> which initially will be 4 columns:
>
> "pid"
> "leader_pid"
> "phase"
> "oid"
>
> and it will have a row for every worker ( or leader ) and the "oid" they are
> processing, which could be a index ( or a heap if we get to that point of
> parallel heap vacuum ).
I don't think a separate view is the right approach at least for two
reasons. One, to know the progress of a vacuum one has to query two
views and relate them. Two, since the leader itself participates in
vacuuming indexes alongside workers (and will also for parallel heap
vacuum), splitting the same command's progress into two views adds
complexity. Keeping everything in a single view with one row per
worker (as discussed upthread) is simpler. Some fields would be null
on worker rows, but documenting this should be sufficient. That said,
I'm open to hear more thoughts on this.
--
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-20 02:37 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
@ 2026-08-20 23:10 ` Michael Paquier <michael@paquier.xyz>
2026-08-20 23:58 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-25 16:34 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
0 siblings, 2 replies; 32+ messages in thread
From: Michael Paquier @ 2026-08-20 23:10 UTC (permalink / raw)
To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: Sami Imseih <samimseih@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
On Wed, Aug 19, 2026 at 07:37:00PM -0700, Bharath Rupireddy wrote:
> I don't think a separate view is the right approach at least for two
> reasons. One, to know the progress of a vacuum one has to query two
> views and relate them. Two, since the leader itself participates in
> vacuuming indexes alongside workers (and will also for parallel heap
> vacuum), splitting the same command's progress into two views adds
> complexity. Keeping everything in a single view with one row per
> worker (as discussed upthread) is simpler. Some fields would be null
> on worker rows, but documenting this should be sufficient. That said,
> I'm open to hear more thoughts on this.
Having a single progress view feels like the natural approach here,
for both the leader and the workers. The leader triggers the
existence of the workers, but both leader and workers may finish by
doing the same job as there could be usually little meaning for a
leader to stand idle, waiting for all the workers to do the work.
Such choices are implementation-agnostic, of course; we should not
lock ourselves.
--
Michael
Attachments:
[application/pgp-signature] signature.asc (832B, ../../aoeJa9Ap4_MyFX3_@paquier.xyz/2-signature.asc)
download
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-20 02:37 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-20 23:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
@ 2026-08-20 23:58 ` Sami Imseih <samimseih@gmail.com>
1 sibling, 0 replies; 32+ messages in thread
From: Sami Imseih @ 2026-08-20 23:58 UTC (permalink / raw)
To: Michael Paquier <michael@paquier.xyz>; +Cc: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
> On Wed, Aug 19, 2026 at 07:37:00PM -0700, Bharath Rupireddy wrote:
> > I don't think a separate view is the right approach at least for two
> > reasons. One, to know the progress of a vacuum one has to query two
> > views and relate them. Two, since the leader itself participates in
> > vacuuming indexes alongside workers (and will also for parallel heap
> > vacuum), splitting the same command's progress into two views adds
> > complexity. Keeping everything in a single view with one row per
> > worker (as discussed upthread) is simpler. Some fields would be null
> > on worker rows, but documenting this should be sufficient. That said,
> > I'm open to hear more thoughts on this.
>
> Having a single progress view feels like the natural approach here,
> for both the leader and the workers. The leader triggers the
> existence of the workers, but both leader and workers may finish by
> doing the same job as there could be usually little meaning for a
> leader to stand idle, waiting for all the workers to do the work.
> Such choices are implementation-agnostic, of course; we should not
> lock ourselves.
This sounds to me that we are ok with breaking the principle design of the
progress views, at least how I understand it, which is that we only have
one row associated with the backend from which the user issued command.
Background workers that are started as a result of that command, and are
transient during the life of the command, don’t fit into that.
Also, If we include workers in a separate row, we will need to mix stats
that are aggregate of the entire command with stats that are specific to
the work that leader done on its own. what would heap_blks_* refer to in
this case? The aggregate of heap blocks accessed by the leader and workers
-or— just the blocks accessed by the leader?
I think monitoring tools/users consuming this view can better deal with the
separation of worker level details ( worker here is also the leader ) in
one view and the aggregate/high level information in the separate existing
view(s) much better than having to deal with documented caveats about what
the stats mean.
--
Sami
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-20 02:37 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-20 23:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
@ 2026-08-25 16:34 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 02:23 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
1 sibling, 1 reply; 32+ messages in thread
From: Bharath Rupireddy @ 2026-08-25 16:34 UTC (permalink / raw)
To: Michael Paquier <michael@paquier.xyz>; +Cc: Sami Imseih <samimseih@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
Hi,
On Thu, Aug 20, 2026 at 4:10 PM Michael Paquier <michael@paquier.xyz> wrote:
>
> On Wed, Aug 19, 2026 at 07:37:00PM -0700, Bharath Rupireddy wrote:
> > I don't think a separate view is the right approach at least for two
> > reasons. One, to know the progress of a vacuum one has to query two
> > views and relate them. Two, since the leader itself participates in
> > vacuuming indexes alongside workers (and will also for parallel heap
> > vacuum), splitting the same command's progress into two views adds
> > complexity. Keeping everything in a single view with one row per
> > worker (as discussed upthread) is simpler. Some fields would be null
> > on worker rows, but documenting this should be sufficient. That said,
> > I'm open to hear more thoughts on this.
>
> Having a single progress view feels like the natural approach here,
> for both the leader and the workers. The leader triggers the
> existence of the workers, but both leader and workers may finish by
> doing the same job as there could be usually little meaning for a
> leader to stand idle, waiting for all the workers to do the work.
> Such choices are implementation-agnostic, of course; we should not
> lock ourselves.
Thanks for reviewing. After thinking about this for a while, I still
think one row per worker is the better choice, with the tradeoff of
some fields being null on worker rows (similar to leader_pid in
pg_stat_activity). The docs also say "the view will contain one row
for each backend that is currently vacuuming," and in that sense,
parallel workers launched for index vacuum are essentially doing
vacuum too.
Here's what I have. 0001 reports the current index being vacuumed, and
0002 reports the total blocks and done blocks for B-tree indexes. This
helps track vacuum progress for large indexes in production.
Please have a look.
Thanks, Sami, for sharing thoughts. I'm open to hearing from others on
the separate view for worker progress reporting approach.
--
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com
Attachments:
[application/octet-stream] v4-0001-Report-the-index-being-vacuumed-in-pg_stat_progre.patch (12.8K, ../../CALj2ACX-49BrUTXaM3HeWVaZr+zGkbcLNJcrQ_dgjYF9AtW-kg@mail.gmail.com/2-v4-0001-Report-the-index-being-vacuumed-in-pg_stat_progre.patch)
download | inline diff:
From d7062b5cf15f208332602d277806bc694c348787 Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Mon, 24 Aug 2026 23:22:22 +0000
Subject: [PATCH v4 1/2] Report the index being vacuumed in
pg_stat_progress_vacuum.
Previously, pg_stat_progress_vacuum showed the total and
processed index counts, but not which index a backend was
actually working on. On a table with many indexes of different
access methods, this made it hard to tell which index a slow or
stuck vacuum was spending its time on.
This commit adds a current_index_relid column reporting the OID
of the index a backend is currently vacuuming or cleaning up. It
is set before each index is processed and reset once that index
is done, so it does not report a stale value after the phase
moves on.
During parallel index vacuum the leader also participates, and
each participant processes a different set of indexes. That
per-index progress cannot be collapsed into a single leader row
without losing the detail that matters, so each participant, the
leader and the workers alike, now reports its own row and shows
the index it is processing.
To map the leader to its workers, the view provides a leader_pid
column, taken from pg_stat_activity by joining on the process ID.
It is null on the leader row (and for a non-parallel vacuum) and
holds the leader's PID on the worker rows. On a worker row only
the phase, the current index and the leader PID are meaningful.
The other columns track table-wide progress that only the leader
maintains. The documentation covers this, along with the fact
that an unprivileged user sees only the process ID for backends
it cannot otherwise inspect.
Author: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Reviewed-by: Michael Paquier <michael@paquier.xyz>
Reviewed-by: Sami Imseih <samimseih@gmail.com>
Reviewed-by: Satyanarayana Narlapuram <satyanarlapuram@gmail.com>
Discussion: https://postgr.es/m/CALj2ACUgwSchK6jQ2CdKLBWUADTOE_zKdTff2Zg3E6hOuXKv-w@mail.gmail.com
---
doc/src/sgml/monitoring.sgml | 56 ++++++++++++++++++++++++++-
src/backend/access/heap/vacuumlazy.c | 16 ++++++++
src/backend/catalog/system_views.sql | 7 +++-
src/backend/commands/vacuumparallel.c | 25 ++++++++++++
src/include/commands/progress.h | 1 +
src/test/regress/expected/rules.out | 9 +++--
6 files changed, 107 insertions(+), 7 deletions(-)
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index 32cb6fdbd76..0d3446d6e86 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -7653,8 +7653,11 @@ FROM pg_stat_get_backend_idset() AS backendid;
<para>
Whenever <command>VACUUM</command> is running, the
<structname>pg_stat_progress_vacuum</structname> view will contain
- one row for each backend (including autovacuum worker processes) that is
- currently vacuuming. The tables below describe the information
+ one row for each backend (including autovacuum worker processes and
+ parallel workers launched for
+ <link linkend="sql-vacuum">parallel vacuum</link>) that is currently
+ vacuuming; see the note following the view for the columns reported on
+ parallel worker rows. The tables below describe the information
that will be reported and provide information about how to interpret it.
Progress for <command>VACUUM FULL</command> commands is reported via
<structname>pg_stat_progress_cluster</structname>
@@ -7910,10 +7913,59 @@ FROM pg_stat_get_backend_idset() AS backendid;
</itemizedlist>
</para></entry>
</row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>current_index_relid</structfield> <type>oid</type>
+ </para>
+ <para>
+ If <command>VACUUM</command> is currently processing an index, this
+ column shows the OID of the index being vacuumed. The value is set
+ when the phase is <literal>vacuuming indexes</literal> or
+ <literal>cleaning up indexes</literal> and is reset to 0 when all
+ indexes have been processed. During parallel index vacuum, each
+ parallel worker row shows the index that particular worker is
+ processing.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>leader_pid</structfield> <type>integer</type>
+ </para>
+ <para>
+ Process ID of the parallel vacuum leader, if this row is for one of its
+ parallel workers; <literal>NULL</literal> on the leader row itself and
+ for a non-parallel vacuum. This column is obtained by joining
+ <link linkend="monitoring-pg-stat-activity-view"><structname>pg_stat_activity</structname></link>,
+ so it follows the same visibility rules as the
+ <structfield>leader_pid</structfield> column there.
+ </para></entry>
+ </row>
+
</tbody>
</tgroup>
</table>
+ <note>
+ <para>
+ During a parallel vacuum, each participating backend (the leader and each
+ parallel worker) reports its own row. Only a subset of the columns is
+ meaningful on a worker row: <structfield>phase</structfield>,
+ <structfield>current_index_relid</structfield> and
+ <structfield>leader_pid</structfield>. The remaining columns track
+ command-level heap progress and are maintained only on the leader row;
+ they read as zero on worker rows. <structfield>leader_pid</structfield> is
+ <literal>NULL</literal> on the leader row (and for a non-parallel vacuum)
+ and the leader's PID on the worker rows, so it distinguishes the two.
+ Because the worker rows are separate backends, a role that does not have
+ privileges of the <literal>pg_read_all_stats</literal> role and does not
+ own those backends sees only their <structfield>pid</structfield>, with the
+ remaining columns <literal>NULL</literal>, just as in
+ <structname>pg_stat_activity</structname>.
+ </para>
+ </note>
+
<table id="vacuum-phases">
<title>VACUUM Phases</title>
<tgroup cols="2">
diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index 063ef2208de..8bc3de84581 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -3048,6 +3048,10 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3069,6 +3073,10 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /* Reset the current index relid to avoid reporting a stale value */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) InvalidOid);
+
return istat;
}
@@ -3099,6 +3107,10 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3118,6 +3130,10 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /* Reset the current index relid to avoid reporting a stale value */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) InvalidOid);
+
return istat;
}
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 8612d99a890..64a627b37b2 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -1353,9 +1353,12 @@ CREATE VIEW pg_stat_progress_vacuum AS
CASE S.param13 WHEN 1 THEN 'manual'
WHEN 2 THEN 'autovacuum'
WHEN 3 THEN 'autovacuum_wraparound'
- ELSE NULL END AS started_by
+ ELSE NULL END AS started_by,
+ CAST(S.param14 AS oid) AS current_index_relid,
+ A.leader_pid AS leader_pid
FROM pg_stat_get_progress_info('VACUUM') AS S
- LEFT JOIN pg_database D ON S.datid = D.oid;
+ LEFT JOIN pg_database D ON S.datid = D.oid
+ LEFT JOIN pg_stat_activity A ON A.pid = S.pid;
CREATE VIEW pg_stat_progress_repack AS
SELECT
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 767d162e578..d6a56f63a0f 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -1076,6 +1076,11 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
IndexBulkDeleteResult *istat = NULL;
IndexBulkDeleteResult *istat_res;
IndexVacuumInfo ivinfo;
+ const int progress_index[] = {
+ PROGRESS_VACUUM_PHASE,
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID
+ };
+ int64 progress_val[2];
/*
* Update the pointer to the corresponding bulk-deletion result if someone
@@ -1097,6 +1102,16 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
pvs->indname = pstrdup(RelationGetRelationName(indrel));
pvs->status = indstats->status;
+ /*
+ * Report the phase and the index we're about to process before we start,
+ * so that it is visible for the whole duration of the index scan.
+ */
+ progress_val[0] = (indstats->status == PARALLEL_INDVAC_STATUS_NEED_BULKDELETE)
+ ? PROGRESS_VACUUM_PHASE_VACUUM_INDEX
+ : PROGRESS_VACUUM_PHASE_INDEX_CLEANUP;
+ progress_val[1] = (int64) RelationGetRelid(indrel);
+ pgstat_progress_update_multi_param(2, progress_index, progress_val);
+
switch (indstats->status)
{
case PARALLEL_INDVAC_STATUS_NEED_BULKDELETE:
@@ -1112,6 +1127,10 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
RelationGetRelationName(indrel));
}
+ /* Reset the current index relid to avoid reporting a stale value */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) InvalidOid);
+
/*
* Copy the index bulk-deletion result returned from ambulkdelete and
* amvacuumcleanup to the DSM segment if it's the first cycle because they
@@ -1315,6 +1334,9 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Prepare to track buffer usage during parallel execution */
InstrStartParallelQuery();
+ /* Register this worker for vacuum progress reporting */
+ pgstat_progress_start_command(PROGRESS_COMMAND_VACUUM, shared->relid);
+
/* Process indexes to perform vacuum/cleanup */
parallel_vacuum_process_safe_indexes(&pvs);
@@ -1334,6 +1356,9 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Pop the error context stack */
error_context_stack = errcallback.previous;
+ /* Unregister this worker from vacuum progress reporting */
+ pgstat_progress_end_command();
+
vac_close_indexes(nindexes, indrels, RowExclusiveLock);
table_close(rel, ShareUpdateExclusiveLock);
FreeAccessStrategy(pvs.bstrategy);
diff --git a/src/include/commands/progress.h b/src/include/commands/progress.h
index 2a12920c75f..bee311955b6 100644
--- a/src/include/commands/progress.h
+++ b/src/include/commands/progress.h
@@ -31,6 +31,7 @@
#define PROGRESS_VACUUM_DELAY_TIME 10
#define PROGRESS_VACUUM_MODE 11
#define PROGRESS_VACUUM_STARTED_BY 12
+#define PROGRESS_VACUUM_CURRENT_INDEX_RELID 13
/* Phases of vacuum (as advertised via PROGRESS_VACUUM_PHASE) */
#define PROGRESS_VACUUM_PHASE_SCAN_HEAP 1
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 1a29d46213e..17caf9418f2 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -2213,9 +2213,12 @@ pg_stat_progress_vacuum| SELECT s.pid,
WHEN 2 THEN 'autovacuum'::text
WHEN 3 THEN 'autovacuum_wraparound'::text
ELSE NULL::text
- END AS started_by
- FROM (pg_stat_get_progress_info('VACUUM'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
- LEFT JOIN pg_database d ON ((s.datid = d.oid)));
+ END AS started_by,
+ (s.param14)::oid AS current_index_relid,
+ a.leader_pid
+ FROM ((pg_stat_get_progress_info('VACUUM'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
+ LEFT JOIN pg_database d ON ((s.datid = d.oid)))
+ LEFT JOIN pg_stat_activity a ON ((a.pid = s.pid)));
pg_stat_recovery| SELECT promote_triggered,
last_replayed_read_lsn,
last_replayed_end_lsn,
--
2.47.3
[application/octet-stream] v4-0002-Report-index-block-progress-in-pg_stat_progress_v.patch (9.5K, ../../CALj2ACX-49BrUTXaM3HeWVaZr+zGkbcLNJcrQ_dgjYF9AtW-kg@mail.gmail.com/3-v4-0002-Report-index-block-progress-in-pg_stat_progress_v.patch)
download | inline diff:
From 4cd92806c6e6a38a9913b9c37c85e498071c1dca Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Tue, 25 Aug 2026 03:20:56 +0000
Subject: [PATCH v4 2/2] Report index block progress in
pg_stat_progress_vacuum.
Previously, pg_stat_progress_vacuum reported which index a
backend was vacuuming but not how far it had gotten through that
index. This matters most for B-tree, the common index type: for a
large B-tree, at the scale of hundreds of GBs to TBs, per-index
progress is what lets one estimate when the index phase, and
together with heap_blks_* the whole vacuum, will finish.
This commit reuses the block counters that CREATE INDEX progress
reporting added in commit ab0dfc961 (PROGRESS_SCAN_BLOCKS_TOTAL
and PROGRESS_SCAN_BLOCKS_DONE), exposing them as new
index_blks_total and index_blks_done columns. B-tree's
bulk-delete scan already knows how to report them, so this commit
only turns that reporting on in the serial and parallel
index-vacuum paths. Like current_index_relid, the counters are
reported per participating backend, so each worker's row shows
progress for the index it is scanning.
These counters are kept separate from heap_blks_*. The heap block
counters must be retained across a multi-pass index vacuum, which
happens when the dead-TID store fills, and in the serial case the
leader does the index vacuuming itself, so reusing the heap
counters for index blocks would corrupt heap progress that is
still needed.
No caller sets IndexVacuumInfo.report_progress to false anymore;
removing it is parked for a follow-up commit.
Author: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Reviewed-by: Michael Paquier <michael@paquier.xyz>
Reviewed-by: Sami Imseih <samimseih@gmail.com>
Discussion: https://postgr.es/m/CALj2ACX6gyBmQfaqbCsycDmaSPbq1=iPJw1OqUT+aLqaKjW8dQ@mail.gmail.com
---
doc/src/sgml/monitoring.sgml | 27 +++++++++++++++++++++-
src/backend/access/heap/vacuumlazy.c | 32 ++++++++++++++++++++-------
src/backend/catalog/system_views.sql | 2 ++
src/backend/commands/vacuumparallel.c | 16 ++++++++++----
src/test/regress/expected/rules.out | 2 ++
5 files changed, 66 insertions(+), 13 deletions(-)
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index 0d3446d6e86..0f5dc69a246 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -7929,6 +7929,29 @@ FROM pg_stat_get_backend_idset() AS backendid;
</para></entry>
</row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>index_blks_total</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Total number of blocks to scan for the B-tree index identified by
+ <structfield>current_index_relid</structfield>. It is reset to 0 once
+ the index has been processed.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>index_blks_done</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of blocks scanned so far for the B-tree index identified by
+ <structfield>current_index_relid</structfield>. Together with
+ <structfield>index_blks_total</structfield> this gives per-index vacuum
+ progress, which is useful for estimating completion of large indexes.
+ </para></entry>
+ </row>
+
<row>
<entry role="catalog_table_entry"><para role="column_definition">
<structfield>leader_pid</structfield> <type>integer</type>
@@ -7952,7 +7975,9 @@ FROM pg_stat_get_backend_idset() AS backendid;
During a parallel vacuum, each participating backend (the leader and each
parallel worker) reports its own row. Only a subset of the columns is
meaningful on a worker row: <structfield>phase</structfield>,
- <structfield>current_index_relid</structfield> and
+ <structfield>current_index_relid</structfield>,
+ <structfield>index_blks_total</structfield>,
+ <structfield>index_blks_done</structfield> and
<structfield>leader_pid</structfield>. The remaining columns track
command-level heap progress and are maintained only on the leader row;
they read as zero on worker rows. <structfield>leader_pid</structfield> is
diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index 8bc3de84581..459fff8df1d 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -3038,11 +3038,17 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
{
IndexVacuumInfo ivinfo;
LVSavedErrInfo saved_err_info;
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.estimated_count = true;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = reltuples;
@@ -3073,9 +3079,11 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
- /* Reset the current index relid to avoid reporting a stale value */
- pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
- (int64) InvalidOid);
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
return istat;
}
@@ -3096,11 +3104,17 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
{
IndexVacuumInfo ivinfo;
LVSavedErrInfo saved_err_info;
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.estimated_count = estimated_count;
ivinfo.message_level = DEBUG2;
@@ -3130,9 +3144,11 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
- /* Reset the current index relid to avoid reporting a stale value */
- pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
- (int64) InvalidOid);
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
return istat;
}
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 64a627b37b2..741e8f4fa36 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -1355,6 +1355,8 @@ CREATE VIEW pg_stat_progress_vacuum AS
WHEN 3 THEN 'autovacuum_wraparound'
ELSE NULL END AS started_by,
CAST(S.param14 AS oid) AS current_index_relid,
+ S.param16 AS index_blks_total,
+ S.param17 AS index_blks_done,
A.leader_pid AS leader_pid
FROM pg_stat_get_progress_info('VACUUM') AS S
LEFT JOIN pg_database D ON S.datid = D.oid
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index d6a56f63a0f..be87b1937ad 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -1081,6 +1081,12 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
PROGRESS_VACUUM_CURRENT_INDEX_RELID
};
int64 progress_val[2];
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
/*
* Update the pointer to the corresponding bulk-deletion result if someone
@@ -1092,7 +1098,7 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
ivinfo.index = indrel;
ivinfo.heaprel = pvs->heaprel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.message_level = DEBUG2;
ivinfo.estimated_count = pvs->shared->estimated_count;
ivinfo.num_heap_tuples = pvs->shared->reltuples;
@@ -1127,9 +1133,11 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
RelationGetRelationName(indrel));
}
- /* Reset the current index relid to avoid reporting a stale value */
- pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
- (int64) InvalidOid);
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
/*
* Copy the index bulk-deletion result returned from ambulkdelete and
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 17caf9418f2..c5ad0881611 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -2215,6 +2215,8 @@ pg_stat_progress_vacuum| SELECT s.pid,
ELSE NULL::text
END AS started_by,
(s.param14)::oid AS current_index_relid,
+ s.param16 AS index_blks_total,
+ s.param17 AS index_blks_done,
a.leader_pid
FROM ((pg_stat_get_progress_info('VACUUM'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
LEFT JOIN pg_database d ON ((s.datid = d.oid)))
--
2.47.3
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-20 02:37 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-20 23:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-25 16:34 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
@ 2026-09-03 02:23 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 05:15 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
0 siblings, 1 reply; 32+ messages in thread
From: Bharath Rupireddy @ 2026-09-03 02:23 UTC (permalink / raw)
To: Michael Paquier <michael@paquier.xyz>; +Cc: Sami Imseih <samimseih@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
Hi,
On Tue, Aug 25, 2026 at 9:34 AM Bharath Rupireddy
<bharath.rupireddyforpostgres@gmail.com> wrote:
>
> Hi,
>
> On Thu, Aug 20, 2026 at 4:10 PM Michael Paquier <michael@paquier.xyz> wrote:
> >
> > On Wed, Aug 19, 2026 at 07:37:00PM -0700, Bharath Rupireddy wrote:
> > > I don't think a separate view is the right approach at least for two
> > > reasons. One, to know the progress of a vacuum one has to query two
> > > views and relate them. Two, since the leader itself participates in
> > > vacuuming indexes alongside workers (and will also for parallel heap
> > > vacuum), splitting the same command's progress into two views adds
> > > complexity. Keeping everything in a single view with one row per
> > > worker (as discussed upthread) is simpler. Some fields would be null
> > > on worker rows, but documenting this should be sufficient. That said,
> > > I'm open to hear more thoughts on this.
> >
> > Having a single progress view feels like the natural approach here,
> > for both the leader and the workers. The leader triggers the
> > existence of the workers, but both leader and workers may finish by
> > doing the same job as there could be usually little meaning for a
> > leader to stand idle, waiting for all the workers to do the work.
> > Such choices are implementation-agnostic, of course; we should not
> > lock ourselves.
>
> Thanks for reviewing. After thinking about this for a while, I still
> think one row per worker is the better choice, with the tradeoff of
> some fields being null on worker rows (similar to leader_pid in
> pg_stat_activity). The docs also say "the view will contain one row
> for each backend that is currently vacuuming," and in that sense,
> parallel workers launched for index vacuum are essentially doing
> vacuum too.
>
> Here's what I have. 0001 reports the current index being vacuumed, and
> 0002 reports the total blocks and done blocks for B-tree indexes. This
> helps track vacuum progress for large indexes in production.
>
> Please have a look.
Thanks Michael for the off-list chat. I agree that emitting leader_pid
via the vacuum progress report is information bloat, since one can
easily identify the workers for a given leader by looking at the
database OID and relation name, and if needed, can also join with
pg_stat_activity.leader_pid. So, I removed leader_pid in the 0001
patch.
Please find the attached v5 patches.
Here is some sample output that I captured:
-- Test 1: manual VACUUM, non-parallel index vacuum.
-- Take 1
pid | datid | datname | relid | phase |
heap_blks_scanned | current_index | index_blks_total | index_blks_done
| started_by
------+-------+----------+----------+-------------------+-------------------+---------------+------------------+-----------------+------------
5786 | 5 | postgres | test_vac | vacuuming indexes |
63695 | test_vac_idx1 | 27422 | 3706 | manual
-- Take 2
5786 | 5 | postgres | test_vac | vacuuming indexes |
63695 | test_vac_idx1 | 27422 | 10275 | manual
-- Test 2: manual VACUUM (PARALLEL 2), parallel index vacuum.
-- Take 1
pid | datid | datname | relid | phase |
heap_blks_scanned | current_index | index_blks_total | index_blks_done
| started_by
------+-------+----------+----------+-------------------+-------------------+---------------+------------------+-----------------+------------
8040 | 5 | postgres | test_vac | vacuuming indexes |
63695 | test_vac_idx1 | 27422 | 463 | manual
8372 | 5 | postgres | test_vac | vacuuming indexes |
0 | test_vac_idx3 | 27422 | 399 |
8373 | 5 | postgres | test_vac | vacuuming indexes |
0 | test_vac_idx2 | 27422 | 351 |
-- Take 2
8040 | 5 | postgres | test_vac | vacuuming indexes |
63695 | test_vac_idx1 | 27422 | 9855 | manual
8372 | 5 | postgres | test_vac | vacuuming indexes |
0 | test_vac_idx3 | 27422 | 9199 |
8373 | 5 | postgres | test_vac | vacuuming indexes |
0 | test_vac_idx2 | 27422 | 9145 |
-- Test 3: autovacuum, non-parallel index vacuum.
-- Take 1
pid | datid | datname | relid | phase |
heap_blks_scanned | current_index | index_blks_total | index_blks_done
| started_by
------+-------+----------+----------+-------------------+-------------------+---------------+------------------+-----------------+------------
9835 | 5 | postgres | test_vac | vacuuming indexes |
63695 | test_vac_idx1 | 27422 | 1023 |
autovacuum
-- Take 2
9835 | 5 | postgres | test_vac | vacuuming indexes |
63695 | test_vac_idx1 | 27422 | 3071 |
autovacuum
-- Test 4: autovacuum, parallel index vacuum (autovacuum_parallel_workers=2).
-- Take 1
pid | datid | datname | relid | phase |
heap_blks_scanned | current_index | index_blks_total | index_blks_done
| started_by
-------+-------+----------+----------+-------------------+-------------------+---------------+------------------+-----------------+------------
31309 | 5 | postgres | test_vac | vacuuming indexes |
63695 | test_vac_idx1 | 27422 | 459 |
autovacuum
650 | 5 | postgres | test_vac | vacuuming indexes |
0 | test_vac_idx2 | 27422 | 476 |
651 | 5 | postgres | test_vac | vacuuming indexes |
0 | test_vac_idx3 | 27422 | 474 |
-- Take 2
31309 | 5 | postgres | test_vac | vacuuming indexes |
63695 | test_vac_idx1 | 27422 | 2115 |
autovacuum
650 | 5 | postgres | test_vac | vacuuming indexes |
0 | test_vac_idx2 | 27422 | 2126 |
651 | 5 | postgres | test_vac | vacuuming indexes |
0 | test_vac_idx3 | 27422 | 2118 |
--
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com
Attachments:
[application/x-patch] v5-0001-Report-the-index-being-vacuumed-in-pg_stat_progre.patch (11.9K, ../../CALj2ACVYtmj=P4u1cBO+MwHFYQw_wzVxhtcV1WWrELbSOvuuzg@mail.gmail.com/2-v5-0001-Report-the-index-being-vacuumed-in-pg_stat_progre.patch)
download | inline diff:
From 40dd523966ec6032acd73382063875fc3b31a46f Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Wed, 2 Sep 2026 23:24:10 +0000
Subject: [PATCH v5 1/2] Report the index being vacuumed in
pg_stat_progress_vacuum.
Previously, pg_stat_progress_vacuum showed the total and processed
index counts, but not which index a backend was actually working on.
On a table with many indexes of different access methods, this made
it hard to tell which index a slow or stuck vacuum was spending its
time on.
This commit adds a current_index_relid column reporting the OID of
the index a backend is currently vacuuming or cleaning up. It is set
before each index is processed and reset once that index is done, so
it does not report a stale value after the phase moves on.
During parallel index vacuum the leader also participates, and each
participant processes a different set of indexes. That per-index
progress cannot be collapsed into a single leader row without losing
the detail that matters, so each participant, the leader and the
workers alike, now reports its own row and shows the index it is
processing. A worker row carries the columns that come from its own
backend state, pid, datid, datname, relid, phase and
current_index_relid; the remaining columns track command-level heap
progress that only the leader maintains and read as zero on worker
rows. Because a table can be vacuumed by only one VACUUM at a time,
the rows sharing a datid and relid make up a single vacuum, so they
can be grouped that way to see the leader together with all of its
workers. If the exact leader-to-worker mapping is wanted, it can be
recovered from leader_pid in pg_stat_activity.
Author: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Reviewed-by: Michael Paquier <michael@paquier.xyz>
Reviewed-by: Sami Imseih <samimseih@gmail.com>
Reviewed-by: Satyanarayana Narlapuram <satyanarlapuram@gmail.com>
Discussion: https://postgr.es/m/CALj2ACUgwSchK6jQ2CdKLBWUADTOE_zKdTff2Zg3E6hOuXKv-w@mail.gmail.com
---
doc/src/sgml/monitoring.sgml | 47 +++++++++++++++++++++++++--
src/backend/access/heap/vacuumlazy.c | 16 +++++++++
src/backend/catalog/system_views.sql | 3 +-
src/backend/commands/vacuumparallel.c | 25 ++++++++++++++
src/include/commands/progress.h | 1 +
src/test/regress/expected/rules.out | 3 +-
6 files changed, 91 insertions(+), 4 deletions(-)
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index b403fb990a7..e39f9637775 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -7651,8 +7651,11 @@ FROM pg_stat_get_backend_idset() AS backendid;
<para>
Whenever <command>VACUUM</command> is running, the
<structname>pg_stat_progress_vacuum</structname> view will contain
- one row for each backend (including autovacuum worker processes) that is
- currently vacuuming. The tables below describe the information
+ one row for each backend (including autovacuum worker processes and
+ parallel workers launched for
+ <link linkend="sql-vacuum">parallel vacuum</link>) that is currently
+ vacuuming; see the note following the view for the columns reported on
+ parallel worker rows. The tables below describe the information
that will be reported and provide information about how to interpret it.
Progress for <command>VACUUM FULL</command> commands is reported via
<structname>pg_stat_progress_repack</structname>, and is also visible via
@@ -7910,10 +7913,50 @@ FROM pg_stat_get_backend_idset() AS backendid;
</itemizedlist>
</para></entry>
</row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>current_index_relid</structfield> <type>oid</type>
+ </para>
+ <para>
+ If <command>VACUUM</command> is currently processing an index, this
+ column shows the OID of the index being vacuumed. The value is set
+ when the phase is <literal>vacuuming indexes</literal> or
+ <literal>cleaning up indexes</literal> and is reset to 0 when all
+ indexes have been processed. During parallel index vacuum, each
+ parallel worker row shows the index that particular worker is
+ processing.
+ </para></entry>
+ </row>
+
</tbody>
</tgroup>
</table>
+ <note>
+ <para>
+ During a parallel vacuum, each participating backend (the leader and each
+ parallel worker) reports its own row. On a worker row the meaningful
+ columns are <structfield>pid</structfield>, <structfield>datid</structfield>,
+ <structfield>datname</structfield>, <structfield>relid</structfield>,
+ <structfield>phase</structfield> and
+ <structfield>current_index_relid</structfield>. The remaining columns track
+ command-level heap progress that only the leader maintains; they read as
+ zero on worker rows. Because a table can be vacuumed by only one
+ <command>VACUUM</command> at a time, the rows sharing a given
+ <structfield>datid</structfield> and <structfield>relid</structfield>
+ together make up a single vacuum, so they can be grouped that way to see
+ the leader and all of its workers. The leader and worker process IDs can be
+ correlated with <structfield>leader_pid</structfield> in
+ <link linkend="monitoring-pg-stat-activity-view"><structname>pg_stat_activity</structname></link>
+ if needed. Because the worker rows are separate backends, a role that does
+ not have privileges of the <literal>pg_read_all_stats</literal> role and
+ does not own those backends sees only their <structfield>pid</structfield>,
+ with the remaining columns <literal>NULL</literal>, just as in
+ <structname>pg_stat_activity</structname>.
+ </para>
+ </note>
+
<table id="vacuum-phases">
<title>VACUUM Phases</title>
<tgroup cols="2">
diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index 063ef2208de..8bc3de84581 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -3048,6 +3048,10 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3069,6 +3073,10 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /* Reset the current index relid to avoid reporting a stale value */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) InvalidOid);
+
return istat;
}
@@ -3099,6 +3107,10 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3118,6 +3130,10 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /* Reset the current index relid to avoid reporting a stale value */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) InvalidOid);
+
return istat;
}
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 8612d99a890..24d2ae8bb44 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -1353,7 +1353,8 @@ CREATE VIEW pg_stat_progress_vacuum AS
CASE S.param13 WHEN 1 THEN 'manual'
WHEN 2 THEN 'autovacuum'
WHEN 3 THEN 'autovacuum_wraparound'
- ELSE NULL END AS started_by
+ ELSE NULL END AS started_by,
+ CAST(S.param14 AS oid) AS current_index_relid
FROM pg_stat_get_progress_info('VACUUM') AS S
LEFT JOIN pg_database D ON S.datid = D.oid;
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 767d162e578..d6a56f63a0f 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -1076,6 +1076,11 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
IndexBulkDeleteResult *istat = NULL;
IndexBulkDeleteResult *istat_res;
IndexVacuumInfo ivinfo;
+ const int progress_index[] = {
+ PROGRESS_VACUUM_PHASE,
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID
+ };
+ int64 progress_val[2];
/*
* Update the pointer to the corresponding bulk-deletion result if someone
@@ -1097,6 +1102,16 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
pvs->indname = pstrdup(RelationGetRelationName(indrel));
pvs->status = indstats->status;
+ /*
+ * Report the phase and the index we're about to process before we start,
+ * so that it is visible for the whole duration of the index scan.
+ */
+ progress_val[0] = (indstats->status == PARALLEL_INDVAC_STATUS_NEED_BULKDELETE)
+ ? PROGRESS_VACUUM_PHASE_VACUUM_INDEX
+ : PROGRESS_VACUUM_PHASE_INDEX_CLEANUP;
+ progress_val[1] = (int64) RelationGetRelid(indrel);
+ pgstat_progress_update_multi_param(2, progress_index, progress_val);
+
switch (indstats->status)
{
case PARALLEL_INDVAC_STATUS_NEED_BULKDELETE:
@@ -1112,6 +1127,10 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
RelationGetRelationName(indrel));
}
+ /* Reset the current index relid to avoid reporting a stale value */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) InvalidOid);
+
/*
* Copy the index bulk-deletion result returned from ambulkdelete and
* amvacuumcleanup to the DSM segment if it's the first cycle because they
@@ -1315,6 +1334,9 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Prepare to track buffer usage during parallel execution */
InstrStartParallelQuery();
+ /* Register this worker for vacuum progress reporting */
+ pgstat_progress_start_command(PROGRESS_COMMAND_VACUUM, shared->relid);
+
/* Process indexes to perform vacuum/cleanup */
parallel_vacuum_process_safe_indexes(&pvs);
@@ -1334,6 +1356,9 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Pop the error context stack */
error_context_stack = errcallback.previous;
+ /* Unregister this worker from vacuum progress reporting */
+ pgstat_progress_end_command();
+
vac_close_indexes(nindexes, indrels, RowExclusiveLock);
table_close(rel, ShareUpdateExclusiveLock);
FreeAccessStrategy(pvs.bstrategy);
diff --git a/src/include/commands/progress.h b/src/include/commands/progress.h
index 2a12920c75f..bee311955b6 100644
--- a/src/include/commands/progress.h
+++ b/src/include/commands/progress.h
@@ -31,6 +31,7 @@
#define PROGRESS_VACUUM_DELAY_TIME 10
#define PROGRESS_VACUUM_MODE 11
#define PROGRESS_VACUUM_STARTED_BY 12
+#define PROGRESS_VACUUM_CURRENT_INDEX_RELID 13
/* Phases of vacuum (as advertised via PROGRESS_VACUUM_PHASE) */
#define PROGRESS_VACUUM_PHASE_SCAN_HEAP 1
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 1a29d46213e..cf9b8c823c8 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -2213,7 +2213,8 @@ pg_stat_progress_vacuum| SELECT s.pid,
WHEN 2 THEN 'autovacuum'::text
WHEN 3 THEN 'autovacuum_wraparound'::text
ELSE NULL::text
- END AS started_by
+ END AS started_by,
+ (s.param14)::oid AS current_index_relid
FROM (pg_stat_get_progress_info('VACUUM'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
LEFT JOIN pg_database d ON ((s.datid = d.oid)));
pg_stat_recovery| SELECT promote_triggered,
--
2.47.3
[application/x-patch] v5-0002-Report-index-block-progress-in-pg_stat_progress_v.patch (9.8K, ../../CALj2ACVYtmj=P4u1cBO+MwHFYQw_wzVxhtcV1WWrELbSOvuuzg@mail.gmail.com/3-v5-0002-Report-index-block-progress-in-pg_stat_progress_v.patch)
download | inline diff:
From 249378140e1016436401da4f117cb5efeab9351e Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Wed, 2 Sep 2026 23:27:23 +0000
Subject: [PATCH v5 2/2] Report index block progress in
pg_stat_progress_vacuum.
Previously, pg_stat_progress_vacuum reported which index a backend
was vacuuming but not how far it had gotten through that index. This
matters most for B-tree, the common index type: for a large B-tree,
at the scale of hundreds of GBs to TBs, per-index progress is what
lets one estimate when the index phase, and together with
heap_blks_* the whole vacuum, will finish.
This commit reuses the block counters that CREATE INDEX progress
reporting added in commit ab0dfc961 (PROGRESS_SCAN_BLOCKS_TOTAL and
PROGRESS_SCAN_BLOCKS_DONE), exposing them as new index_blks_total
and index_blks_done columns. B-tree's bulk-delete scan already knows
how to report them, so this commit only turns that reporting on in
the serial and parallel index-vacuum paths. Like current_index_relid,
the counters are reported per participating backend, so each worker's
row shows progress for the index it is scanning.
These counters are kept separate from heap_blks_*. The heap block
counters must be retained across a multi-pass index vacuum, which
happens when the dead-TID store fills, and in the serial case the
leader does the index vacuuming itself, so reusing the heap counters
for index blocks would corrupt heap progress that is still needed.
No caller sets IndexVacuumInfo.report_progress to false anymore;
removing it is parked for a follow-up commit.
Author: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Reviewed-by: Michael Paquier <michael@paquier.xyz>
Reviewed-by: Sami Imseih <samimseih@gmail.com>
Discussion: https://postgr.es/m/CALj2ACX6gyBmQfaqbCsycDmaSPbq1=iPJw1OqUT+aLqaKjW8dQ@mail.gmail.com
---
doc/src/sgml/monitoring.sgml | 28 +++++++++++++++++++++--
src/backend/access/heap/vacuumlazy.c | 32 ++++++++++++++++++++-------
src/backend/catalog/system_views.sql | 4 +++-
src/backend/commands/vacuumparallel.c | 16 ++++++++++----
src/test/regress/expected/rules.out | 4 +++-
5 files changed, 68 insertions(+), 16 deletions(-)
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index e39f9637775..d7f6a83edfd 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -7929,6 +7929,29 @@ FROM pg_stat_get_backend_idset() AS backendid;
</para></entry>
</row>
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>index_blks_total</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Total number of blocks to scan for the B-tree index identified by
+ <structfield>current_index_relid</structfield>. It is reset to 0 once
+ the index has been processed.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>index_blks_done</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of blocks scanned so far for the B-tree index identified by
+ <structfield>current_index_relid</structfield>. Together with
+ <structfield>index_blks_total</structfield> this gives per-index vacuum
+ progress, which is useful for estimating completion of large indexes.
+ </para></entry>
+ </row>
+
</tbody>
</tgroup>
</table>
@@ -7939,8 +7962,9 @@ FROM pg_stat_get_backend_idset() AS backendid;
parallel worker) reports its own row. On a worker row the meaningful
columns are <structfield>pid</structfield>, <structfield>datid</structfield>,
<structfield>datname</structfield>, <structfield>relid</structfield>,
- <structfield>phase</structfield> and
- <structfield>current_index_relid</structfield>. The remaining columns track
+ <structfield>phase</structfield>, <structfield>current_index_relid</structfield>,
+ <structfield>index_blks_total</structfield> and
+ <structfield>index_blks_done</structfield>. The remaining columns track
command-level heap progress that only the leader maintains; they read as
zero on worker rows. Because a table can be vacuumed by only one
<command>VACUUM</command> at a time, the rows sharing a given
diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index 8bc3de84581..459fff8df1d 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -3038,11 +3038,17 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
{
IndexVacuumInfo ivinfo;
LVSavedErrInfo saved_err_info;
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.estimated_count = true;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = reltuples;
@@ -3073,9 +3079,11 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
- /* Reset the current index relid to avoid reporting a stale value */
- pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
- (int64) InvalidOid);
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
return istat;
}
@@ -3096,11 +3104,17 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
{
IndexVacuumInfo ivinfo;
LVSavedErrInfo saved_err_info;
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.estimated_count = estimated_count;
ivinfo.message_level = DEBUG2;
@@ -3130,9 +3144,11 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
- /* Reset the current index relid to avoid reporting a stale value */
- pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
- (int64) InvalidOid);
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
return istat;
}
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index 24d2ae8bb44..f29e2767154 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -1354,7 +1354,9 @@ CREATE VIEW pg_stat_progress_vacuum AS
WHEN 2 THEN 'autovacuum'
WHEN 3 THEN 'autovacuum_wraparound'
ELSE NULL END AS started_by,
- CAST(S.param14 AS oid) AS current_index_relid
+ CAST(S.param14 AS oid) AS current_index_relid,
+ S.param16 AS index_blks_total,
+ S.param17 AS index_blks_done
FROM pg_stat_get_progress_info('VACUUM') AS S
LEFT JOIN pg_database D ON S.datid = D.oid;
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index d6a56f63a0f..be87b1937ad 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -1081,6 +1081,12 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
PROGRESS_VACUUM_CURRENT_INDEX_RELID
};
int64 progress_val[2];
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
/*
* Update the pointer to the corresponding bulk-deletion result if someone
@@ -1092,7 +1098,7 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
ivinfo.index = indrel;
ivinfo.heaprel = pvs->heaprel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.message_level = DEBUG2;
ivinfo.estimated_count = pvs->shared->estimated_count;
ivinfo.num_heap_tuples = pvs->shared->reltuples;
@@ -1127,9 +1133,11 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
RelationGetRelationName(indrel));
}
- /* Reset the current index relid to avoid reporting a stale value */
- pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
- (int64) InvalidOid);
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
/*
* Copy the index bulk-deletion result returned from ambulkdelete and
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index cf9b8c823c8..475604289de 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -2214,7 +2214,9 @@ pg_stat_progress_vacuum| SELECT s.pid,
WHEN 3 THEN 'autovacuum_wraparound'::text
ELSE NULL::text
END AS started_by,
- (s.param14)::oid AS current_index_relid
+ (s.param14)::oid AS current_index_relid,
+ s.param16 AS index_blks_total,
+ s.param17 AS index_blks_done
FROM (pg_stat_get_progress_info('VACUUM'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
LEFT JOIN pg_database d ON ((s.datid = d.oid)));
pg_stat_recovery| SELECT promote_triggered,
--
2.47.3
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-20 02:37 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-20 23:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-25 16:34 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 02:23 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
@ 2026-09-03 05:15 ` Michael Paquier <michael@paquier.xyz>
2026-09-03 16:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih.pg@gmail.com>
0 siblings, 1 reply; 32+ messages in thread
From: Michael Paquier @ 2026-09-03 05:15 UTC (permalink / raw)
To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: Sami Imseih <samimseih@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
On Wed, Sep 02, 2026 at 07:23:00PM -0700, Bharath Rupireddy wrote:
> Thanks Michael for the off-list chat. I agree that emitting leader_pid
> via the vacuum progress report is information bloat, since one can
> easily identify the workers for a given leader by looking at the
> database OID and relation name, and if needed, can also join with
> pg_stat_activity.leader_pid. So, I removed leader_pid in the 0001
> patch.
Full disclosure. I have discussed this patch set a bit with Bharath.
The discussion can be summed up like:
- leader_pid in the progress view with a JOIN to pg_stat_activity
feels like bloating the view with duplicated information.
- The database OID, the relation OID and the index OID gain in
visibility by being specified in the lines for the workers. Note: it
looks like we are doing so based on your output posted upthread,
missed that during our discussion, initially.
- The two new fields for total index blocks and index blocks processed
make more sense than trying to reuse the heap attributes because a
leader may do itself some of the cleanup. Multiple passes are less
likely lately, but could still be possible, and we want to know where
the leader is at for the heap part while working on the indexes.
Reading through v5-0001 and v5-0002, it looks like all these check
boxes are ticked. In terms of review clarity, splitting both patches
slightly helps, but I'd rather merge both things together at the end:
the first patch gains a lot in value thanks to the second patch where
the two block aggregates are added.
Perhaps I am missing something else? In this case, please feel free
to overwrite my words..
--
Michael
Attachments:
[application/pgp-signature] signature.asc (832B, ../../apkCcttASAC6d6J8@paquier.xyz/2-signature.asc)
download
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-20 02:37 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-20 23:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-25 16:34 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 02:23 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 05:15 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
@ 2026-09-03 16:49 ` Sami Imseih <samimseih.pg@gmail.com>
2026-09-09 19:01 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Masahiko Sawada <sawada.mshk@gmail.com>
0 siblings, 1 reply; 32+ messages in thread
From: Sami Imseih @ 2026-09-03 16:49 UTC (permalink / raw)
To: Michael Paquier <michael@paquier.xyz>; +Cc: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Sami Imseih <samimseih@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
Sorry, but I'm still not convinced this is the right path here.
I just want to put my concerns about this approach in a single reply,
including some that I do not think were really addressed earlier.
My concern is that this changes the semantics of a row in
pg_stat_progress_vacuum in a way that feels pretty confusing. The same
row now mixes command-level aggregate fields with backend-local fields,
and many of the command-level fields are effectively no-op on worker
rows.
In the output I am looking at, heap_blks_total, heap_blks_scanned,
heap_blks_vacuumed, index_vacuum_count, max_dead_tuple_bytes,
dead_tuple_bytes, num_dead_item_ids, indexes_total, indexes_processed,
and delay_time are all really leader-level command state. On the other
hand, current_index_relid and index_blks_* are backend-local. mode and
started_by are unset on worker rows. So to me a row no longer
represents one coherent kind of progress information.
We do not have precedent for this elsewhere in the progress views, and
I think it will be confusing for users and monitoring tools.
Documenting the caveats does not really solve that problem.
I also worry about the precedent this sets for other progress views that
may want worker reporting later, for example
pg_stat_progress_create_index. If we applied the same pattern there,
we would immediately get more no-op fields on worker rows. So I do not
think this is just a VACUUM specific awkwardness.
Putting my monitoring-tool hat on, I do not see how this is a cleaner
interface. If the worker detail lives in the same view, a tool
now has to reconstruct one logical VACUUM by grouping or self-joining
pg_stat_progress_vacuum on (datid, relid), and then also know which
fields are command-level aggregates and which are backend-local. If we
really want to separate worker-level detail from leader-level progress,
I think a secondary view with an explicit join key such as leader_pid is
a cleaner approach.
There is also a broader problem with this approach. If we want to add
more worker-specific fields, the view gets wider, which will also add to
the incoherence of it.
I do not think we need to solve this part immediately, because it is not
really a problem yet. But I still think it is worth calling out that
the underlying progress tracking only gives us 20 generic slots in
st_progress_param[]. This patch works today because it uses slots that
were previously unused for VACUUM, but we will run out of those if we
want to expose more worker detail. At that point, I think we need a clear
mechanism for re-using progress fields differently on leader and worker rows,
and I do not think we have that.
--
Sami Imseih
Amazon Web Services (AWS)
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-20 02:37 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-20 23:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-25 16:34 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 02:23 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 05:15 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-09-03 16:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih.pg@gmail.com>
@ 2026-09-09 19:01 ` Masahiko Sawada <sawada.mshk@gmail.com>
2026-09-14 20:28 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
0 siblings, 1 reply; 32+ messages in thread
From: Masahiko Sawada @ 2026-09-09 19:01 UTC (permalink / raw)
To: Sami Imseih <samimseih.pg@gmail.com>; +Cc: Michael Paquier <michael@paquier.xyz>; Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Sami Imseih <samimseih@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
Hi,
On Thu, Sep 3, 2026 at 9:49 AM Sami Imseih <samimseih.pg@gmail.com> wrote:
>
> Sorry, but I'm still not convinced this is the right path here.
>
> I just want to put my concerns about this approach in a single reply,
> including some that I do not think were really addressed earlier.
>
> My concern is that this changes the semantics of a row in
> pg_stat_progress_vacuum in a way that feels pretty confusing. The same
> row now mixes command-level aggregate fields with backend-local fields,
> and many of the command-level fields are effectively no-op on worker
> rows.
>
> In the output I am looking at, heap_blks_total, heap_blks_scanned,
> heap_blks_vacuumed, index_vacuum_count, max_dead_tuple_bytes,
> dead_tuple_bytes, num_dead_item_ids, indexes_total, indexes_processed,
> and delay_time are all really leader-level command state. On the other
> hand, current_index_relid and index_blks_* are backend-local. mode and
> started_by are unset on worker rows. So to me a row no longer
> represents one coherent kind of progress information.
>
> We do not have precedent for this elsewhere in the progress views, and
> I think it will be confusing for users and monitoring tools.
> Documenting the caveats does not really solve that problem.
>
> I also worry about the precedent this sets for other progress views that
> may want worker reporting later, for example
> pg_stat_progress_create_index. If we applied the same pattern there,
> we would immediately get more no-op fields on worker rows. So I do not
> think this is just a VACUUM specific awkwardness.
>
> Putting my monitoring-tool hat on, I do not see how this is a cleaner
> interface. If the worker detail lives in the same view, a tool
> now has to reconstruct one logical VACUUM by grouping or self-joining
> pg_stat_progress_vacuum on (datid, relid), and then also know which
> fields are command-level aggregates and which are backend-local. If we
> really want to separate worker-level detail from leader-level progress,
> I think a secondary view with an explicit join key such as leader_pid is
> a cleaner approach.
>
> There is also a broader problem with this approach. If we want to add
> more worker-specific fields, the view gets wider, which will also add to
> the incoherence of it.
I'm studying the patch and discussion so I might be missing something,
but let me share my thoughts on this patch:
I think it's fine for a progress view to have an entry per backend,
including parallel workers, so more than one row for a single command.
If we ever implement parallel index rebuilding for REINDEX TABLE,
where different workers rebuild different indexes, having one entry
per worker in pg_stat_progress_create_index seems the straightforward
thing to do. A dedicated view for workers, say
pg_stat_progress_create_index_worker, would end up with mostly the
same columns as pg_stat_progress_create_index.
I can also see Sami's point that the patch makes many of the existing
columns of pg_stat_progress_vacuum no-op. But I'm not sure a dedicated
worker view really avoids that. If a user wants the progress of one
vacuum command they would join the two views, get one row per
participant anyway, and the leader's heap columns would be repeated on
every row. That's better than reporting them as 0, but the user still
has to know which columns are command-level and which are
backend-local. It might be worth clarifying the actual query and its
output for each approach and comparing them.
If we do want a separate view to avoid meaningless columns, I think it
should be split by the job, not by who does it. For example a
pg_stat_progress_index_vacuum view showing index bulkdelete and index
cleanup per index vacuum work, where both the leader and the parallel
workers have their own entry. One benefit is that GIN's pending list
cleanup invoked from autoanalyze could report there instead of adding
more columns to pg_stat_progress_analyze. The same probably applies to
REPACK, which might want to use pg_stat_progress_create_index for its
index rebuild. This needs infrastructure changes so that a backend can
report more than one progress at a time, though, so it isn't something
for this patch.
Regards,
--
Masahiko Sawada
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-20 02:37 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-20 23:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-25 16:34 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 02:23 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 05:15 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-09-03 16:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih.pg@gmail.com>
2026-09-09 19:01 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Masahiko Sawada <sawada.mshk@gmail.com>
@ 2026-09-14 20:28 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-15 17:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih.pg@gmail.com>
0 siblings, 1 reply; 32+ messages in thread
From: Bharath Rupireddy @ 2026-09-14 20:28 UTC (permalink / raw)
To: Masahiko Sawada <sawada.mshk@gmail.com>; +Cc: Sami Imseih <samimseih.pg@gmail.com>; Michael Paquier <michael@paquier.xyz>; Sami Imseih <samimseih@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
Hi,
On Wed, Sep 9, 2026 at 12:01 PM Masahiko Sawada <sawada.mshk@gmail.com> wrote:
>
> I'm studying the patch and discussion so I might be missing something,
> but let me share my thoughts on this patch:
>
> I think it's fine for a progress view to have an entry per backend,
> including parallel workers, so more than one row for a single command.
> If we ever implement parallel index rebuilding for REINDEX TABLE,
> where different workers rebuild different indexes, having one entry
> per worker in pg_stat_progress_create_index seems the straightforward
> thing to do. A dedicated view for workers, say
> pg_stat_progress_create_index_worker, would end up with mostly the
> same columns as pg_stat_progress_create_index.
Thanks for sharing the thoughts.
> I can also see Sami's point that the patch makes many of the existing
> columns of pg_stat_progress_vacuum no-op. But I'm not sure a dedicated
> worker view really avoids that. If a user wants the progress of one
> vacuum command they would join the two views, get one row per
> participant anyway, and the leader's heap columns would be repeated on
> every row. That's better than reporting them as 0, but the user still
> has to know which columns are command-level and which are
> backend-local. It might be worth clarifying the actual query and its
> output for each approach and comparing them.
Here's the sample output [1] with the two approaches, and yes, the
separate view for workers has repeated columns when joined to get the
progress report of a single vacuum command.
That said, I would like to mention an interesting and closely related
discussion on single view vs. separate view to show wait event info of
different processes in pg_stat_activity. The alignment there (almost
unanimously) was to go with the single view even if some of the
columns are not meaningful for certain processes (for example, most of
the pg_stat_activity columns are not applicable to auxiliary processes
such as the checkpointer, background writer, autovacuum launcher,
etc.): https://www.postgresql.org/message-id/CA%2BTgmoYES5nhkEGw9nZXU8\_FhA8XEm8NTm3-SO%2B3ML1B81Hkww%40mail.gmail.com.
As suggested upthread by Michael, I have merged the two patches (index
being vacuumed and index progress report) into a single patch since
they are closely related and there is no strong reason to keep them
separate. I also attached a 0002 patch that removes the now-unused
IndexVacuumInfo.report_progress field. Please find the attached v6
patch.
Note that I haven't added tests for pg_stat_progress_vacuum. It looks
like we have lived without them so far. I'm happy to add a simple
progress report test with parallel index vacuuming (I haven't checked,
but we may need an injection point to hold the parallel workers during
index vacuuming).
[1]
-- One row per worker
select * from pg_stat_progress_vacuum;
pid | datid | datname | relid | phase |
heap_blks_total | heap_blks_scanned | heap_blks_vacuumed |
index_vacuum_count | max_dead_tuple_bytes | dead_tuple_bytes |
num_dead_item_ids | indexes_total | indexes_processed | delay_time |
mode | started_by | current_index_relid | index_blks_total |
index_blks_done
-------+-------+----------+-------+-------------------+-----------------+-------------------+--------------------+--------------------+----------------------+------------------+-------------------+---------------+-------------------+------------+--------+------------+---------------------+------------------+-----------------
24598 | 5 | postgres | 16408 | vacuuming indexes |
16667 | 12862 | 0 | 0 |
1048576 | 1310720 | 352797 |
4 | 0 | 0 | normal | manual |
16413 | 5487 | 303
24600 | 5 | postgres | 16408 | vacuuming indexes |
0 | 0 | 0 | 0 |
0 | 0 | 0 | 0
| 0 | 0 | (null) | (null) |
16416 | 1754 | 130
24601 | 5 | postgres | 16408 | vacuuming indexes |
0 | 0 | 0 | 0 |
0 | 0 | 0 | 0
| 0 | 0 | (null) | (null) |
16414 | 5487 | 335
24602 | 5 | postgres | 16408 | vacuuming indexes |
0 | 0 | 0 | 0 |
0 | 0 | 0 | 0 |
0 | 0 | (null) | (null) |
16415 | 5487 | 291
(4 rows)
-- Separate views for leader and worker vacuum progress
select * from pg_stat_progress_vacuum v
left join pg_stat_progress_vacuum_worker w
on w.leader_pid = v.pid;
pid | datid | datname | relid | phase |
heap_blks_total | heap_blks_scanned | heap_blks_vacuumed |
index_vacuum_count | max_dead_tuple_bytes | dead_tuple_bytes |
num_dead_item_ids | indexes_total | indexes_processed | delay_time |
mode | started_by | current_index_relid | index_blks_total |
index_blks_done | pid | leader_pid | datid | datname | relid |
phase | current_index_relid | index_blks_total |
index_blks_done
-------+-------+----------+-------+-------------------+-----------------+-------------------+--------------------+--------------------+----------------------+------------------+-------------------+---------------+-------------------+------------+--------+------------+---------------------+------------------+-----------------+-------+------------+-------+----------+-------+-------------------+---------------------+------------------+-----------------
27109 | 5 | postgres | 16384 | vacuuming indexes |
16667 | 12862 | 0 | 0 |
1048576 | 1310720 | 205920 |
4 | 0 | 0 | normal | manual |
16389 | 5487 | 1577 | 27110 | 27109 |
5 | postgres | 16384 | vacuuming indexes | 16392 |
1754 | 1070
27109 | 5 | postgres | 16384 | vacuuming indexes |
16667 | 12862 | 0 | 0 |
1048576 | 1310720 | 205920 |
4 | 0 | 0 | normal | manual |
16389 | 5487 | 1577 | 27111 | 27109 |
5 | postgres | 16384 | vacuuming indexes | 16390 |
5487 | 1583
27109 | 5 | postgres | 16384 | vacuuming indexes |
16667 | 12862 | 0 | 0 |
1048576 | 1310720 | 205920 |
4 | 0 | 0 | normal | manual |
16389 | 5487 | 1577 | 27112 | 27109 |
5 | postgres | 16384 | vacuuming indexes | 16391 |
5487 | 1599
(3 rows)
--
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com
Attachments:
[application/x-patch] v6-0001-Report-per-index-vacuum-progress-in-pg_stat_progr.patch (17.2K, ../../CALj2ACV6vuAYVnLqHR62ggMtDfD0YTomW_YxoXL0dMMneq0bxQ@mail.gmail.com/2-v6-0001-Report-per-index-vacuum-progress-in-pg_stat_progr.patch)
download | inline diff:
From 39725d292975a40ba1245d4b9e254d16ecc31f59 Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Sun, 13 Sep 2026 05:19:26 +0000
Subject: [PATCH v6 1/2] Report per-index vacuum progress in
pg_stat_progress_vacuum.
Previously, pg_stat_progress_vacuum showed the total and
processed index counts, but neither which index a backend was
working on nor how far it had gotten through it. On a table with
many indexes of different access methods, that made it hard to
tell which index a slow or stuck vacuum was spending its time on,
and for a large B-tree, at the scale of hundreds of GBs to TBs,
there was no way to estimate when the index phase, and together
with heap_blks_*, the whole vacuum would finish.
This commit adds three columns. current_index_relid reports the
OID of the index a backend is currently vacuuming or cleaning up.
index_blks_total and index_blks_done report block progress within
that index. All three are set before an index is processed and
reset once that index is done, so they do not report a stale
value after the phase moves on.
During parallel index vacuum the leader also participates, and
each participant processes a different set of indexes. That
per-index progress cannot be collapsed into a single leader row
without losing the detail that matters, so each participant, the
leader and the workers alike, now reports its own row and shows
the index it is processing and how far along it is. A worker row
carries the columns that come from its own backend state, pid,
datid, datname, relid, phase, current_index_relid,
index_blks_total and index_blks_done; the remaining columns track
command-level heap progress that only the leader maintains and
read as zero on worker rows. Because a table can be vacuumed by
only one VACUUM at a time, the rows sharing a datid and relid
make up a single vacuum, so they can be grouped that way to see
the leader together with all of its workers. If the exact
leader-to-worker mapping is wanted, it can be recovered from
leader_pid in pg_stat_activity.
The block counters are the ones CREATE INDEX progress reporting
added in commit ab0dfc961b6a, PROGRESS_SCAN_BLOCKS_TOTAL and
PROGRESS_SCAN_BLOCKS_DONE. B-tree's index scan already knows how
to report them, so all that is needed here is turning that
reporting on in the serial and parallel index vacuum paths; other
access methods report nothing and leave the two columns at zero.
They are deliberately kept separate from heap_blks_*, which must
be retained across a multi-pass index vacuum, the case where the
dead-TID store fills, and which in the serial case belong to the
same backend that is doing the index vacuuming, so reusing them
for index blocks would destroy heap progress that is still needed.
No caller sets IndexVacuumInfo.report_progress to false anymore;
the next commit removes the field.
XXX: Bump catalog version.
Author: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Reviewed-by: Michael Paquier <michael@paquier.xyz>
Reviewed-by: Sami Imseih <samimseih@gmail.com>
Reviewed-by: Masahiko Sawada <sawada.mshk@gmail.com>
Discussion: https://postgr.es/m/CALj2ACUgwSchK6jQ2CdKLBWUADTOE_zKdTff2Zg3E6hOuXKv-w@mail.gmail.com
---
doc/src/sgml/monitoring.sgml | 93 ++++++++++++++++++++++++++-
src/backend/access/heap/vacuumlazy.c | 36 ++++++++++-
src/backend/catalog/system_views.sql | 5 +-
src/backend/commands/vacuumparallel.c | 35 +++++++++-
src/include/commands/progress.h | 2 +
src/test/regress/expected/rules.out | 5 +-
6 files changed, 169 insertions(+), 7 deletions(-)
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index 49bf6b51c49..a699d95514e 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -7670,8 +7670,11 @@ FROM pg_stat_get_backend_idset() AS backendid;
<para>
Whenever <command>VACUUM</command> is running, the
<structname>pg_stat_progress_vacuum</structname> view will contain
- one row for each backend (including autovacuum worker processes) that is
- currently vacuuming. The tables below describe the information
+ one row for each backend (including autovacuum worker processes and
+ parallel workers launched for
+ <link linkend="sql-vacuum">parallel vacuum</link>) that is currently
+ vacuuming; see the note following the view for the columns reported on
+ parallel worker rows. The tables below describe the information
that will be reported and provide information about how to interpret it.
Progress for <command>VACUUM FULL</command> commands is reported via
<structname>pg_stat_progress_repack</structname>, and is also visible via
@@ -7929,10 +7932,96 @@ FROM pg_stat_get_backend_idset() AS backendid;
</itemizedlist>
</para></entry>
</row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>current_index_relid</structfield> <type>oid</type>
+ </para>
+ <para>
+ If <command>VACUUM</command> is currently processing an index, this
+ column shows the OID of the index being vacuumed. The value is set
+ when the phase is <literal>vacuuming indexes</literal> or
+ <literal>cleaning up indexes</literal>, and is reset to 0 once that
+ index has been processed, so it does not show a stale index while the
+ backend is between indexes or has moved on to another phase. During
+ parallel index vacuum, each parallel worker row shows the index that
+ particular worker is processing.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>index_blks_total</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Total number of blocks in the index identified by
+ <structfield>current_index_relid</structfield>. This is reported only
+ for B-tree indexes, and only while the index is being scanned; it is 0
+ for other index access methods, for a B-tree whose scan
+ <command>VACUUM</command> was able to skip during cleanup, and once
+ the index has been processed.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>index_blks_done</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of blocks of the index identified by
+ <structfield>current_index_relid</structfield> scanned so far. This is
+ reported only for B-tree indexes, and only while the index is being
+ scanned; it is 0 for other index access methods, for a B-tree whose
+ scan <command>VACUUM</command> was able to skip during cleanup, and
+ once the index has been processed. Together with
+ <structfield>index_blks_total</structfield> this gives per-index vacuum
+ progress, which is useful for estimating completion of large indexes.
+ The index metapage is counted in
+ <structfield>index_blks_total</structfield> but is never scanned, so
+ this column stops one block short of it.
+ </para></entry>
+ </row>
+
</tbody>
</tgroup>
</table>
+ <note>
+ <para>
+ During a parallel vacuum, each participating backend (the leader and each
+ parallel worker) reports its own row. On a worker row the meaningful
+ columns are <structfield>pid</structfield>, <structfield>datid</structfield>,
+ <structfield>datname</structfield>, <structfield>relid</structfield>,
+ <structfield>phase</structfield>, <structfield>current_index_relid</structfield>,
+ <structfield>index_blks_total</structfield> and
+ <structfield>index_blks_done</structfield>. The remaining columns track
+ command-level heap progress that only the leader maintains; they read as
+ zero on worker rows. A participant row showing
+ <literal>vacuuming indexes</literal> or
+ <literal>cleaning up indexes</literal> with a zero
+ <structfield>current_index_relid</structfield> is between indexes, or has
+ finished its share of the indexes while other participants are still
+ working. A worker row with a <literal>NULL</literal>
+ <structfield>phase</structfield> is a worker that has been launched but has
+ not started on an index yet, which also happens when every index was
+ claimed by another participant before this worker got to it. Because a
+ table can be vacuumed by only one
+ <command>VACUUM</command> at a time, the rows sharing a given
+ <structfield>datid</structfield> and <structfield>relid</structfield>
+ together make up a single vacuum, so they can be grouped that way to see
+ the leader and all of its workers. The leader and worker process IDs can be
+ correlated with <structfield>leader_pid</structfield> in
+ <link linkend="monitoring-pg-stat-activity-view"><structname>pg_stat_activity</structname></link>
+ if needed. Because the worker rows are separate backends, a role that does
+ not have privileges of the <literal>pg_read_all_stats</literal> role and
+ does not own those backends sees only their <structfield>pid</structfield>,
+ <structfield>datid</structfield> and <structfield>datname</structfield>,
+ with the remaining columns <literal>NULL</literal>, so such a role can tell
+ that a parallel vacuum is running in a database but not which table or
+ index it is working on.
+ </para>
+ </note>
+
<table id="vacuum-phases">
<title>VACUUM Phases</title>
<tgroup cols="2">
diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index 063ef2208de..459fff8df1d 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -3038,16 +3038,26 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
{
IndexVacuumInfo ivinfo;
LVSavedErrInfo saved_err_info;
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.estimated_count = true;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3069,6 +3079,12 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
+
return istat;
}
@@ -3088,17 +3104,27 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
{
IndexVacuumInfo ivinfo;
LVSavedErrInfo saved_err_info;
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.estimated_count = estimated_count;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3118,6 +3144,12 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
+
return istat;
}
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index ad340887f54..809b9c0f1e4 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -1353,7 +1353,10 @@ CREATE VIEW pg_stat_progress_vacuum AS
CASE S.param13 WHEN 1 THEN 'manual'
WHEN 2 THEN 'autovacuum'
WHEN 3 THEN 'autovacuum_wraparound'
- ELSE NULL END AS started_by
+ ELSE NULL END AS started_by,
+ CAST(S.param14 AS oid) AS current_index_relid,
+ S.param16 AS index_blks_total,
+ S.param17 AS index_blks_done
FROM pg_stat_get_progress_info('VACUUM') AS S
LEFT JOIN pg_database D ON S.datid = D.oid;
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 767d162e578..be87b1937ad 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -1076,6 +1076,17 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
IndexBulkDeleteResult *istat = NULL;
IndexBulkDeleteResult *istat_res;
IndexVacuumInfo ivinfo;
+ const int progress_index[] = {
+ PROGRESS_VACUUM_PHASE,
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID
+ };
+ int64 progress_val[2];
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
/*
* Update the pointer to the corresponding bulk-deletion result if someone
@@ -1087,7 +1098,7 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
ivinfo.index = indrel;
ivinfo.heaprel = pvs->heaprel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.message_level = DEBUG2;
ivinfo.estimated_count = pvs->shared->estimated_count;
ivinfo.num_heap_tuples = pvs->shared->reltuples;
@@ -1097,6 +1108,16 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
pvs->indname = pstrdup(RelationGetRelationName(indrel));
pvs->status = indstats->status;
+ /*
+ * Report the phase and the index we're about to process before we start,
+ * so that it is visible for the whole duration of the index scan.
+ */
+ progress_val[0] = (indstats->status == PARALLEL_INDVAC_STATUS_NEED_BULKDELETE)
+ ? PROGRESS_VACUUM_PHASE_VACUUM_INDEX
+ : PROGRESS_VACUUM_PHASE_INDEX_CLEANUP;
+ progress_val[1] = (int64) RelationGetRelid(indrel);
+ pgstat_progress_update_multi_param(2, progress_index, progress_val);
+
switch (indstats->status)
{
case PARALLEL_INDVAC_STATUS_NEED_BULKDELETE:
@@ -1112,6 +1133,12 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
RelationGetRelationName(indrel));
}
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
+
/*
* Copy the index bulk-deletion result returned from ambulkdelete and
* amvacuumcleanup to the DSM segment if it's the first cycle because they
@@ -1315,6 +1342,9 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Prepare to track buffer usage during parallel execution */
InstrStartParallelQuery();
+ /* Register this worker for vacuum progress reporting */
+ pgstat_progress_start_command(PROGRESS_COMMAND_VACUUM, shared->relid);
+
/* Process indexes to perform vacuum/cleanup */
parallel_vacuum_process_safe_indexes(&pvs);
@@ -1334,6 +1364,9 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Pop the error context stack */
error_context_stack = errcallback.previous;
+ /* Unregister this worker from vacuum progress reporting */
+ pgstat_progress_end_command();
+
vac_close_indexes(nindexes, indrels, RowExclusiveLock);
table_close(rel, ShareUpdateExclusiveLock);
FreeAccessStrategy(pvs.bstrategy);
diff --git a/src/include/commands/progress.h b/src/include/commands/progress.h
index 2a12920c75f..bf91455eff9 100644
--- a/src/include/commands/progress.h
+++ b/src/include/commands/progress.h
@@ -31,6 +31,8 @@
#define PROGRESS_VACUUM_DELAY_TIME 10
#define PROGRESS_VACUUM_MODE 11
#define PROGRESS_VACUUM_STARTED_BY 12
+#define PROGRESS_VACUUM_CURRENT_INDEX_RELID 13
+/* 15 and 16 reserved for "block number" metrics */
/* Phases of vacuum (as advertised via PROGRESS_VACUUM_PHASE) */
#define PROGRESS_VACUUM_PHASE_SCAN_HEAP 1
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 4a8cc759d7b..0addd043e68 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -2214,7 +2214,10 @@ pg_stat_progress_vacuum| SELECT s.pid,
WHEN 2 THEN 'autovacuum'::text
WHEN 3 THEN 'autovacuum_wraparound'::text
ELSE NULL::text
- END AS started_by
+ END AS started_by,
+ (s.param14)::oid AS current_index_relid,
+ s.param16 AS index_blks_total,
+ s.param17 AS index_blks_done
FROM (pg_stat_get_progress_info('VACUUM'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
LEFT JOIN pg_database d ON ((s.datid = d.oid)));
pg_stat_recovery| SELECT promote_triggered,
--
2.47.3
[application/x-patch] v6-0002-Remove-IndexVacuumInfo.report_progress.patch (4.7K, ../../CALj2ACV6vuAYVnLqHR62ggMtDfD0YTomW_YxoXL0dMMneq0bxQ@mail.gmail.com/3-v6-0002-Remove-IndexVacuumInfo.report_progress.patch)
download | inline diff:
From e25111f8b4142e2087ee0e38b23b406bef74f871 Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Sun, 13 Sep 2026 06:58:05 +0000
Subject: [PATCH v6 2/2] Remove IndexVacuumInfo.report_progress.
Commit ab0dfc961b6a, which added progress reporting for CREATE
INDEX, introduced report_progress so that the block counters
PROGRESS_SCAN_BLOCKS_TOTAL and PROGRESS_SCAN_BLOCKS_DONE were
reported only during index validation. Commit XXXX now enables
vacuum to report them as well, so every caller sets the field to
true. B-tree is the only access method that reads it.
Remove the field and report the counters unconditionally in
btvacuumscan(). An out-of-tree index access method that sets
report_progress needs a trivial adjustment.
Author: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Discussion: https://postgr.es/m/CALj2ACUgwSchK6jQ2CdKLBWUADTOE_zKdTff2Zg3E6hOuXKv-w@mail.gmail.com
---
src/backend/access/heap/vacuumlazy.c | 2 --
src/backend/access/nbtree/nbtree.c | 9 +++------
src/backend/catalog/index.c | 1 -
src/backend/commands/vacuumparallel.c | 1 -
src/include/access/genam.h | 1 -
5 files changed, 3 insertions(+), 11 deletions(-)
diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index 459fff8df1d..dd6a37f65c1 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -3048,7 +3048,6 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = true;
ivinfo.estimated_count = true;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = reltuples;
@@ -3114,7 +3113,6 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = true;
ivinfo.estimated_count = estimated_count;
ivinfo.message_level = DEBUG2;
diff --git a/src/backend/access/nbtree/nbtree.c b/src/backend/access/nbtree/nbtree.c
index 0abdd7b49f5..a4c3ad1b0f6 100644
--- a/src/backend/access/nbtree/nbtree.c
+++ b/src/backend/access/nbtree/nbtree.c
@@ -1339,9 +1339,7 @@ btvacuumscan(IndexVacuumInfo *info, IndexBulkDeleteResult *stats,
if (needLock)
UnlockRelationForExtension(rel, ExclusiveLock);
- if (info->report_progress)
- pgstat_progress_update_param(PROGRESS_SCAN_BLOCKS_TOTAL,
- num_pages);
+ pgstat_progress_update_param(PROGRESS_SCAN_BLOCKS_TOTAL, num_pages);
/* Quit if we've scanned the whole relation */
if (p.current_blocknum >= num_pages)
@@ -1365,9 +1363,8 @@ btvacuumscan(IndexVacuumInfo *info, IndexBulkDeleteResult *stats,
current_block = btvacuumpage(&vstate, buf);
- if (info->report_progress)
- pgstat_progress_update_param(PROGRESS_SCAN_BLOCKS_DONE,
- current_block);
+ pgstat_progress_update_param(PROGRESS_SCAN_BLOCKS_DONE,
+ current_block);
}
/*
diff --git a/src/backend/catalog/index.c b/src/backend/catalog/index.c
index ec21b83b6b8..a3768e8daf7 100644
--- a/src/backend/catalog/index.c
+++ b/src/backend/catalog/index.c
@@ -3455,7 +3455,6 @@ validate_index(Oid heapId, Oid indexId, Snapshot snapshot)
ivinfo.index = indexRelation;
ivinfo.heaprel = heapRelation;
ivinfo.analyze_only = false;
- ivinfo.report_progress = true;
ivinfo.estimated_count = true;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = heapRelation->rd_rel->reltuples;
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index be87b1937ad..74f749de064 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -1098,7 +1098,6 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
ivinfo.index = indrel;
ivinfo.heaprel = pvs->heaprel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = true;
ivinfo.message_level = DEBUG2;
ivinfo.estimated_count = pvs->shared->estimated_count;
ivinfo.num_heap_tuples = pvs->shared->reltuples;
diff --git a/src/include/access/genam.h b/src/include/access/genam.h
index 68bfe405db3..64a44889ceb 100644
--- a/src/include/access/genam.h
+++ b/src/include/access/genam.h
@@ -54,7 +54,6 @@ typedef struct IndexVacuumInfo
Relation index; /* the index being vacuumed */
Relation heaprel; /* the heap relation the index belongs to */
bool analyze_only; /* ANALYZE (without any actual vacuum) */
- bool report_progress; /* emit progress.h status reports */
bool estimated_count; /* num_heap_tuples is an estimate */
int message_level; /* ereport level for progress messages */
double num_heap_tuples; /* tuples remaining in heap */
--
2.47.3
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-20 02:37 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-20 23:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-25 16:34 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 02:23 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 05:15 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-09-03 16:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih.pg@gmail.com>
2026-09-09 19:01 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Masahiko Sawada <sawada.mshk@gmail.com>
2026-09-14 20:28 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
@ 2026-09-15 17:10 ` Sami Imseih <samimseih.pg@gmail.com>
2026-09-15 22:28 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
0 siblings, 1 reply; 32+ messages in thread
From: Sami Imseih @ 2026-09-15 17:10 UTC (permalink / raw)
To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: Masahiko Sawada <sawada.mshk@gmail.com>; Michael Paquier <michael@paquier.xyz>; Sami Imseih <samimseih@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
Hi,
> I think it's fine for a progress view to have an entry per backend,
> including parallel workers, so more than one row for a single command.
My worry is that if we do not separate high level command stats from
per-worker (or leader) stats, we will end up with a confusing view and
tie our hands in the future.
For example, in v6 we have heap_blks_total and heap_blks_scanned, which
represent aggregate progress for how much of the table has been scanned,
while index_blks_total and index_blks_done represent how much one worker
has processed for a specific index. This is probably OK today only
because we do not support parallel heap scan.
But if we ever support parallel heap scan during VACUUM, what would
heap_blks_total and heap_blks_scanned represent at that point? Would
they remain command-level aggregate progress, or become per-worker
progress?
My point is that the existing progress views have historically exposed
command-level aggregate stats. Trying to also use them for per-worker
stats in the same view does not seem right to me.
> If we ever implement parallel index rebuilding for REINDEX TABLE,
> where different workers rebuild different indexes, having one entry
> per worker in pg_stat_progress_create_index seems the straightforward
> thing to do.
For parallel REINDEX TABLE, a command-level progress model with
indexes_total and
indexes_done would seem more natural, similar to VACUUM.
One separate question is whether each individual index rebuild would
itself run in parallel.
Today, an individual REINDEX INDEX may already use parallel build, but
that is leader-driven
parallelism for one index. Parallel workers do not in turn spawn more
parallel workers. So
if REINDEX TABLE were implemented by assigning different indexes to
different workers, It seems
likely each worker’s assigned index rebuild to be serial unless we
introduced a more complex
nested-parallel design.
> A dedicated view for workers, say
> pg_stat_progress_create_index_worker, would end up with mostly the
> same columns as pg_stat_progress_create_index.
That may be the case, but at least the meaning of each view will be well
understood and clear. Not mixing per-command vs per-worker stats.
> I can also see Sami's point that the patch makes many of the existing
> columns of pg_stat_progress_vacuum no-op. But I'm not sure a dedicated
> worker view really avoids that. If a user wants the progress of one
> vacuum command they would join the two views, get one row per
> participant anyway, and the leader's heap columns would be repeated on
> every row.
I don't really think most users will actually perform this join. They would
start with the command-level view and if they need to understand more
of what the worker is doing they can query a separate worker specific view.
So my hang up is the semantic differences each row will carry and
how hard it may be for a (monitoring) user to reason about this. It is
also clear that others are not as worried about this as me, so I will
concede, because getting this information is important.
As far as v6: The code looks overall good to me, but I have some
comments.
1/
I do think it will be better to do one more split of v6-0001 to separate
current_index_relid and index_blks_*; they are 2 distinct features. Also
you can fold the 0002 into the index_blks_* commit.
Also, there are some comment updates still needed:
2/
+ <structfield>index_blks_done</structfield>. The remaining columns track
+ command-level heap progress that only the leader maintains; they read as
+ zero on worker rows.
We should mention mode and started_by as NULL, such as:
".... read as zero on worker rows, except that
<structfield>mode</structfield> and
<structfield>started_by</structfield> appear as <literal>NULL</literal>."
3/
+ working. A worker row with a <literal>NULL</literal>
+ <structfield>phase</structfield> is a worker that has been launched but has
+ not started on an index yet, which also happens when every index was
+ claimed by another participant before this worker got to it. Because a
+ table can be vacuumed by only one
I don't think phase will ever be NULL here since phase will be set to
"initializing" at minimum.
4/
/*
* Perform work within a launched parallel process.
*
* Since parallel vacuum workers perform only index vacuum or index cleanup,
* we don't need to report progress information.
*/
void
parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
{
This comment is now out-of-date and should be updated.
--
Sami Imseih
Amazon Web Services (AWS)
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-20 02:37 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-20 23:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-25 16:34 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 02:23 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 05:15 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-09-03 16:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih.pg@gmail.com>
2026-09-09 19:01 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Masahiko Sawada <sawada.mshk@gmail.com>
2026-09-14 20:28 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-15 17:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih.pg@gmail.com>
@ 2026-09-15 22:28 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-17 02:23 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum shihao zhong <zhong950419@gmail.com>
2026-09-17 23:00 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
0 siblings, 2 replies; 32+ messages in thread
From: Bharath Rupireddy @ 2026-09-15 22:28 UTC (permalink / raw)
To: Sami Imseih <samimseih.pg@gmail.com>; +Cc: Masahiko Sawada <sawada.mshk@gmail.com>; Michael Paquier <michael@paquier.xyz>; Sami Imseih <samimseih@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
Hi,
On Tue, Sep 15, 2026 at 10:11 AM Sami Imseih <samimseih.pg@gmail.com> wrote:
>
> As far as v6: The code looks overall good to me, but I have some
> comments.
Thanks for reviewing.
> 1/
>
> I do think it will be better to do one more split of v6-0001 to separate
> current_index_relid and index_blks_*; they are 2 distinct features. Also
> you can fold the 0002 into the index_blks_* commit.
I prefer to keep them as one patch as I see them as closely related,
but I'm open to splitting if anyone thinks otherwise.
> Also, there are some comment updates still needed:
> 2/
>
> + <structfield>index_blks_done</structfield>. The remaining columns track
> + command-level heap progress that only the leader maintains; they read as
> + zero on worker rows.
>
> We should mention mode and started_by as NULL, such as:
>
> ".... read as zero on worker rows, except that
> <structfield>mode</structfield> and
> <structfield>started_by</structfield> appear as <literal>NULL</literal>."
>
> 3/
>
> + working. A worker row with a <literal>NULL</literal>
> + <structfield>phase</structfield> is a worker that has been launched but has
> + not started on an index yet, which also happens when every index was
> + claimed by another participant before this worker got to it. Because a
> + table can be vacuumed by only one
>
> I don't think phase will ever be NULL here since phase will be set to
> "initializing" at minimum.
>
> 4/
>
> /*
> * Perform work within a launched parallel process.
> *
> * Since parallel vacuum workers perform only index vacuum or index cleanup,
> * we don't need to report progress information.
> */
> void
> parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
> {
>
> This comment is now out-of-date and should be updated.
I agree with all three comments above and have updated the v6 patches
accordingly. Please have a look.
--
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com
Attachments:
[application/x-patch] v7-0001-Report-per-index-vacuum-progress-in-pg_stat_progr.patch (17.8K, ../../CALj2ACXabJxqApcYA-ghLuhg3iryyi-9GHVFwXh27B3TRWSnLQ@mail.gmail.com/2-v7-0001-Report-per-index-vacuum-progress-in-pg_stat_progr.patch)
download | inline diff:
From 4acb49d16e43f7f19b7c2cdd38be6001d43debf4 Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Sun, 13 Sep 2026 05:19:26 +0000
Subject: [PATCH v7 1/2] Report per-index vacuum progress in
pg_stat_progress_vacuum.
Previously, pg_stat_progress_vacuum showed the total and
processed index counts, but neither which index a backend was
working on nor how far it had gotten through it. On a table with
many indexes of different access methods, that made it hard to
tell which index a slow or stuck vacuum was spending its time on,
and for a large B-tree, at the scale of hundreds of GBs to TBs,
there was no way to estimate when the index phase, and together
with heap_blks_*, the whole vacuum would finish.
This commit adds three columns. current_index_relid reports the
OID of the index a backend is currently vacuuming or cleaning up.
index_blks_total and index_blks_done report block progress within
that index. All three are set before an index is processed and
reset once that index is done, so they do not report a stale
value after the phase moves on.
During parallel index vacuum the leader also participates, and
each participant processes a different set of indexes. That
per-index progress cannot be collapsed into a single leader row
without losing the detail that matters, so each participant, the
leader and the workers alike, now reports its own row and shows
the index it is processing and how far along it is. A worker row
carries the columns that come from its own backend state, pid,
datid, datname, relid, phase, current_index_relid,
index_blks_total and index_blks_done; the remaining columns track
command-level progress that the leader maintains and read as zero
on worker rows, or NULL for mode and started_by. Because a table
can be vacuumed by only one VACUUM at a time, the rows sharing a
datid and relid make up a single vacuum, so they can be grouped
that way to see the leader together with all of its workers. If
the exact leader-to-worker mapping is wanted, it can be recovered
from leader_pid in pg_stat_activity.
The block counters are the ones CREATE INDEX progress reporting
added in commit ab0dfc961b6a, PROGRESS_SCAN_BLOCKS_TOTAL and
PROGRESS_SCAN_BLOCKS_DONE. B-tree's index scan already knows how
to report them, so all that is needed here is turning that
reporting on in the serial and parallel index vacuum paths; other
access methods report nothing and leave the two columns at zero.
They are deliberately kept separate from heap_blks_*, which must
be retained across a multi-pass index vacuum, the case where the
dead-TID store fills, and which in the serial case belong to the
same backend that is doing the index vacuuming, so reusing them
for index blocks would destroy heap progress that is still needed.
No caller sets IndexVacuumInfo.report_progress to false anymore;
the next commit removes the field.
XXX: Bump catalog version.
Author: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Reviewed-by: Michael Paquier <michael@paquier.xyz>
Reviewed-by: Sami Imseih <samimseih@gmail.com>
Reviewed-by: Masahiko Sawada <sawada.mshk@gmail.com>
Discussion: https://postgr.es/m/CALj2ACUgwSchK6jQ2CdKLBWUADTOE_zKdTff2Zg3E6hOuXKv-w@mail.gmail.com
---
doc/src/sgml/monitoring.sgml | 95 ++++++++++++++++++++++++++-
src/backend/access/heap/vacuumlazy.c | 36 +++++++++-
src/backend/catalog/system_views.sql | 5 +-
src/backend/commands/vacuumparallel.c | 37 ++++++++++-
src/include/commands/progress.h | 2 +
src/test/regress/expected/rules.out | 5 +-
6 files changed, 172 insertions(+), 8 deletions(-)
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index 62dadf3e86c..4c49aac7003 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -7670,8 +7670,11 @@ FROM pg_stat_get_backend_idset() AS backendid;
<para>
Whenever <command>VACUUM</command> is running, the
<structname>pg_stat_progress_vacuum</structname> view will contain
- one row for each backend (including autovacuum worker processes) that is
- currently vacuuming. The tables below describe the information
+ one row for each backend (including autovacuum worker processes and
+ parallel workers launched for
+ <link linkend="sql-vacuum">parallel vacuum</link>) that is currently
+ vacuuming; see the note following the view for the columns reported on
+ parallel worker rows. The tables below describe the information
that will be reported and provide information about how to interpret it.
Progress for <command>VACUUM FULL</command> commands is reported via
<structname>pg_stat_progress_repack</structname>, and is also visible via
@@ -7929,10 +7932,98 @@ FROM pg_stat_get_backend_idset() AS backendid;
</itemizedlist>
</para></entry>
</row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>current_index_relid</structfield> <type>oid</type>
+ </para>
+ <para>
+ If <command>VACUUM</command> is currently processing an index, this
+ column shows the OID of the index being vacuumed. The value is set
+ when the phase is <literal>vacuuming indexes</literal> or
+ <literal>cleaning up indexes</literal>, and is reset to 0 once that
+ index has been processed, so it does not show a stale index while the
+ backend is between indexes or has moved on to another phase. During
+ parallel index vacuum, each parallel worker row shows the index that
+ particular worker is processing.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>index_blks_total</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Total number of blocks in the index identified by
+ <structfield>current_index_relid</structfield>. This is reported only
+ for B-tree indexes, and only while the index is being scanned; it is 0
+ for other index access methods, for a B-tree whose scan
+ <command>VACUUM</command> was able to skip during cleanup, and once
+ the index has been processed.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>index_blks_done</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of blocks of the index identified by
+ <structfield>current_index_relid</structfield> scanned so far. This is
+ reported only for B-tree indexes, and only while the index is being
+ scanned; it is 0 for other index access methods, for a B-tree whose
+ scan <command>VACUUM</command> was able to skip during cleanup, and
+ once the index has been processed. Together with
+ <structfield>index_blks_total</structfield> this gives per-index vacuum
+ progress, which is useful for estimating completion of large indexes.
+ The index metapage is counted in
+ <structfield>index_blks_total</structfield> but is never scanned, so
+ this column stops one block short of it.
+ </para></entry>
+ </row>
+
</tbody>
</tgroup>
</table>
+ <note>
+ <para>
+ During a parallel vacuum, each participating backend (the leader and each
+ parallel worker) reports its own row. On a worker row the meaningful
+ columns are <structfield>pid</structfield>, <structfield>datid</structfield>,
+ <structfield>datname</structfield>, <structfield>relid</structfield>,
+ <structfield>phase</structfield>, <structfield>current_index_relid</structfield>,
+ <structfield>index_blks_total</structfield> and
+ <structfield>index_blks_done</structfield>. The remaining columns track
+ command-level heap progress that only the leader maintains; they read as
+ zero on worker rows, except that <structfield>mode</structfield> and
+ <structfield>started_by</structfield> appear as <literal>NULL</literal>.
+ A participant row showing
+ <literal>vacuuming indexes</literal> or
+ <literal>cleaning up indexes</literal> with a zero
+ <structfield>current_index_relid</structfield> is between indexes, or has
+ finished its share of the indexes while other participants are still
+ working. A worker row whose <structfield>phase</structfield> is
+ <literal>initializing</literal> is a worker that has been launched but has
+ not started on an index yet, which also happens when every index was
+ claimed by another participant before this worker got to it. Because a
+ table can be vacuumed by only one
+ <command>VACUUM</command> at a time, the rows sharing a given
+ <structfield>datid</structfield> and <structfield>relid</structfield>
+ together make up a single vacuum, so they can be grouped that way to see
+ the leader and all of its workers. The leader and worker process IDs can be
+ correlated with <structfield>leader_pid</structfield> in
+ <link linkend="monitoring-pg-stat-activity-view"><structname>pg_stat_activity</structname></link>
+ if needed. Because the worker rows are separate backends, a role that does
+ not have privileges of the <literal>pg_read_all_stats</literal> role and
+ does not own those backends sees only their <structfield>pid</structfield>,
+ <structfield>datid</structfield> and <structfield>datname</structfield>,
+ with the remaining columns <literal>NULL</literal>, so such a role can tell
+ that a parallel vacuum is running in a database but not which table or
+ index it is working on.
+ </para>
+ </note>
+
<table id="vacuum-phases">
<title>VACUUM Phases</title>
<tgroup cols="2">
diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index 8e1f660bc2f..a23af936b7f 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -3038,16 +3038,26 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
{
IndexVacuumInfo ivinfo;
LVSavedErrInfo saved_err_info;
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.estimated_count = true;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3069,6 +3079,12 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
+
return istat;
}
@@ -3088,17 +3104,27 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
{
IndexVacuumInfo ivinfo;
LVSavedErrInfo saved_err_info;
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.estimated_count = estimated_count;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3118,6 +3144,12 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
+
return istat;
}
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index ad340887f54..809b9c0f1e4 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -1353,7 +1353,10 @@ CREATE VIEW pg_stat_progress_vacuum AS
CASE S.param13 WHEN 1 THEN 'manual'
WHEN 2 THEN 'autovacuum'
WHEN 3 THEN 'autovacuum_wraparound'
- ELSE NULL END AS started_by
+ ELSE NULL END AS started_by,
+ CAST(S.param14 AS oid) AS current_index_relid,
+ S.param16 AS index_blks_total,
+ S.param17 AS index_blks_done
FROM pg_stat_get_progress_info('VACUUM') AS S
LEFT JOIN pg_database D ON S.datid = D.oid;
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 767d162e578..607c57c8eba 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -1076,6 +1076,17 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
IndexBulkDeleteResult *istat = NULL;
IndexBulkDeleteResult *istat_res;
IndexVacuumInfo ivinfo;
+ const int progress_index[] = {
+ PROGRESS_VACUUM_PHASE,
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID
+ };
+ int64 progress_val[2];
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
/*
* Update the pointer to the corresponding bulk-deletion result if someone
@@ -1087,7 +1098,7 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
ivinfo.index = indrel;
ivinfo.heaprel = pvs->heaprel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.message_level = DEBUG2;
ivinfo.estimated_count = pvs->shared->estimated_count;
ivinfo.num_heap_tuples = pvs->shared->reltuples;
@@ -1097,6 +1108,16 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
pvs->indname = pstrdup(RelationGetRelationName(indrel));
pvs->status = indstats->status;
+ /*
+ * Report the phase and the index we're about to process before we start,
+ * so that it is visible for the whole duration of the index scan.
+ */
+ progress_val[0] = (indstats->status == PARALLEL_INDVAC_STATUS_NEED_BULKDELETE)
+ ? PROGRESS_VACUUM_PHASE_VACUUM_INDEX
+ : PROGRESS_VACUUM_PHASE_INDEX_CLEANUP;
+ progress_val[1] = (int64) RelationGetRelid(indrel);
+ pgstat_progress_update_multi_param(2, progress_index, progress_val);
+
switch (indstats->status)
{
case PARALLEL_INDVAC_STATUS_NEED_BULKDELETE:
@@ -1112,6 +1133,12 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
RelationGetRelationName(indrel));
}
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
+
/*
* Copy the index bulk-deletion result returned from ambulkdelete and
* amvacuumcleanup to the DSM segment if it's the first cycle because they
@@ -1191,7 +1218,7 @@ parallel_vacuum_index_is_parallel_safe(Relation indrel, int num_index_scans,
* Perform work within a launched parallel process.
*
* Since parallel vacuum workers perform only index vacuum or index cleanup,
- * we don't need to report progress information.
+ * they report progress for the index they are processing.
*/
void
parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
@@ -1315,6 +1342,9 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Prepare to track buffer usage during parallel execution */
InstrStartParallelQuery();
+ /* Register this worker for vacuum progress reporting */
+ pgstat_progress_start_command(PROGRESS_COMMAND_VACUUM, shared->relid);
+
/* Process indexes to perform vacuum/cleanup */
parallel_vacuum_process_safe_indexes(&pvs);
@@ -1334,6 +1364,9 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Pop the error context stack */
error_context_stack = errcallback.previous;
+ /* Unregister this worker from vacuum progress reporting */
+ pgstat_progress_end_command();
+
vac_close_indexes(nindexes, indrels, RowExclusiveLock);
table_close(rel, ShareUpdateExclusiveLock);
FreeAccessStrategy(pvs.bstrategy);
diff --git a/src/include/commands/progress.h b/src/include/commands/progress.h
index 2a12920c75f..bf91455eff9 100644
--- a/src/include/commands/progress.h
+++ b/src/include/commands/progress.h
@@ -31,6 +31,8 @@
#define PROGRESS_VACUUM_DELAY_TIME 10
#define PROGRESS_VACUUM_MODE 11
#define PROGRESS_VACUUM_STARTED_BY 12
+#define PROGRESS_VACUUM_CURRENT_INDEX_RELID 13
+/* 15 and 16 reserved for "block number" metrics */
/* Phases of vacuum (as advertised via PROGRESS_VACUUM_PHASE) */
#define PROGRESS_VACUUM_PHASE_SCAN_HEAP 1
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 4a8cc759d7b..0addd043e68 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -2214,7 +2214,10 @@ pg_stat_progress_vacuum| SELECT s.pid,
WHEN 2 THEN 'autovacuum'::text
WHEN 3 THEN 'autovacuum_wraparound'::text
ELSE NULL::text
- END AS started_by
+ END AS started_by,
+ (s.param14)::oid AS current_index_relid,
+ s.param16 AS index_blks_total,
+ s.param17 AS index_blks_done
FROM (pg_stat_get_progress_info('VACUUM'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
LEFT JOIN pg_database d ON ((s.datid = d.oid)));
pg_stat_recovery| SELECT promote_triggered,
--
2.47.3
[application/x-patch] v7-0002-Remove-IndexVacuumInfo.report_progress.patch (4.7K, ../../CALj2ACXabJxqApcYA-ghLuhg3iryyi-9GHVFwXh27B3TRWSnLQ@mail.gmail.com/3-v7-0002-Remove-IndexVacuumInfo.report_progress.patch)
download | inline diff:
From 0a7a0dec24b3f1143d2c3c240b0f823d750b86a6 Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Sun, 13 Sep 2026 06:58:05 +0000
Subject: [PATCH v7 2/2] Remove IndexVacuumInfo.report_progress.
Commit ab0dfc961b6a, which added progress reporting for CREATE
INDEX, introduced report_progress so that the block counters
PROGRESS_SCAN_BLOCKS_TOTAL and PROGRESS_SCAN_BLOCKS_DONE were
reported only during index validation. Commit XXXX now enables
vacuum to report them as well, so every caller sets the field to
true. B-tree is the only access method that reads it.
Remove the field and report the counters unconditionally in
btvacuumscan(). An out-of-tree index access method that sets
report_progress needs a trivial adjustment.
Author: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Discussion: https://postgr.es/m/CALj2ACUgwSchK6jQ2CdKLBWUADTOE_zKdTff2Zg3E6hOuXKv-w@mail.gmail.com
---
src/backend/access/heap/vacuumlazy.c | 2 --
src/backend/access/nbtree/nbtree.c | 9 +++------
src/backend/catalog/index.c | 1 -
src/backend/commands/vacuumparallel.c | 1 -
src/include/access/genam.h | 1 -
5 files changed, 3 insertions(+), 11 deletions(-)
diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index a23af936b7f..44a7466def4 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -3048,7 +3048,6 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = true;
ivinfo.estimated_count = true;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = reltuples;
@@ -3114,7 +3113,6 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = true;
ivinfo.estimated_count = estimated_count;
ivinfo.message_level = DEBUG2;
diff --git a/src/backend/access/nbtree/nbtree.c b/src/backend/access/nbtree/nbtree.c
index 0abdd7b49f5..a4c3ad1b0f6 100644
--- a/src/backend/access/nbtree/nbtree.c
+++ b/src/backend/access/nbtree/nbtree.c
@@ -1339,9 +1339,7 @@ btvacuumscan(IndexVacuumInfo *info, IndexBulkDeleteResult *stats,
if (needLock)
UnlockRelationForExtension(rel, ExclusiveLock);
- if (info->report_progress)
- pgstat_progress_update_param(PROGRESS_SCAN_BLOCKS_TOTAL,
- num_pages);
+ pgstat_progress_update_param(PROGRESS_SCAN_BLOCKS_TOTAL, num_pages);
/* Quit if we've scanned the whole relation */
if (p.current_blocknum >= num_pages)
@@ -1365,9 +1363,8 @@ btvacuumscan(IndexVacuumInfo *info, IndexBulkDeleteResult *stats,
current_block = btvacuumpage(&vstate, buf);
- if (info->report_progress)
- pgstat_progress_update_param(PROGRESS_SCAN_BLOCKS_DONE,
- current_block);
+ pgstat_progress_update_param(PROGRESS_SCAN_BLOCKS_DONE,
+ current_block);
}
/*
diff --git a/src/backend/catalog/index.c b/src/backend/catalog/index.c
index ec21b83b6b8..a3768e8daf7 100644
--- a/src/backend/catalog/index.c
+++ b/src/backend/catalog/index.c
@@ -3455,7 +3455,6 @@ validate_index(Oid heapId, Oid indexId, Snapshot snapshot)
ivinfo.index = indexRelation;
ivinfo.heaprel = heapRelation;
ivinfo.analyze_only = false;
- ivinfo.report_progress = true;
ivinfo.estimated_count = true;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = heapRelation->rd_rel->reltuples;
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 607c57c8eba..fe388d99dc3 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -1098,7 +1098,6 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
ivinfo.index = indrel;
ivinfo.heaprel = pvs->heaprel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = true;
ivinfo.message_level = DEBUG2;
ivinfo.estimated_count = pvs->shared->estimated_count;
ivinfo.num_heap_tuples = pvs->shared->reltuples;
diff --git a/src/include/access/genam.h b/src/include/access/genam.h
index 72b256ecf43..07447bf1c4f 100644
--- a/src/include/access/genam.h
+++ b/src/include/access/genam.h
@@ -54,7 +54,6 @@ typedef struct IndexVacuumInfo
Relation index; /* the index being vacuumed */
Relation heaprel; /* the heap relation the index belongs to */
bool analyze_only; /* ANALYZE (without any actual vacuum) */
- bool report_progress; /* emit progress.h status reports */
bool estimated_count; /* num_heap_tuples is an estimate */
int message_level; /* ereport level for progress messages */
double num_heap_tuples; /* tuples remaining in heap */
--
2.47.3
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-20 02:37 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-20 23:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-25 16:34 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 02:23 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 05:15 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-09-03 16:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih.pg@gmail.com>
2026-09-09 19:01 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Masahiko Sawada <sawada.mshk@gmail.com>
2026-09-14 20:28 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-15 17:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih.pg@gmail.com>
2026-09-15 22:28 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
@ 2026-09-17 02:23 ` shihao zhong <zhong950419@gmail.com>
1 sibling, 0 replies; 32+ messages in thread
From: shihao zhong @ 2026-09-17 02:23 UTC (permalink / raw)
To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: Sami Imseih <samimseih.pg@gmail.com>; Masahiko Sawada <sawada.mshk@gmail.com>; Michael Paquier <michael@paquier.xyz>; Sami Imseih <samimseih@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
Hi Bharath,
v7 looks good here. Tested sequential, parallel on btree and GIN, all as the
docs describe, including GIN reporting zero blocks and the reset
clearing the previous index. Applies cleanly on 5e595d659fd, suites
pass.
You were right that this needs an injection point, polling is not
reliable when the value is set and reset around every index. One thing
to add: the sequential path needs one too, not just the parallel
workers, since lazy_vacuum_one_index sets and resets just as tightly. A
single point in the per index loop covers both.
Attached as two patches, take either half. v7-0003 is the injection
point, v7-0004 a TAP test in test_misc. Runs in about a second, under
both meson and make. Sanity check, it does fail if you take out the
progress update.
It does not cover the reset to 0, only that the value moves on. That
needs a second point after the reset.
Thanks,
Shihao
Attachments:
[application/octet-stream] v7-0003-Add-an-injection-point-in-the-per-index-vacuum-lo.patch (1.2K, ../../CAGRkXqSYB-LBEZmfu9qmj75CC_3gthKqD3mYw1yYZWRSWoOopQ@mail.gmail.com/3-v7-0003-Add-an-injection-point-in-the-per-index-vacuum-lo.patch)
download | inline diff:
From dacd39c1fd9583a6592b8ffb7e5f616e4a05c069 Mon Sep 17 00:00:00 2001
From: Shihao Zhong <zhong950419@gmail.com>
Date: Wed, 16 Sep 2026 21:56:16 -0400
Subject: [PATCH v7 3/4] Add an injection point in the per index vacuum loop
current_index_relid in pg_stat_progress_vacuum is set and reset around
each index, so a test that polls the view can miss it entirely. Add an
injection point just after the value is set, so a test can stop with a
known index in progress.
---
src/backend/access/heap/vacuumlazy.c | 9 +++++++++
1 file changed, 9 insertions(+)
diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index 44a7466def4..14e85e55077 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -3057,6 +3057,15 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
(int64) RelationGetRelid(indrel));
+#ifdef USE_INJECTION_POINTS
+
+ /*
+ * Used by tests to inspect pg_stat_progress_vacuum while a known index
+ * is in progress.
+ */
+ INJECTION_POINT("vacuum-index-in-progress", NULL);
+#endif
+
/*
* Update error traceback information.
*
--
2.37.1 (Apple Git-137.1)
[application/octet-stream] v7-0004-Add-a-TAP-test-for-per-index-vacuum-progress-repo.patch (3.7K, ../../CAGRkXqSYB-LBEZmfu9qmj75CC_3gthKqD3mYw1yYZWRSWoOopQ@mail.gmail.com/4-v7-0004-Add-a-TAP-test-for-per-index-vacuum-progress-repo.patch)
download | inline diff:
From bb32940d12df1cadcd94bb7b4a071044c837cc9b Mon Sep 17 00:00:00 2001
From: Shihao Zhong <zhong950419@gmail.com>
Date: Wed, 16 Sep 2026 21:56:17 -0400
Subject: [PATCH v7 4/4] Add a TAP test for per-index vacuum progress reporting
Use the injection point added by the previous commit to stop with a
known index in progress, and check that pg_stat_progress_vacuum names
that index and then moves on to the next one rather than carrying a
stale value over.
---
src/test/modules/test_misc/meson.build | 1 +
.../test_misc/t/016_vacuum_progress.pl | 69 +++++++++++++++++++
2 files changed, 70 insertions(+)
create mode 100644 src/test/modules/test_misc/t/016_vacuum_progress.pl
diff --git a/src/test/modules/test_misc/meson.build b/src/test/modules/test_misc/meson.build
index 5d81f5b13be..27b33a2c61e 100644
--- a/src/test/modules/test_misc/meson.build
+++ b/src/test/modules/test_misc/meson.build
@@ -24,6 +24,7 @@ tests += {
't/013_temp_obj_multisession.pl',
't/014_log_statement_max_length.pl',
't/015_temp_schema_exit_deferrable.pl',
+ 't/016_vacuum_progress.pl',
],
# The injection points are cluster-wide, so disable installcheck
'runningcheck': false,
diff --git a/src/test/modules/test_misc/t/016_vacuum_progress.pl b/src/test/modules/test_misc/t/016_vacuum_progress.pl
new file mode 100644
index 00000000000..525e7621afd
--- /dev/null
+++ b/src/test/modules/test_misc/t/016_vacuum_progress.pl
@@ -0,0 +1,69 @@
+# Copyright (c) 2026, PostgreSQL Global Development Group
+#
+# Check that pg_stat_progress_vacuum reports the index being vacuumed.
+
+use strict;
+use warnings FATAL => 'all';
+use PostgreSQL::Test::Cluster;
+use PostgreSQL::Test::Utils;
+use Test::More;
+
+if (($ENV{enable_injection_points} // 'no') ne 'yes')
+{
+ plan skip_all => 'Injection points not supported by this build';
+}
+
+my $node = PostgreSQL::Test::Cluster->new('vacprog');
+$node->init;
+$node->append_conf('postgresql.conf', 'autovacuum = off');
+$node->start;
+
+if (!$node->check_extension('injection_points'))
+{
+ plan skip_all => 'Extension injection_points not installed';
+}
+
+$node->safe_psql(
+ 'postgres', qq{
+ CREATE EXTENSION injection_points;
+ CREATE TABLE vacprog (a int, b int);
+ INSERT INTO vacprog SELECT g, g FROM generate_series(1, 20000) g;
+ CREATE INDEX vacprog_a ON vacprog (a);
+ CREATE INDEX vacprog_b ON vacprog (b);
+ DELETE FROM vacprog WHERE a % 2 = 0;
+ SELECT injection_points_attach('vacuum-index-in-progress', 'wait');
+});
+
+# Stop inside the per index loop, with one index known to be in progress.
+my $vac = $node->background_psql('postgres');
+$vac->query_until(qr//, "VACUUM vacprog;\n");
+$node->wait_for_event('client backend', 'vacuum-index-in-progress');
+
+my $first = $node->safe_psql(
+ 'postgres', q{
+ SELECT c.relname FROM pg_stat_progress_vacuum v
+ JOIN pg_class c ON c.oid = v.current_index_relid
+ WHERE v.relid = 'vacprog'::regclass});
+like($first, qr/^vacprog_[ab]$/, "reports the index being vacuumed, $first");
+
+$node->safe_psql('postgres',
+ "SELECT injection_points_wakeup('vacuum-index-in-progress');");
+
+# The other index must be reported in its turn. A value left over from the
+# first index would never satisfy this.
+ok( $node->poll_query_until(
+ 'postgres', qq{
+ SELECT count(*) = 1 FROM pg_stat_progress_vacuum v
+ JOIN pg_class c ON c.oid = v.current_index_relid
+ WHERE v.relid = 'vacprog'::regclass AND c.relname <> '$first'}),
+ 'reports the next index in turn');
+
+$node->safe_psql(
+ 'postgres', q{
+ SELECT injection_points_detach('vacuum-index-in-progress');
+ SELECT injection_points_wakeup('vacuum-index-in-progress');
+});
+$vac->quit;
+$node->stop;
+
+done_testing();
--
2.37.1 (Apple Git-137.1)
[application/octet-stream] v7-0001-Report-per-index-vacuum-progress-in-pg_stat_progr.patch (17.8K, ../../CAGRkXqSYB-LBEZmfu9qmj75CC_3gthKqD3mYw1yYZWRSWoOopQ@mail.gmail.com/5-v7-0001-Report-per-index-vacuum-progress-in-pg_stat_progr.patch)
download | inline diff:
From 4acb49d16e43f7f19b7c2cdd38be6001d43debf4 Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Sun, 13 Sep 2026 05:19:26 +0000
Subject: [PATCH v7 1/2] Report per-index vacuum progress in
pg_stat_progress_vacuum.
Previously, pg_stat_progress_vacuum showed the total and
processed index counts, but neither which index a backend was
working on nor how far it had gotten through it. On a table with
many indexes of different access methods, that made it hard to
tell which index a slow or stuck vacuum was spending its time on,
and for a large B-tree, at the scale of hundreds of GBs to TBs,
there was no way to estimate when the index phase, and together
with heap_blks_*, the whole vacuum would finish.
This commit adds three columns. current_index_relid reports the
OID of the index a backend is currently vacuuming or cleaning up.
index_blks_total and index_blks_done report block progress within
that index. All three are set before an index is processed and
reset once that index is done, so they do not report a stale
value after the phase moves on.
During parallel index vacuum the leader also participates, and
each participant processes a different set of indexes. That
per-index progress cannot be collapsed into a single leader row
without losing the detail that matters, so each participant, the
leader and the workers alike, now reports its own row and shows
the index it is processing and how far along it is. A worker row
carries the columns that come from its own backend state, pid,
datid, datname, relid, phase, current_index_relid,
index_blks_total and index_blks_done; the remaining columns track
command-level progress that the leader maintains and read as zero
on worker rows, or NULL for mode and started_by. Because a table
can be vacuumed by only one VACUUM at a time, the rows sharing a
datid and relid make up a single vacuum, so they can be grouped
that way to see the leader together with all of its workers. If
the exact leader-to-worker mapping is wanted, it can be recovered
from leader_pid in pg_stat_activity.
The block counters are the ones CREATE INDEX progress reporting
added in commit ab0dfc961b6a, PROGRESS_SCAN_BLOCKS_TOTAL and
PROGRESS_SCAN_BLOCKS_DONE. B-tree's index scan already knows how
to report them, so all that is needed here is turning that
reporting on in the serial and parallel index vacuum paths; other
access methods report nothing and leave the two columns at zero.
They are deliberately kept separate from heap_blks_*, which must
be retained across a multi-pass index vacuum, the case where the
dead-TID store fills, and which in the serial case belong to the
same backend that is doing the index vacuuming, so reusing them
for index blocks would destroy heap progress that is still needed.
No caller sets IndexVacuumInfo.report_progress to false anymore;
the next commit removes the field.
XXX: Bump catalog version.
Author: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Reviewed-by: Michael Paquier <michael@paquier.xyz>
Reviewed-by: Sami Imseih <samimseih@gmail.com>
Reviewed-by: Masahiko Sawada <sawada.mshk@gmail.com>
Discussion: https://postgr.es/m/CALj2ACUgwSchK6jQ2CdKLBWUADTOE_zKdTff2Zg3E6hOuXKv-w@mail.gmail.com
---
doc/src/sgml/monitoring.sgml | 95 ++++++++++++++++++++++++++-
src/backend/access/heap/vacuumlazy.c | 36 +++++++++-
src/backend/catalog/system_views.sql | 5 +-
src/backend/commands/vacuumparallel.c | 37 ++++++++++-
src/include/commands/progress.h | 2 +
src/test/regress/expected/rules.out | 5 +-
6 files changed, 172 insertions(+), 8 deletions(-)
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index 62dadf3e86c..4c49aac7003 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -7670,8 +7670,11 @@ FROM pg_stat_get_backend_idset() AS backendid;
<para>
Whenever <command>VACUUM</command> is running, the
<structname>pg_stat_progress_vacuum</structname> view will contain
- one row for each backend (including autovacuum worker processes) that is
- currently vacuuming. The tables below describe the information
+ one row for each backend (including autovacuum worker processes and
+ parallel workers launched for
+ <link linkend="sql-vacuum">parallel vacuum</link>) that is currently
+ vacuuming; see the note following the view for the columns reported on
+ parallel worker rows. The tables below describe the information
that will be reported and provide information about how to interpret it.
Progress for <command>VACUUM FULL</command> commands is reported via
<structname>pg_stat_progress_repack</structname>, and is also visible via
@@ -7929,10 +7932,98 @@ FROM pg_stat_get_backend_idset() AS backendid;
</itemizedlist>
</para></entry>
</row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>current_index_relid</structfield> <type>oid</type>
+ </para>
+ <para>
+ If <command>VACUUM</command> is currently processing an index, this
+ column shows the OID of the index being vacuumed. The value is set
+ when the phase is <literal>vacuuming indexes</literal> or
+ <literal>cleaning up indexes</literal>, and is reset to 0 once that
+ index has been processed, so it does not show a stale index while the
+ backend is between indexes or has moved on to another phase. During
+ parallel index vacuum, each parallel worker row shows the index that
+ particular worker is processing.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>index_blks_total</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Total number of blocks in the index identified by
+ <structfield>current_index_relid</structfield>. This is reported only
+ for B-tree indexes, and only while the index is being scanned; it is 0
+ for other index access methods, for a B-tree whose scan
+ <command>VACUUM</command> was able to skip during cleanup, and once
+ the index has been processed.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>index_blks_done</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of blocks of the index identified by
+ <structfield>current_index_relid</structfield> scanned so far. This is
+ reported only for B-tree indexes, and only while the index is being
+ scanned; it is 0 for other index access methods, for a B-tree whose
+ scan <command>VACUUM</command> was able to skip during cleanup, and
+ once the index has been processed. Together with
+ <structfield>index_blks_total</structfield> this gives per-index vacuum
+ progress, which is useful for estimating completion of large indexes.
+ The index metapage is counted in
+ <structfield>index_blks_total</structfield> but is never scanned, so
+ this column stops one block short of it.
+ </para></entry>
+ </row>
+
</tbody>
</tgroup>
</table>
+ <note>
+ <para>
+ During a parallel vacuum, each participating backend (the leader and each
+ parallel worker) reports its own row. On a worker row the meaningful
+ columns are <structfield>pid</structfield>, <structfield>datid</structfield>,
+ <structfield>datname</structfield>, <structfield>relid</structfield>,
+ <structfield>phase</structfield>, <structfield>current_index_relid</structfield>,
+ <structfield>index_blks_total</structfield> and
+ <structfield>index_blks_done</structfield>. The remaining columns track
+ command-level heap progress that only the leader maintains; they read as
+ zero on worker rows, except that <structfield>mode</structfield> and
+ <structfield>started_by</structfield> appear as <literal>NULL</literal>.
+ A participant row showing
+ <literal>vacuuming indexes</literal> or
+ <literal>cleaning up indexes</literal> with a zero
+ <structfield>current_index_relid</structfield> is between indexes, or has
+ finished its share of the indexes while other participants are still
+ working. A worker row whose <structfield>phase</structfield> is
+ <literal>initializing</literal> is a worker that has been launched but has
+ not started on an index yet, which also happens when every index was
+ claimed by another participant before this worker got to it. Because a
+ table can be vacuumed by only one
+ <command>VACUUM</command> at a time, the rows sharing a given
+ <structfield>datid</structfield> and <structfield>relid</structfield>
+ together make up a single vacuum, so they can be grouped that way to see
+ the leader and all of its workers. The leader and worker process IDs can be
+ correlated with <structfield>leader_pid</structfield> in
+ <link linkend="monitoring-pg-stat-activity-view"><structname>pg_stat_activity</structname></link>
+ if needed. Because the worker rows are separate backends, a role that does
+ not have privileges of the <literal>pg_read_all_stats</literal> role and
+ does not own those backends sees only their <structfield>pid</structfield>,
+ <structfield>datid</structfield> and <structfield>datname</structfield>,
+ with the remaining columns <literal>NULL</literal>, so such a role can tell
+ that a parallel vacuum is running in a database but not which table or
+ index it is working on.
+ </para>
+ </note>
+
<table id="vacuum-phases">
<title>VACUUM Phases</title>
<tgroup cols="2">
diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index 8e1f660bc2f..a23af936b7f 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -3038,16 +3038,26 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
{
IndexVacuumInfo ivinfo;
LVSavedErrInfo saved_err_info;
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.estimated_count = true;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3069,6 +3079,12 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
+
return istat;
}
@@ -3088,17 +3104,27 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
{
IndexVacuumInfo ivinfo;
LVSavedErrInfo saved_err_info;
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.estimated_count = estimated_count;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3118,6 +3144,12 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
+
return istat;
}
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index ad340887f54..809b9c0f1e4 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -1353,7 +1353,10 @@ CREATE VIEW pg_stat_progress_vacuum AS
CASE S.param13 WHEN 1 THEN 'manual'
WHEN 2 THEN 'autovacuum'
WHEN 3 THEN 'autovacuum_wraparound'
- ELSE NULL END AS started_by
+ ELSE NULL END AS started_by,
+ CAST(S.param14 AS oid) AS current_index_relid,
+ S.param16 AS index_blks_total,
+ S.param17 AS index_blks_done
FROM pg_stat_get_progress_info('VACUUM') AS S
LEFT JOIN pg_database D ON S.datid = D.oid;
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 767d162e578..607c57c8eba 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -1076,6 +1076,17 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
IndexBulkDeleteResult *istat = NULL;
IndexBulkDeleteResult *istat_res;
IndexVacuumInfo ivinfo;
+ const int progress_index[] = {
+ PROGRESS_VACUUM_PHASE,
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID
+ };
+ int64 progress_val[2];
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
/*
* Update the pointer to the corresponding bulk-deletion result if someone
@@ -1087,7 +1098,7 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
ivinfo.index = indrel;
ivinfo.heaprel = pvs->heaprel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.message_level = DEBUG2;
ivinfo.estimated_count = pvs->shared->estimated_count;
ivinfo.num_heap_tuples = pvs->shared->reltuples;
@@ -1097,6 +1108,16 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
pvs->indname = pstrdup(RelationGetRelationName(indrel));
pvs->status = indstats->status;
+ /*
+ * Report the phase and the index we're about to process before we start,
+ * so that it is visible for the whole duration of the index scan.
+ */
+ progress_val[0] = (indstats->status == PARALLEL_INDVAC_STATUS_NEED_BULKDELETE)
+ ? PROGRESS_VACUUM_PHASE_VACUUM_INDEX
+ : PROGRESS_VACUUM_PHASE_INDEX_CLEANUP;
+ progress_val[1] = (int64) RelationGetRelid(indrel);
+ pgstat_progress_update_multi_param(2, progress_index, progress_val);
+
switch (indstats->status)
{
case PARALLEL_INDVAC_STATUS_NEED_BULKDELETE:
@@ -1112,6 +1133,12 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
RelationGetRelationName(indrel));
}
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
+
/*
* Copy the index bulk-deletion result returned from ambulkdelete and
* amvacuumcleanup to the DSM segment if it's the first cycle because they
@@ -1191,7 +1218,7 @@ parallel_vacuum_index_is_parallel_safe(Relation indrel, int num_index_scans,
* Perform work within a launched parallel process.
*
* Since parallel vacuum workers perform only index vacuum or index cleanup,
- * we don't need to report progress information.
+ * they report progress for the index they are processing.
*/
void
parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
@@ -1315,6 +1342,9 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Prepare to track buffer usage during parallel execution */
InstrStartParallelQuery();
+ /* Register this worker for vacuum progress reporting */
+ pgstat_progress_start_command(PROGRESS_COMMAND_VACUUM, shared->relid);
+
/* Process indexes to perform vacuum/cleanup */
parallel_vacuum_process_safe_indexes(&pvs);
@@ -1334,6 +1364,9 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Pop the error context stack */
error_context_stack = errcallback.previous;
+ /* Unregister this worker from vacuum progress reporting */
+ pgstat_progress_end_command();
+
vac_close_indexes(nindexes, indrels, RowExclusiveLock);
table_close(rel, ShareUpdateExclusiveLock);
FreeAccessStrategy(pvs.bstrategy);
diff --git a/src/include/commands/progress.h b/src/include/commands/progress.h
index 2a12920c75f..bf91455eff9 100644
--- a/src/include/commands/progress.h
+++ b/src/include/commands/progress.h
@@ -31,6 +31,8 @@
#define PROGRESS_VACUUM_DELAY_TIME 10
#define PROGRESS_VACUUM_MODE 11
#define PROGRESS_VACUUM_STARTED_BY 12
+#define PROGRESS_VACUUM_CURRENT_INDEX_RELID 13
+/* 15 and 16 reserved for "block number" metrics */
/* Phases of vacuum (as advertised via PROGRESS_VACUUM_PHASE) */
#define PROGRESS_VACUUM_PHASE_SCAN_HEAP 1
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 4a8cc759d7b..0addd043e68 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -2214,7 +2214,10 @@ pg_stat_progress_vacuum| SELECT s.pid,
WHEN 2 THEN 'autovacuum'::text
WHEN 3 THEN 'autovacuum_wraparound'::text
ELSE NULL::text
- END AS started_by
+ END AS started_by,
+ (s.param14)::oid AS current_index_relid,
+ s.param16 AS index_blks_total,
+ s.param17 AS index_blks_done
FROM (pg_stat_get_progress_info('VACUUM'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
LEFT JOIN pg_database d ON ((s.datid = d.oid)));
pg_stat_recovery| SELECT promote_triggered,
--
2.47.3
[application/octet-stream] v7-0002-Remove-IndexVacuumInfo.report_progress.patch (4.7K, ../../CAGRkXqSYB-LBEZmfu9qmj75CC_3gthKqD3mYw1yYZWRSWoOopQ@mail.gmail.com/6-v7-0002-Remove-IndexVacuumInfo.report_progress.patch)
download | inline diff:
From 0a7a0dec24b3f1143d2c3c240b0f823d750b86a6 Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Sun, 13 Sep 2026 06:58:05 +0000
Subject: [PATCH v7 2/2] Remove IndexVacuumInfo.report_progress.
Commit ab0dfc961b6a, which added progress reporting for CREATE
INDEX, introduced report_progress so that the block counters
PROGRESS_SCAN_BLOCKS_TOTAL and PROGRESS_SCAN_BLOCKS_DONE were
reported only during index validation. Commit XXXX now enables
vacuum to report them as well, so every caller sets the field to
true. B-tree is the only access method that reads it.
Remove the field and report the counters unconditionally in
btvacuumscan(). An out-of-tree index access method that sets
report_progress needs a trivial adjustment.
Author: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Discussion: https://postgr.es/m/CALj2ACUgwSchK6jQ2CdKLBWUADTOE_zKdTff2Zg3E6hOuXKv-w@mail.gmail.com
---
src/backend/access/heap/vacuumlazy.c | 2 --
src/backend/access/nbtree/nbtree.c | 9 +++------
src/backend/catalog/index.c | 1 -
src/backend/commands/vacuumparallel.c | 1 -
src/include/access/genam.h | 1 -
5 files changed, 3 insertions(+), 11 deletions(-)
diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index a23af936b7f..44a7466def4 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -3048,7 +3048,6 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = true;
ivinfo.estimated_count = true;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = reltuples;
@@ -3114,7 +3113,6 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = true;
ivinfo.estimated_count = estimated_count;
ivinfo.message_level = DEBUG2;
diff --git a/src/backend/access/nbtree/nbtree.c b/src/backend/access/nbtree/nbtree.c
index 0abdd7b49f5..a4c3ad1b0f6 100644
--- a/src/backend/access/nbtree/nbtree.c
+++ b/src/backend/access/nbtree/nbtree.c
@@ -1339,9 +1339,7 @@ btvacuumscan(IndexVacuumInfo *info, IndexBulkDeleteResult *stats,
if (needLock)
UnlockRelationForExtension(rel, ExclusiveLock);
- if (info->report_progress)
- pgstat_progress_update_param(PROGRESS_SCAN_BLOCKS_TOTAL,
- num_pages);
+ pgstat_progress_update_param(PROGRESS_SCAN_BLOCKS_TOTAL, num_pages);
/* Quit if we've scanned the whole relation */
if (p.current_blocknum >= num_pages)
@@ -1365,9 +1363,8 @@ btvacuumscan(IndexVacuumInfo *info, IndexBulkDeleteResult *stats,
current_block = btvacuumpage(&vstate, buf);
- if (info->report_progress)
- pgstat_progress_update_param(PROGRESS_SCAN_BLOCKS_DONE,
- current_block);
+ pgstat_progress_update_param(PROGRESS_SCAN_BLOCKS_DONE,
+ current_block);
}
/*
diff --git a/src/backend/catalog/index.c b/src/backend/catalog/index.c
index ec21b83b6b8..a3768e8daf7 100644
--- a/src/backend/catalog/index.c
+++ b/src/backend/catalog/index.c
@@ -3455,7 +3455,6 @@ validate_index(Oid heapId, Oid indexId, Snapshot snapshot)
ivinfo.index = indexRelation;
ivinfo.heaprel = heapRelation;
ivinfo.analyze_only = false;
- ivinfo.report_progress = true;
ivinfo.estimated_count = true;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = heapRelation->rd_rel->reltuples;
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 607c57c8eba..fe388d99dc3 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -1098,7 +1098,6 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
ivinfo.index = indrel;
ivinfo.heaprel = pvs->heaprel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = true;
ivinfo.message_level = DEBUG2;
ivinfo.estimated_count = pvs->shared->estimated_count;
ivinfo.num_heap_tuples = pvs->shared->reltuples;
diff --git a/src/include/access/genam.h b/src/include/access/genam.h
index 72b256ecf43..07447bf1c4f 100644
--- a/src/include/access/genam.h
+++ b/src/include/access/genam.h
@@ -54,7 +54,6 @@ typedef struct IndexVacuumInfo
Relation index; /* the index being vacuumed */
Relation heaprel; /* the heap relation the index belongs to */
bool analyze_only; /* ANALYZE (without any actual vacuum) */
- bool report_progress; /* emit progress.h status reports */
bool estimated_count; /* num_heap_tuples is an estimate */
int message_level; /* ereport level for progress messages */
double num_heap_tuples; /* tuples remaining in heap */
--
2.47.3
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-20 02:37 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-20 23:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-25 16:34 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 02:23 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 05:15 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-09-03 16:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih.pg@gmail.com>
2026-09-09 19:01 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Masahiko Sawada <sawada.mshk@gmail.com>
2026-09-14 20:28 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-15 17:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih.pg@gmail.com>
2026-09-15 22:28 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
@ 2026-09-17 23:00 ` Michael Paquier <michael@paquier.xyz>
2026-09-18 01:00 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum shihao zhong <zhong950419@gmail.com>
1 sibling, 1 reply; 32+ messages in thread
From: Michael Paquier @ 2026-09-17 23:00 UTC (permalink / raw)
To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: Sami Imseih <samimseih.pg@gmail.com>; Masahiko Sawada <sawada.mshk@gmail.com>; Sami Imseih <samimseih@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
On Tue, Sep 15, 2026 at 03:28:00PM -0700, Bharath Rupireddy wrote:
> On Tue, Sep 15, 2026 at 10:11 AM Sami Imseih <samimseih.pg@gmail.com> wrote:
>> I do think it will be better to do one more split of v6-0001 to separate
>> current_index_relid and index_blks_*; they are 2 distinct features. Also
>> you can fold the 0002 into the index_blks_* commit.
>
> I prefer to keep them as one patch as I see them as closely related,
> but I'm open to splitting if anyone thinks otherwise.
In terms of the "core" patch that introduces the counters, grouping
them together is fine by me.
> I agree with all three comments above and have updated the v6 patches
> accordingly. Please have a look.
Reading through the patches. 0002 is a nice cleanup.
0003 and 0004 are not things I can get much into, echoing with the
recent following message:
https://www.postgresql.org/message-id/aqqmGGS6qqm5LI0Z%40alvherre.pgsql
I think that we should try to think harder regarding what kind of
facility we are looking for:
https://www.postgresql.org/message-id/aqqmGGS6qqm5LI0Z%40alvherre.pgsql
Your proposal based on TAP is not as heavy as the other message, but
I'm worried about the bloat this could create long-term if the same
pattern flourishes in more code paths of the tree (aka 0004 seems
AI-generated to me based on the current facilities we have, I'd
consider more building pieces before that). Switching to python may
make these test patterns much easier compared to perl, though I have
to be honest I have not looked at the recent proposals based on pypi
and such.
--
Michael
Attachments:
[application/pgp-signature] signature.asc (832B, ../../aqxxGzx6Ll_sbIMH@paquier.xyz/2-signature.asc)
download
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-20 02:37 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-20 23:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-25 16:34 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 02:23 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 05:15 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-09-03 16:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih.pg@gmail.com>
2026-09-09 19:01 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Masahiko Sawada <sawada.mshk@gmail.com>
2026-09-14 20:28 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-15 17:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih.pg@gmail.com>
2026-09-15 22:28 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-17 23:00 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
@ 2026-09-18 01:00 ` shihao zhong <zhong950419@gmail.com>
2026-09-18 18:23 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
0 siblings, 1 reply; 32+ messages in thread
From: shihao zhong @ 2026-09-18 01:00 UTC (permalink / raw)
To: Michael Paquier <michael@paquier.xyz>; +Cc: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; Sami Imseih <samimseih.pg@gmail.com>; Masahiko Sawada <sawada.mshk@gmail.com>; Sami Imseih <samimseih@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
Hi Michael,
> 0003 and 0004 are not things I can get much into, echoing with the
> recent following message:
> https://www.postgresql.org/message-id/aqqmGGS6qqm5LI0Z%40alvherre.pgsql
Thanks for the pointer. I had not seen that thread, withdraw 0003 and 0004.
I will work with Manu and Alvaro on the progress reporting test
framework in that thread. This patch would make a good first user.
Thanks,
Shihao
Attachments:
[application/octet-stream] v7-0002-Remove-IndexVacuumInfo.report_progress.patch (4.7K, ../../CAGRkXqRdus_6EfVeeuvNtSWEo1Gf5E3T3RmDsL14Ami=3CC4OA@mail.gmail.com/3-v7-0002-Remove-IndexVacuumInfo.report_progress.patch)
download | inline diff:
From 0a7a0dec24b3f1143d2c3c240b0f823d750b86a6 Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Sun, 13 Sep 2026 06:58:05 +0000
Subject: [PATCH v7 2/2] Remove IndexVacuumInfo.report_progress.
Commit ab0dfc961b6a, which added progress reporting for CREATE
INDEX, introduced report_progress so that the block counters
PROGRESS_SCAN_BLOCKS_TOTAL and PROGRESS_SCAN_BLOCKS_DONE were
reported only during index validation. Commit XXXX now enables
vacuum to report them as well, so every caller sets the field to
true. B-tree is the only access method that reads it.
Remove the field and report the counters unconditionally in
btvacuumscan(). An out-of-tree index access method that sets
report_progress needs a trivial adjustment.
Author: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Discussion: https://postgr.es/m/CALj2ACUgwSchK6jQ2CdKLBWUADTOE_zKdTff2Zg3E6hOuXKv-w@mail.gmail.com
---
src/backend/access/heap/vacuumlazy.c | 2 --
src/backend/access/nbtree/nbtree.c | 9 +++------
src/backend/catalog/index.c | 1 -
src/backend/commands/vacuumparallel.c | 1 -
src/include/access/genam.h | 1 -
5 files changed, 3 insertions(+), 11 deletions(-)
diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index a23af936b7f..44a7466def4 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -3048,7 +3048,6 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = true;
ivinfo.estimated_count = true;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = reltuples;
@@ -3114,7 +3113,6 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = true;
ivinfo.estimated_count = estimated_count;
ivinfo.message_level = DEBUG2;
diff --git a/src/backend/access/nbtree/nbtree.c b/src/backend/access/nbtree/nbtree.c
index 0abdd7b49f5..a4c3ad1b0f6 100644
--- a/src/backend/access/nbtree/nbtree.c
+++ b/src/backend/access/nbtree/nbtree.c
@@ -1339,9 +1339,7 @@ btvacuumscan(IndexVacuumInfo *info, IndexBulkDeleteResult *stats,
if (needLock)
UnlockRelationForExtension(rel, ExclusiveLock);
- if (info->report_progress)
- pgstat_progress_update_param(PROGRESS_SCAN_BLOCKS_TOTAL,
- num_pages);
+ pgstat_progress_update_param(PROGRESS_SCAN_BLOCKS_TOTAL, num_pages);
/* Quit if we've scanned the whole relation */
if (p.current_blocknum >= num_pages)
@@ -1365,9 +1363,8 @@ btvacuumscan(IndexVacuumInfo *info, IndexBulkDeleteResult *stats,
current_block = btvacuumpage(&vstate, buf);
- if (info->report_progress)
- pgstat_progress_update_param(PROGRESS_SCAN_BLOCKS_DONE,
- current_block);
+ pgstat_progress_update_param(PROGRESS_SCAN_BLOCKS_DONE,
+ current_block);
}
/*
diff --git a/src/backend/catalog/index.c b/src/backend/catalog/index.c
index ec21b83b6b8..a3768e8daf7 100644
--- a/src/backend/catalog/index.c
+++ b/src/backend/catalog/index.c
@@ -3455,7 +3455,6 @@ validate_index(Oid heapId, Oid indexId, Snapshot snapshot)
ivinfo.index = indexRelation;
ivinfo.heaprel = heapRelation;
ivinfo.analyze_only = false;
- ivinfo.report_progress = true;
ivinfo.estimated_count = true;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = heapRelation->rd_rel->reltuples;
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 607c57c8eba..fe388d99dc3 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -1098,7 +1098,6 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
ivinfo.index = indrel;
ivinfo.heaprel = pvs->heaprel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = true;
ivinfo.message_level = DEBUG2;
ivinfo.estimated_count = pvs->shared->estimated_count;
ivinfo.num_heap_tuples = pvs->shared->reltuples;
diff --git a/src/include/access/genam.h b/src/include/access/genam.h
index 72b256ecf43..07447bf1c4f 100644
--- a/src/include/access/genam.h
+++ b/src/include/access/genam.h
@@ -54,7 +54,6 @@ typedef struct IndexVacuumInfo
Relation index; /* the index being vacuumed */
Relation heaprel; /* the heap relation the index belongs to */
bool analyze_only; /* ANALYZE (without any actual vacuum) */
- bool report_progress; /* emit progress.h status reports */
bool estimated_count; /* num_heap_tuples is an estimate */
int message_level; /* ereport level for progress messages */
double num_heap_tuples; /* tuples remaining in heap */
--
2.47.3
[application/octet-stream] v7-0001-Report-per-index-vacuum-progress-in-pg_stat_progr.patch (17.8K, ../../CAGRkXqRdus_6EfVeeuvNtSWEo1Gf5E3T3RmDsL14Ami=3CC4OA@mail.gmail.com/4-v7-0001-Report-per-index-vacuum-progress-in-pg_stat_progr.patch)
download | inline diff:
From 4acb49d16e43f7f19b7c2cdd38be6001d43debf4 Mon Sep 17 00:00:00 2001
From: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Date: Sun, 13 Sep 2026 05:19:26 +0000
Subject: [PATCH v7 1/2] Report per-index vacuum progress in
pg_stat_progress_vacuum.
Previously, pg_stat_progress_vacuum showed the total and
processed index counts, but neither which index a backend was
working on nor how far it had gotten through it. On a table with
many indexes of different access methods, that made it hard to
tell which index a slow or stuck vacuum was spending its time on,
and for a large B-tree, at the scale of hundreds of GBs to TBs,
there was no way to estimate when the index phase, and together
with heap_blks_*, the whole vacuum would finish.
This commit adds three columns. current_index_relid reports the
OID of the index a backend is currently vacuuming or cleaning up.
index_blks_total and index_blks_done report block progress within
that index. All three are set before an index is processed and
reset once that index is done, so they do not report a stale
value after the phase moves on.
During parallel index vacuum the leader also participates, and
each participant processes a different set of indexes. That
per-index progress cannot be collapsed into a single leader row
without losing the detail that matters, so each participant, the
leader and the workers alike, now reports its own row and shows
the index it is processing and how far along it is. A worker row
carries the columns that come from its own backend state, pid,
datid, datname, relid, phase, current_index_relid,
index_blks_total and index_blks_done; the remaining columns track
command-level progress that the leader maintains and read as zero
on worker rows, or NULL for mode and started_by. Because a table
can be vacuumed by only one VACUUM at a time, the rows sharing a
datid and relid make up a single vacuum, so they can be grouped
that way to see the leader together with all of its workers. If
the exact leader-to-worker mapping is wanted, it can be recovered
from leader_pid in pg_stat_activity.
The block counters are the ones CREATE INDEX progress reporting
added in commit ab0dfc961b6a, PROGRESS_SCAN_BLOCKS_TOTAL and
PROGRESS_SCAN_BLOCKS_DONE. B-tree's index scan already knows how
to report them, so all that is needed here is turning that
reporting on in the serial and parallel index vacuum paths; other
access methods report nothing and leave the two columns at zero.
They are deliberately kept separate from heap_blks_*, which must
be retained across a multi-pass index vacuum, the case where the
dead-TID store fills, and which in the serial case belong to the
same backend that is doing the index vacuuming, so reusing them
for index blocks would destroy heap progress that is still needed.
No caller sets IndexVacuumInfo.report_progress to false anymore;
the next commit removes the field.
XXX: Bump catalog version.
Author: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
Reviewed-by: Michael Paquier <michael@paquier.xyz>
Reviewed-by: Sami Imseih <samimseih@gmail.com>
Reviewed-by: Masahiko Sawada <sawada.mshk@gmail.com>
Discussion: https://postgr.es/m/CALj2ACUgwSchK6jQ2CdKLBWUADTOE_zKdTff2Zg3E6hOuXKv-w@mail.gmail.com
---
doc/src/sgml/monitoring.sgml | 95 ++++++++++++++++++++++++++-
src/backend/access/heap/vacuumlazy.c | 36 +++++++++-
src/backend/catalog/system_views.sql | 5 +-
src/backend/commands/vacuumparallel.c | 37 ++++++++++-
src/include/commands/progress.h | 2 +
src/test/regress/expected/rules.out | 5 +-
6 files changed, 172 insertions(+), 8 deletions(-)
diff --git a/doc/src/sgml/monitoring.sgml b/doc/src/sgml/monitoring.sgml
index 62dadf3e86c..4c49aac7003 100644
--- a/doc/src/sgml/monitoring.sgml
+++ b/doc/src/sgml/monitoring.sgml
@@ -7670,8 +7670,11 @@ FROM pg_stat_get_backend_idset() AS backendid;
<para>
Whenever <command>VACUUM</command> is running, the
<structname>pg_stat_progress_vacuum</structname> view will contain
- one row for each backend (including autovacuum worker processes) that is
- currently vacuuming. The tables below describe the information
+ one row for each backend (including autovacuum worker processes and
+ parallel workers launched for
+ <link linkend="sql-vacuum">parallel vacuum</link>) that is currently
+ vacuuming; see the note following the view for the columns reported on
+ parallel worker rows. The tables below describe the information
that will be reported and provide information about how to interpret it.
Progress for <command>VACUUM FULL</command> commands is reported via
<structname>pg_stat_progress_repack</structname>, and is also visible via
@@ -7929,10 +7932,98 @@ FROM pg_stat_get_backend_idset() AS backendid;
</itemizedlist>
</para></entry>
</row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>current_index_relid</structfield> <type>oid</type>
+ </para>
+ <para>
+ If <command>VACUUM</command> is currently processing an index, this
+ column shows the OID of the index being vacuumed. The value is set
+ when the phase is <literal>vacuuming indexes</literal> or
+ <literal>cleaning up indexes</literal>, and is reset to 0 once that
+ index has been processed, so it does not show a stale index while the
+ backend is between indexes or has moved on to another phase. During
+ parallel index vacuum, each parallel worker row shows the index that
+ particular worker is processing.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>index_blks_total</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Total number of blocks in the index identified by
+ <structfield>current_index_relid</structfield>. This is reported only
+ for B-tree indexes, and only while the index is being scanned; it is 0
+ for other index access methods, for a B-tree whose scan
+ <command>VACUUM</command> was able to skip during cleanup, and once
+ the index has been processed.
+ </para></entry>
+ </row>
+
+ <row>
+ <entry role="catalog_table_entry"><para role="column_definition">
+ <structfield>index_blks_done</structfield> <type>bigint</type>
+ </para>
+ <para>
+ Number of blocks of the index identified by
+ <structfield>current_index_relid</structfield> scanned so far. This is
+ reported only for B-tree indexes, and only while the index is being
+ scanned; it is 0 for other index access methods, for a B-tree whose
+ scan <command>VACUUM</command> was able to skip during cleanup, and
+ once the index has been processed. Together with
+ <structfield>index_blks_total</structfield> this gives per-index vacuum
+ progress, which is useful for estimating completion of large indexes.
+ The index metapage is counted in
+ <structfield>index_blks_total</structfield> but is never scanned, so
+ this column stops one block short of it.
+ </para></entry>
+ </row>
+
</tbody>
</tgroup>
</table>
+ <note>
+ <para>
+ During a parallel vacuum, each participating backend (the leader and each
+ parallel worker) reports its own row. On a worker row the meaningful
+ columns are <structfield>pid</structfield>, <structfield>datid</structfield>,
+ <structfield>datname</structfield>, <structfield>relid</structfield>,
+ <structfield>phase</structfield>, <structfield>current_index_relid</structfield>,
+ <structfield>index_blks_total</structfield> and
+ <structfield>index_blks_done</structfield>. The remaining columns track
+ command-level heap progress that only the leader maintains; they read as
+ zero on worker rows, except that <structfield>mode</structfield> and
+ <structfield>started_by</structfield> appear as <literal>NULL</literal>.
+ A participant row showing
+ <literal>vacuuming indexes</literal> or
+ <literal>cleaning up indexes</literal> with a zero
+ <structfield>current_index_relid</structfield> is between indexes, or has
+ finished its share of the indexes while other participants are still
+ working. A worker row whose <structfield>phase</structfield> is
+ <literal>initializing</literal> is a worker that has been launched but has
+ not started on an index yet, which also happens when every index was
+ claimed by another participant before this worker got to it. Because a
+ table can be vacuumed by only one
+ <command>VACUUM</command> at a time, the rows sharing a given
+ <structfield>datid</structfield> and <structfield>relid</structfield>
+ together make up a single vacuum, so they can be grouped that way to see
+ the leader and all of its workers. The leader and worker process IDs can be
+ correlated with <structfield>leader_pid</structfield> in
+ <link linkend="monitoring-pg-stat-activity-view"><structname>pg_stat_activity</structname></link>
+ if needed. Because the worker rows are separate backends, a role that does
+ not have privileges of the <literal>pg_read_all_stats</literal> role and
+ does not own those backends sees only their <structfield>pid</structfield>,
+ <structfield>datid</structfield> and <structfield>datname</structfield>,
+ with the remaining columns <literal>NULL</literal>, so such a role can tell
+ that a parallel vacuum is running in a database but not which table or
+ index it is working on.
+ </para>
+ </note>
+
<table id="vacuum-phases">
<title>VACUUM Phases</title>
<tgroup cols="2">
diff --git a/src/backend/access/heap/vacuumlazy.c b/src/backend/access/heap/vacuumlazy.c
index 8e1f660bc2f..a23af936b7f 100644
--- a/src/backend/access/heap/vacuumlazy.c
+++ b/src/backend/access/heap/vacuumlazy.c
@@ -3038,16 +3038,26 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
{
IndexVacuumInfo ivinfo;
LVSavedErrInfo saved_err_info;
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.estimated_count = true;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3069,6 +3079,12 @@ lazy_vacuum_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
+
return istat;
}
@@ -3088,17 +3104,27 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
{
IndexVacuumInfo ivinfo;
LVSavedErrInfo saved_err_info;
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
ivinfo.index = indrel;
ivinfo.heaprel = vacrel->rel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.estimated_count = estimated_count;
ivinfo.message_level = DEBUG2;
ivinfo.num_heap_tuples = reltuples;
ivinfo.strategy = vacrel->bstrategy;
+ /* Report which index we're currently processing */
+ pgstat_progress_update_param(PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ (int64) RelationGetRelid(indrel));
+
/*
* Update error traceback information.
*
@@ -3118,6 +3144,12 @@ lazy_cleanup_one_index(Relation indrel, IndexBulkDeleteResult *istat,
pfree(vacrel->indname);
vacrel->indname = NULL;
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
+
return istat;
}
diff --git a/src/backend/catalog/system_views.sql b/src/backend/catalog/system_views.sql
index ad340887f54..809b9c0f1e4 100644
--- a/src/backend/catalog/system_views.sql
+++ b/src/backend/catalog/system_views.sql
@@ -1353,7 +1353,10 @@ CREATE VIEW pg_stat_progress_vacuum AS
CASE S.param13 WHEN 1 THEN 'manual'
WHEN 2 THEN 'autovacuum'
WHEN 3 THEN 'autovacuum_wraparound'
- ELSE NULL END AS started_by
+ ELSE NULL END AS started_by,
+ CAST(S.param14 AS oid) AS current_index_relid,
+ S.param16 AS index_blks_total,
+ S.param17 AS index_blks_done
FROM pg_stat_get_progress_info('VACUUM') AS S
LEFT JOIN pg_database D ON S.datid = D.oid;
diff --git a/src/backend/commands/vacuumparallel.c b/src/backend/commands/vacuumparallel.c
index 767d162e578..607c57c8eba 100644
--- a/src/backend/commands/vacuumparallel.c
+++ b/src/backend/commands/vacuumparallel.c
@@ -1076,6 +1076,17 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
IndexBulkDeleteResult *istat = NULL;
IndexBulkDeleteResult *istat_res;
IndexVacuumInfo ivinfo;
+ const int progress_index[] = {
+ PROGRESS_VACUUM_PHASE,
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID
+ };
+ int64 progress_val[2];
+ const int reset_index[] = {
+ PROGRESS_VACUUM_CURRENT_INDEX_RELID,
+ PROGRESS_SCAN_BLOCKS_TOTAL,
+ PROGRESS_SCAN_BLOCKS_DONE
+ };
+ const int64 reset_val[] = {(int64) InvalidOid, 0, 0};
/*
* Update the pointer to the corresponding bulk-deletion result if someone
@@ -1087,7 +1098,7 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
ivinfo.index = indrel;
ivinfo.heaprel = pvs->heaprel;
ivinfo.analyze_only = false;
- ivinfo.report_progress = false;
+ ivinfo.report_progress = true;
ivinfo.message_level = DEBUG2;
ivinfo.estimated_count = pvs->shared->estimated_count;
ivinfo.num_heap_tuples = pvs->shared->reltuples;
@@ -1097,6 +1108,16 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
pvs->indname = pstrdup(RelationGetRelationName(indrel));
pvs->status = indstats->status;
+ /*
+ * Report the phase and the index we're about to process before we start,
+ * so that it is visible for the whole duration of the index scan.
+ */
+ progress_val[0] = (indstats->status == PARALLEL_INDVAC_STATUS_NEED_BULKDELETE)
+ ? PROGRESS_VACUUM_PHASE_VACUUM_INDEX
+ : PROGRESS_VACUUM_PHASE_INDEX_CLEANUP;
+ progress_val[1] = (int64) RelationGetRelid(indrel);
+ pgstat_progress_update_multi_param(2, progress_index, progress_val);
+
switch (indstats->status)
{
case PARALLEL_INDVAC_STATUS_NEED_BULKDELETE:
@@ -1112,6 +1133,12 @@ parallel_vacuum_process_one_index(ParallelVacuumState *pvs, Relation indrel,
RelationGetRelationName(indrel));
}
+ /*
+ * Reset the current index progress parameters to avoid reporting stale
+ * values.
+ */
+ pgstat_progress_update_multi_param(3, reset_index, reset_val);
+
/*
* Copy the index bulk-deletion result returned from ambulkdelete and
* amvacuumcleanup to the DSM segment if it's the first cycle because they
@@ -1191,7 +1218,7 @@ parallel_vacuum_index_is_parallel_safe(Relation indrel, int num_index_scans,
* Perform work within a launched parallel process.
*
* Since parallel vacuum workers perform only index vacuum or index cleanup,
- * we don't need to report progress information.
+ * they report progress for the index they are processing.
*/
void
parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
@@ -1315,6 +1342,9 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Prepare to track buffer usage during parallel execution */
InstrStartParallelQuery();
+ /* Register this worker for vacuum progress reporting */
+ pgstat_progress_start_command(PROGRESS_COMMAND_VACUUM, shared->relid);
+
/* Process indexes to perform vacuum/cleanup */
parallel_vacuum_process_safe_indexes(&pvs);
@@ -1334,6 +1364,9 @@ parallel_vacuum_main(dsm_segment *seg, shm_toc *toc)
/* Pop the error context stack */
error_context_stack = errcallback.previous;
+ /* Unregister this worker from vacuum progress reporting */
+ pgstat_progress_end_command();
+
vac_close_indexes(nindexes, indrels, RowExclusiveLock);
table_close(rel, ShareUpdateExclusiveLock);
FreeAccessStrategy(pvs.bstrategy);
diff --git a/src/include/commands/progress.h b/src/include/commands/progress.h
index 2a12920c75f..bf91455eff9 100644
--- a/src/include/commands/progress.h
+++ b/src/include/commands/progress.h
@@ -31,6 +31,8 @@
#define PROGRESS_VACUUM_DELAY_TIME 10
#define PROGRESS_VACUUM_MODE 11
#define PROGRESS_VACUUM_STARTED_BY 12
+#define PROGRESS_VACUUM_CURRENT_INDEX_RELID 13
+/* 15 and 16 reserved for "block number" metrics */
/* Phases of vacuum (as advertised via PROGRESS_VACUUM_PHASE) */
#define PROGRESS_VACUUM_PHASE_SCAN_HEAP 1
diff --git a/src/test/regress/expected/rules.out b/src/test/regress/expected/rules.out
index 4a8cc759d7b..0addd043e68 100644
--- a/src/test/regress/expected/rules.out
+++ b/src/test/regress/expected/rules.out
@@ -2214,7 +2214,10 @@ pg_stat_progress_vacuum| SELECT s.pid,
WHEN 2 THEN 'autovacuum'::text
WHEN 3 THEN 'autovacuum_wraparound'::text
ELSE NULL::text
- END AS started_by
+ END AS started_by,
+ (s.param14)::oid AS current_index_relid,
+ s.param16 AS index_blks_total,
+ s.param17 AS index_blks_done
FROM (pg_stat_get_progress_info('VACUUM'::text) s(pid, datid, relid, param1, param2, param3, param4, param5, param6, param7, param8, param9, param10, param11, param12, param13, param14, param15, param16, param17, param18, param19, param20)
LEFT JOIN pg_database d ON ((s.datid = d.oid)));
pg_stat_recovery| SELECT promote_triggered,
--
2.47.3
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-20 02:37 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-20 23:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-25 16:34 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 02:23 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 05:15 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-09-03 16:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih.pg@gmail.com>
2026-09-09 19:01 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Masahiko Sawada <sawada.mshk@gmail.com>
2026-09-14 20:28 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-15 17:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih.pg@gmail.com>
2026-09-15 22:28 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-17 23:00 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-09-18 01:00 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum shihao zhong <zhong950419@gmail.com>
@ 2026-09-18 18:23 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
0 siblings, 0 replies; 32+ messages in thread
From: Bharath Rupireddy @ 2026-09-18 18:23 UTC (permalink / raw)
To: shihao zhong <zhong950419@gmail.com>; +Cc: Michael Paquier <michael@paquier.xyz>; Sami Imseih <samimseih.pg@gmail.com>; Masahiko Sawada <sawada.mshk@gmail.com>; Sami Imseih <samimseih@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
Hi,
On Thu, Sep 17, 2026 at 6:00 PM shihao zhong <zhong950419@gmail.com> wrote:
>
> v7 looks good here. Tested sequential, parallel on btree and GIN, all as the
> docs describe, including GIN reporting zero blocks and the reset
> clearing the previous index. Applies cleanly on 5e595d659fd, suites
> pass.
Thanks for testing v7 on these and for confirming the behavior matches the docs.
> > 0003 and 0004 are not things I can get much into, echoing with the
> > recent following message:
> > https://www.postgresql.org/message-id/aqqmGGS6qqm5LI0Z%40alvherre.pgsql
>
> Thanks for the pointer. I had not seen that thread, withdraw 0003 and 0004.
>
> I will work with Manu and Alvaro on the progress reporting test
> framework in that thread. This patch would make a good first user.
+1 to discussing a general approach for progress reporting tests
separately. We haven't had progress reporting tests so far, but that
doesn't mean we must not have them. One idea is to look at whether
there are any bugs or inconsistencies reported for any of the progress
reporting commands and start from there (but that's a discussion for
the other thread).
--
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
@ 2026-08-19 18:10 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-19 18:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
1 sibling, 1 reply; 32+ messages in thread
From: Bharath Rupireddy @ 2026-08-19 18:10 UTC (permalink / raw)
To: Michael Paquier <michael@paquier.xyz>; +Cc: Sami Imseih <samimseih@gmail.com>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
Hi,
On Thu, Aug 13, 2026 at 8:48 PM Michael Paquier <michael@paquier.xyz> wrote:
>
> On Tue, Aug 04, 2026 at 03:20:00PM -0700, Bharath Rupireddy wrote:
> > I spent some more time thinking about using arrays here, and about the
> > one-row-per-command policy. I still think emitting the index OIDs and
> > worker PIDs as position-aligned arrays (like the existing
> > pg_stats.most_common_vals/most_common_freqs columns) is the simple
> > solution. I appreciate any thoughts or other ways here.
>
> Exposing the information of an index a worker is processing is a good
> idea, but I think that this choice lacks a long-term vision. I think
> that we should expose one row for each worker rather than an array of
> PIDs and index OIDs in the row of a leader.
> The main issue for me is
> the granularity of the information provided, where it would actually
> make sense to provide more information for each worker.
The position-aligned array will only grow in the future, making it
hard to add more per-worker information later (index blocks total vs
done for parallel index vacuum, heap blocks total vs done for parallel
heap vacuum, per index dead rows cleaned up and deduped, etc.).
> Choosing how
> an index clean works in vacuum for parallel workers is an
> implementation choice, where we could think about approaches like:
> - Distribute the workload of one index across N workers (for a 1TB
> index, spawn N workers each sharing 1/N TB of data to clean)
> - Have each worker do one index.
> - Or more strategies, etc.
> My point is not the strategy or the design we choose, which could vary
> depending on an index AM. It's that for any design, any strategy or
> any index AM, at the end it is going to be way more important for the
> end-user how *each* individual worker behaves.
I found ab0dfc961 which adds AM-agnostic and AM-specific fields for
the indexes supported in core (see below). When multiple indexes want
a common thing to be reported, we can add such a thing to the core and
let the specific index AM report that information.
> One thing could be for
> example reusing heap_blks_total and heap_blks_scanned for indexes, so
> as it is possible how much each worker has done (let's perhaps rename
> them).
heap_blks_* cannot be reused for tracking index blocks total and
scanned, because those values must be retained across multi-pass index
vacuuming (which triggers when the dead-TID store fills). Although
this is rare after the radix-tree based TID store optimizations, it's
still possible. The leader itself does index vacuuming in the serial
case, so overwriting those fields would corrupt the heap progress
that's still needed.
I came across two AM-specific fields (used for BTree and GIN for now):
PROGRESS_SCAN_BLOCKS_TOTAL/DONE, added by ab0dfc961 for create index
progress reporting. For example, btvacuumscan reports index progress
during concurrent BTree index creation/recreation cases, something
like the following:
phase | blocks_done | blocks_total
----------------------------------+-------------+--------------
index validation: scanning index | 959 | 27422
index validation: scanning index | 13350 | 27422
index validation: scanning index | 25973 | 27422
This, combined with per-worker vacuum progress reporting, lets us
report index vacuum progress nicely for BTree. Although this only
covers BTree for now (others will continue to report as NULL), it's a
good starting point since the majority of indexes are BTree. It also
gives visibility into how the index vacuum is progressing towards its
goal and lets one estimate the vacuum finish time (along with
heap_blks_*), particularly with hundreds of GBs and TBs of indexes at
scale.
> Being able to map a leader with its worker is an information
> already provided by pg_stat_activity, adding this information in the
> progress view seems unnecessary for me to add here as a JOIN is
> already able to solve that anyway. If extra SQL knowledge is
> necessary, that's more a documentation problem to me, adding more
> fields for data that's already available is just more information
> bloat.
Although we could get leader_pid almost for free in the parallel index
vacuum cases, I agree that having it there is not only information
bloat but also eats up a fixed progress reporting slot in shared
memory (we only have 20, and I expect that to grow in the future). We
can leave a note in the docs that leader_pid being NULL in
pg_stat_activity, especially with roles not having pg_read_all_stats
or roles not owning the backends, means they won't see the worker
rows.
In short, I tend to agree with having one row per worker in the vacuum
progress report, joining pg_stat_activity's leader_pid for simpler
usability, extensibility, and less information bloat, along with doc
changes to explain this. One concern is that some progress fields
would be null on worker rows, but documenting this should be
sufficient. I could be missing something here, so I would like to hear
some thoughts before coming up with a patch.
--
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com
^ permalink raw reply [nested|flat] 32+ messages in thread
* Re: Report index currently being vacuumed in pg_stat_progress_vacuum
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-05 17:54 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 03:48 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Michael Paquier <michael@paquier.xyz>
2026-08-19 18:10 ` Re: Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
@ 2026-08-19 18:31 ` Sami Imseih <samimseih@gmail.com>
0 siblings, 0 replies; 32+ messages in thread
From: Sami Imseih @ 2026-08-19 18:31 UTC (permalink / raw)
To: Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>; +Cc: Michael Paquier <michael@paquier.xyz>; PostgreSQL Hackers <pgsql-hackers@lists.postgresql.org>; SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
> In short, I tend to agree with having one row per worker in the vacuum
> progress report, joining pg_stat_activity's leader_pid for simpler
> usability, extensibility, and less information bloat, along with doc
> changes to explain this. One concern is that some progress fields
> would be null on worker rows, but documenting this should be
> sufficient. I could be missing something here, so I would like to hear
> some thoughts before coming up with a patch.
>
I will look at the rest of the points later in detail, but it does sound
like to me that a new worker view will be a better place to hold extra per
worker ( or leader ) details and the current view will remain high
level/progress data. Unused progress fields do not sound right to me, and I
also worry we will bloat the existing view over time if we want to add more
per worker fields.
--
Sami
^ permalink raw reply [nested|flat] 32+ messages in thread
end of thread, other threads:[~2026-09-18 18:23 UTC | newest]
Thread overview: 32+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2026-05-04 02:00 Report index currently being vacuumed in pg_stat_progress_vacuum Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-05-04 04:53 ` SATYANARAYANA NARLAPURAM <satyanarlapuram@gmail.com>
2026-05-04 09:52 ` Antonin Houska <ah@cybertec.at>
2026-05-05 17:54 ` Sami Imseih <samimseih@gmail.com>
2026-06-29 15:31 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-07-16 20:49 ` Sami Imseih <samimseih@gmail.com>
2026-07-17 19:24 ` Sami Imseih <samimseih@gmail.com>
2026-08-04 22:20 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-12 21:45 ` Sami Imseih <samimseih@gmail.com>
2026-08-13 23:30 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-14 00:47 ` Sami Imseih <samimseih@gmail.com>
2026-08-14 03:33 ` Michael Paquier <michael@paquier.xyz>
2026-08-14 03:48 ` Michael Paquier <michael@paquier.xyz>
2026-08-14 21:26 ` Sami Imseih <samimseih@gmail.com>
2026-08-14 23:35 ` Sami Imseih <samimseih@gmail.com>
2026-08-20 02:37 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-20 23:10 ` Michael Paquier <michael@paquier.xyz>
2026-08-20 23:58 ` Sami Imseih <samimseih@gmail.com>
2026-08-25 16:34 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 02:23 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-03 05:15 ` Michael Paquier <michael@paquier.xyz>
2026-09-03 16:49 ` Sami Imseih <samimseih.pg@gmail.com>
2026-09-09 19:01 ` Masahiko Sawada <sawada.mshk@gmail.com>
2026-09-14 20:28 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-15 17:10 ` Sami Imseih <samimseih.pg@gmail.com>
2026-09-15 22:28 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-09-17 02:23 ` shihao zhong <zhong950419@gmail.com>
2026-09-17 23:00 ` Michael Paquier <michael@paquier.xyz>
2026-09-18 01:00 ` shihao zhong <zhong950419@gmail.com>
2026-09-18 18:23 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-19 18:10 ` Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>
2026-08-19 18:31 ` Sami Imseih <samimseih@gmail.com>
This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox