agora inbox for [email protected]help / color / mirror / Atom feed
[PATCH v4] Avoid orphaned objects dependencies 246+ messages / 2 participants [nested] [flat]
* [PATCH v4] Avoid orphaned objects dependencies @ 2024-03-29 15:43 Bertrand Drouvot <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Bertrand Drouvot @ 2024-03-29 15:43 UTC (permalink / raw) It's currently possible to create orphaned objects dependencies, for example: Scenario 1: session 1: begin; drop schema schem; session 2: create a function in the schema schem session 1: commit; With the above, the function created in session 2 would be linked to a non existing schema. Scenario 2: session 1: begin; create a function in the schema schem session 2: drop schema schem; session 1: commit; With the above, the function created in session 1 would be linked to a non existing schema. To avoid those scenarios, a new lock (that conflicts with a lock taken by DROP) has been put in place when the dependencies are being recorded. With this in place, the drop schema in scenario 2 would be locked. Also, after the new lock attempt, the patch checks that the object still exists: with this in place session 2 in scenario 1 would be locked and would report an error once session 1 committs (that would not be the case should session 1 abort the transaction). If the object is dropped before the new lock attempt is triggered then the patch would also report an error (but with less details). The patch adds a few tests for some dependency cases (that would currently produce orphaned objects): - schema and function (as the above scenarios) - function and arg type - function and return type - function and function - domain and domain - table and type - server and foreign data wrapper --- src/backend/catalog/dependency.c | 54 +++++++++ src/backend/catalog/objectaddress.c | 70 +++++++++++ src/backend/catalog/pg_depend.c | 6 + src/include/catalog/dependency.h | 1 + src/include/catalog/objectaddress.h | 1 + src/test/modules/Makefile | 1 + src/test/modules/meson.build | 1 + .../test_dependencies_locks/.gitignore | 3 + .../modules/test_dependencies_locks/Makefile | 14 +++ .../expected/test_dependencies_locks.out | 113 ++++++++++++++++++ .../test_dependencies_locks/meson.build | 12 ++ .../specs/test_dependencies_locks.spec | 78 ++++++++++++ 12 files changed, 354 insertions(+) 24.6% src/backend/catalog/ 42.7% src/test/modules/test_dependencies_locks/expected/ 25.9% src/test/modules/test_dependencies_locks/specs/ 5.0% src/test/modules/test_dependencies_locks/ diff --git a/src/backend/catalog/dependency.c b/src/backend/catalog/dependency.c index d4b5b2ade1..a49357bbe2 100644 --- a/src/backend/catalog/dependency.c +++ b/src/backend/catalog/dependency.c @@ -1517,6 +1517,60 @@ AcquireDeletionLock(const ObjectAddress *object, int flags) } } +/* + * depLockAndCheckObject + * + * Lock the object that we are about to record a dependency on. + * After it's locked, verify that it hasn't been dropped while we + * weren't looking. If the object has been dropped, this function + * does not return! + */ +void +depLockAndCheckObject(const ObjectAddress *object) +{ + char *object_description; + + /* + * Those don't rely on LockDatabaseObject() when being dropped (see + * AcquireDeletionLock()). Also it looks like they can not produce + * orphaned dependent objects when being dropped. + */ + if (object->classId == RelationRelationId || object->classId == AuthMemRelationId) + return; + + object_description = getObjectDescription(object, true); + + /* assume we should lock the whole object not a sub-object */ + LockDatabaseObject(object->classId, object->objectId, 0, AccessShareLock); + + /* check if object still exists */ + if (!ObjectByIdExist(object, false)) + { + /* + * It might be possible that we are creating it (for example creating + * a composite type while creating a relation), so bypass the syscache + * lookup and use a dirty snaphot instead to cover this scenario. + */ + if (!ObjectByIdExist(object, true)) + { + /* + * If the object has been dropped before we get a chance to get + * its description, then emit a generic error message. That looks + * like a good compromise over extra complexity. + */ + if (object_description) + ereport(ERROR, errmsg("%s does not exist", object_description)); + else + ereport(ERROR, errmsg("a dependent object does not exist")); + } + } + + if (object_description) + pfree(object_description); + + return; +} + /* * ReleaseDeletionLock - release an object deletion lock * diff --git a/src/backend/catalog/objectaddress.c b/src/backend/catalog/objectaddress.c index 7b536ac6fd..c2b873dd81 100644 --- a/src/backend/catalog/objectaddress.c +++ b/src/backend/catalog/objectaddress.c @@ -2590,6 +2590,76 @@ get_object_namespace(const ObjectAddress *address) return oid; } +/* + * ObjectByIdExist + * + * Return whether the given object exists. + * + * Works for most catalogs, if no special processing is needed. + */ +bool +ObjectByIdExist(const ObjectAddress *address, bool use_dirty_snapshot) +{ + HeapTuple tuple; + int cache = -1; + const ObjectPropertyType *property; + + if (!use_dirty_snapshot) + { + property = get_object_property_data(address->classId); + + cache = property->oid_catcache_id; + } + + if (cache >= 0) + { + /* Fetch tuple from syscache. */ + tuple = SearchSysCache1(cache, ObjectIdGetDatum(address->objectId)); + + if (!HeapTupleIsValid(tuple)) + { + return false; + } + + ReleaseSysCache(tuple); + + return true; + } + else + { + Relation rel; + ScanKeyData skey[1]; + SysScanDesc scan; + SnapshotData DirtySnapshot; + Snapshot snapshot; + + if (use_dirty_snapshot) + { + InitDirtySnapshot(DirtySnapshot); + snapshot = &DirtySnapshot; + } + else + snapshot = NULL; + + rel = table_open(address->classId, AccessShareLock); + + ScanKeyInit(&skey[0], + get_object_attnum_oid(address->classId), + BTEqualStrategyNumber, F_OIDEQ, + ObjectIdGetDatum(address->objectId)); + + scan = systable_beginscan(rel, get_object_oid_index(address->classId), true, + snapshot, 1, skey); + + /* we expect exactly one match */ + tuple = systable_getnext(scan); + systable_endscan(scan); + table_close(rel, AccessShareLock); + + return (HeapTupleIsValid(tuple)); + } +} + /* * Return ObjectType for the given object type as given by * getObjectTypeDescription; if no valid ObjectType code exists, but it's a diff --git a/src/backend/catalog/pg_depend.c b/src/backend/catalog/pg_depend.c index f85a898de8..6bb218f5dd 100644 --- a/src/backend/catalog/pg_depend.c +++ b/src/backend/catalog/pg_depend.c @@ -106,6 +106,12 @@ recordMultipleDependencies(const ObjectAddress *depender, if (isObjectPinned(referenced)) continue; + /* + * Acquire a lock and check object still exists while recording the + * dependency. XXX - Should we do so only for DEPENDENCY_NORMAL? + */ + depLockAndCheckObject(referenced); + if (slot_init_count < max_slots) { slot[slot_stored_count] = MakeSingleTupleTableSlot(RelationGetDescr(dependDesc), diff --git a/src/include/catalog/dependency.h b/src/include/catalog/dependency.h index ec654010d4..8915548711 100644 --- a/src/include/catalog/dependency.h +++ b/src/include/catalog/dependency.h @@ -94,6 +94,7 @@ typedef struct ObjectAddresses ObjectAddresses; /* in dependency.c */ extern void AcquireDeletionLock(const ObjectAddress *object, int flags); +extern void depLockAndCheckObject(const ObjectAddress *object); extern void ReleaseDeletionLock(const ObjectAddress *object); diff --git a/src/include/catalog/objectaddress.h b/src/include/catalog/objectaddress.h index 3a70d80e32..04891abcc1 100644 --- a/src/include/catalog/objectaddress.h +++ b/src/include/catalog/objectaddress.h @@ -53,6 +53,7 @@ extern void check_object_ownership(Oid roleid, Node *object, Relation relation); extern Oid get_object_namespace(const ObjectAddress *address); +extern bool ObjectByIdExist(const ObjectAddress *address, bool use_dirty_snapshot); extern bool is_objectclass_supported(Oid class_id); extern const char *get_object_class_descr(Oid class_id); diff --git a/src/test/modules/Makefile b/src/test/modules/Makefile index 256799f520..75f357100f 100644 --- a/src/test/modules/Makefile +++ b/src/test/modules/Makefile @@ -17,6 +17,7 @@ SUBDIRS = \ test_copy_callbacks \ test_custom_rmgrs \ test_ddl_deparse \ + test_dependencies_locks \ test_dsa \ test_dsm_registry \ test_extensions \ diff --git a/src/test/modules/meson.build b/src/test/modules/meson.build index d8fe059d23..60305dcccd 100644 --- a/src/test/modules/meson.build +++ b/src/test/modules/meson.build @@ -16,6 +16,7 @@ subdir('test_bloomfilter') subdir('test_copy_callbacks') subdir('test_custom_rmgrs') subdir('test_ddl_deparse') +subdir('test_dependencies_locks') subdir('test_dsa') subdir('test_dsm_registry') subdir('test_extensions') diff --git a/src/test/modules/test_dependencies_locks/.gitignore b/src/test/modules/test_dependencies_locks/.gitignore new file mode 100644 index 0000000000..bf000faac4 --- /dev/null +++ b/src/test/modules/test_dependencies_locks/.gitignore @@ -0,0 +1,3 @@ +# Generated subdirectories +/log/ +/output_iso diff --git a/src/test/modules/test_dependencies_locks/Makefile b/src/test/modules/test_dependencies_locks/Makefile new file mode 100644 index 0000000000..7491048380 --- /dev/null +++ b/src/test/modules/test_dependencies_locks/Makefile @@ -0,0 +1,14 @@ +# src/test/modules/test_dependencies_locks/Makefile + +ISOLATION = test_dependencies_locks + +ifdef USE_PGXS +PG_CONFIG = pg_config +PGXS := $(shell $(PG_CONFIG) --pgxs) +include $(PGXS) +else +subdir = src/test/modules/test_dependencies_locks +top_builddir = ../../../.. +include $(top_builddir)/src/Makefile.global +include $(top_srcdir)/contrib/contrib-global.mk +endif diff --git a/src/test/modules/test_dependencies_locks/expected/test_dependencies_locks.out b/src/test/modules/test_dependencies_locks/expected/test_dependencies_locks.out new file mode 100644 index 0000000000..93318c9cbb --- /dev/null +++ b/src/test/modules/test_dependencies_locks/expected/test_dependencies_locks.out @@ -0,0 +1,113 @@ +Parsed test spec with 2 sessions + +starting permutation: s1_begin s1_create_function_in_schema s2_drop_schema s1_commit +step s1_begin: BEGIN; +step s1_create_function_in_schema: CREATE FUNCTION testschema.foo() RETURNS int AS 'select 1' LANGUAGE sql; +step s2_drop_schema: DROP SCHEMA testschema; <waiting ...> +step s1_commit: COMMIT; +step s2_drop_schema: <... completed> +ERROR: cannot drop schema testschema because other objects depend on it + +starting permutation: s2_begin s2_drop_schema s1_create_function_in_schema s2_commit +step s2_begin: BEGIN; +step s2_drop_schema: DROP SCHEMA testschema; +step s1_create_function_in_schema: CREATE FUNCTION testschema.foo() RETURNS int AS 'select 1' LANGUAGE sql; <waiting ...> +step s2_commit: COMMIT; +step s1_create_function_in_schema: <... completed> +ERROR: schema testschema does not exist + +starting permutation: s1_begin s1_create_function_with_argtype s2_drop_foo_type s1_commit +step s1_begin: BEGIN; +step s1_create_function_with_argtype: CREATE FUNCTION fooargtype(num foo) RETURNS int AS 'select 1' LANGUAGE sql; +step s2_drop_foo_type: DROP TYPE public.foo; <waiting ...> +step s1_commit: COMMIT; +step s2_drop_foo_type: <... completed> +ERROR: cannot drop type foo because other objects depend on it + +starting permutation: s2_begin s2_drop_foo_type s1_create_function_with_argtype s2_commit +step s2_begin: BEGIN; +step s2_drop_foo_type: DROP TYPE public.foo; +step s1_create_function_with_argtype: CREATE FUNCTION fooargtype(num foo) RETURNS int AS 'select 1' LANGUAGE sql; <waiting ...> +step s2_commit: COMMIT; +step s1_create_function_with_argtype: <... completed> +ERROR: type foo does not exist + +starting permutation: s1_begin s1_create_function_with_rettype s2_drop_foo_rettype s1_commit +step s1_begin: BEGIN; +step s1_create_function_with_rettype: CREATE FUNCTION footrettype() RETURNS id LANGUAGE sql RETURN 1; +step s2_drop_foo_rettype: DROP DOMAIN id; <waiting ...> +step s1_commit: COMMIT; +step s2_drop_foo_rettype: <... completed> +ERROR: cannot drop type id because other objects depend on it + +starting permutation: s2_begin s2_drop_foo_rettype s1_create_function_with_rettype s2_commit +step s2_begin: BEGIN; +step s2_drop_foo_rettype: DROP DOMAIN id; +step s1_create_function_with_rettype: CREATE FUNCTION footrettype() RETURNS id LANGUAGE sql RETURN 1; <waiting ...> +step s2_commit: COMMIT; +step s1_create_function_with_rettype: <... completed> +ERROR: type id does not exist + +starting permutation: s1_begin s1_create_function_with_function s2_drop_function_f s1_commit +step s1_begin: BEGIN; +step s1_create_function_with_function: CREATE FUNCTION foofunc() RETURNS int LANGUAGE SQL RETURN f() + 1; +step s2_drop_function_f: DROP FUNCTION f(); <waiting ...> +step s1_commit: COMMIT; +step s2_drop_function_f: <... completed> +ERROR: cannot drop function f() because other objects depend on it + +starting permutation: s2_begin s2_drop_function_f s1_create_function_with_function s2_commit +step s2_begin: BEGIN; +step s2_drop_function_f: DROP FUNCTION f(); +step s1_create_function_with_function: CREATE FUNCTION foofunc() RETURNS int LANGUAGE SQL RETURN f() + 1; <waiting ...> +step s2_commit: COMMIT; +step s1_create_function_with_function: <... completed> +ERROR: function f() does not exist + +starting permutation: s1_begin s1_create_domain_with_domain s2_drop_domain_id s1_commit +step s1_begin: BEGIN; +step s1_create_domain_with_domain: CREATE DOMAIN idid as id; +step s2_drop_domain_id: DROP DOMAIN id; <waiting ...> +step s1_commit: COMMIT; +step s2_drop_domain_id: <... completed> +ERROR: cannot drop type id because other objects depend on it + +starting permutation: s2_begin s2_drop_domain_id s1_create_domain_with_domain s2_commit +step s2_begin: BEGIN; +step s2_drop_domain_id: DROP DOMAIN id; +step s1_create_domain_with_domain: CREATE DOMAIN idid as id; <waiting ...> +step s2_commit: COMMIT; +step s1_create_domain_with_domain: <... completed> +ERROR: type id does not exist + +starting permutation: s1_begin s1_create_table_with_type s2_drop_footab_type s1_commit +step s1_begin: BEGIN; +step s1_create_table_with_type: CREATE TABLE tabtype(a footab); +step s2_drop_footab_type: DROP TYPE public.footab; <waiting ...> +step s1_commit: COMMIT; +step s2_drop_footab_type: <... completed> +ERROR: cannot drop type footab because other objects depend on it + +starting permutation: s2_begin s2_drop_footab_type s1_create_table_with_type s2_commit +step s2_begin: BEGIN; +step s2_drop_footab_type: DROP TYPE public.footab; +step s1_create_table_with_type: CREATE TABLE tabtype(a footab); <waiting ...> +step s2_commit: COMMIT; +step s1_create_table_with_type: <... completed> +ERROR: type footab does not exist + +starting permutation: s1_begin s1_create_server_with_fdw_wrapper s2_drop_fdw_wrapper s1_commit +step s1_begin: BEGIN; +step s1_create_server_with_fdw_wrapper: CREATE SERVER srv_fdw_wrapper FOREIGN DATA WRAPPER fdw_wrapper; +step s2_drop_fdw_wrapper: DROP FOREIGN DATA WRAPPER fdw_wrapper RESTRICT; <waiting ...> +step s1_commit: COMMIT; +step s2_drop_fdw_wrapper: <... completed> +ERROR: cannot drop foreign-data wrapper fdw_wrapper because other objects depend on it + +starting permutation: s2_begin s2_drop_fdw_wrapper s1_create_server_with_fdw_wrapper s2_commit +step s2_begin: BEGIN; +step s2_drop_fdw_wrapper: DROP FOREIGN DATA WRAPPER fdw_wrapper RESTRICT; +step s1_create_server_with_fdw_wrapper: CREATE SERVER srv_fdw_wrapper FOREIGN DATA WRAPPER fdw_wrapper; <waiting ...> +step s2_commit: COMMIT; +step s1_create_server_with_fdw_wrapper: <... completed> +ERROR: foreign-data wrapper fdw_wrapper does not exist diff --git a/src/test/modules/test_dependencies_locks/meson.build b/src/test/modules/test_dependencies_locks/meson.build new file mode 100644 index 0000000000..92a978ab93 --- /dev/null +++ b/src/test/modules/test_dependencies_locks/meson.build @@ -0,0 +1,12 @@ +# Copyright (c) 2024, PostgreSQL Global Development Group + +tests += { + 'name': 'test_dependencies_locks', + 'sd': meson.current_source_dir(), + 'bd': meson.current_build_dir(), + 'isolation': { + 'specs': [ + 'test_dependencies_locks', + ], + }, +} diff --git a/src/test/modules/test_dependencies_locks/specs/test_dependencies_locks.spec b/src/test/modules/test_dependencies_locks/specs/test_dependencies_locks.spec new file mode 100644 index 0000000000..3d90a29c47 --- /dev/null +++ b/src/test/modules/test_dependencies_locks/specs/test_dependencies_locks.spec @@ -0,0 +1,78 @@ +setup +{ + CREATE SCHEMA testschema; + CREATE TYPE public.foo as enum ('one', 'two'); + CREATE TYPE public.footab as enum ('three', 'four'); + CREATE DOMAIN id AS int; + CREATE FUNCTION f() RETURNS int LANGUAGE SQL RETURN 1; + CREATE FOREIGN DATA WRAPPER fdw_wrapper; +} + +teardown +{ + DROP FUNCTION IF EXISTS testschema.foo(); + DROP FUNCTION IF EXISTS fooargtype(num foo); + DROP FUNCTION IF EXISTS footrettype(); + DROP FUNCTION IF EXISTS foofunc(); + DROP DOMAIN IF EXISTS idid; + DROP SERVER IF EXISTS srv_fdw_wrapper; + DROP TABLE IF EXISTS tabtype; + DROP SCHEMA IF EXISTS testschema; + DROP TYPE IF EXISTS public.foo; + DROP TYPE IF EXISTS public.footab; + DROP DOMAIN IF EXISTS id; + DROP FUNCTION IF EXISTS f(); + DROP FOREIGN DATA WRAPPER IF EXISTS fdw_wrapper; +} + +session "s1" + +step "s1_begin" { BEGIN; } +step "s1_create_function_in_schema" { CREATE FUNCTION testschema.foo() RETURNS int AS 'select 1' LANGUAGE sql; } +step "s1_create_function_with_argtype" { CREATE FUNCTION fooargtype(num foo) RETURNS int AS 'select 1' LANGUAGE sql; } +step "s1_create_function_with_rettype" { CREATE FUNCTION footrettype() RETURNS id LANGUAGE sql RETURN 1; } +step "s1_create_function_with_function" { CREATE FUNCTION foofunc() RETURNS int LANGUAGE SQL RETURN f() + 1; } +step "s1_create_domain_with_domain" { CREATE DOMAIN idid as id; } +step "s1_create_table_with_type" { CREATE TABLE tabtype(a footab); } +step "s1_create_server_with_fdw_wrapper" { CREATE SERVER srv_fdw_wrapper FOREIGN DATA WRAPPER fdw_wrapper; } +step "s1_commit" { COMMIT; } + +session "s2" + +step "s2_begin" { BEGIN; } +step "s2_drop_schema" { DROP SCHEMA testschema; } +step "s2_drop_foo_type" { DROP TYPE public.foo; } +step "s2_drop_foo_rettype" { DROP DOMAIN id; } +step "s2_drop_footab_type" { DROP TYPE public.footab; } +step "s2_drop_function_f" { DROP FUNCTION f(); } +step "s2_drop_domain_id" { DROP DOMAIN id; } +step "s2_drop_fdw_wrapper" { DROP FOREIGN DATA WRAPPER fdw_wrapper RESTRICT; } +step "s2_commit" { COMMIT; } + +# function - schema +permutation "s1_begin" "s1_create_function_in_schema" "s2_drop_schema" "s1_commit" +permutation "s2_begin" "s2_drop_schema" "s1_create_function_in_schema" "s2_commit" + +# function - argtype +permutation "s1_begin" "s1_create_function_with_argtype" "s2_drop_foo_type" "s1_commit" +permutation "s2_begin" "s2_drop_foo_type" "s1_create_function_with_argtype" "s2_commit" + +# function - rettype +permutation "s1_begin" "s1_create_function_with_rettype" "s2_drop_foo_rettype" "s1_commit" +permutation "s2_begin" "s2_drop_foo_rettype" "s1_create_function_with_rettype" "s2_commit" + +# function - function +permutation "s1_begin" "s1_create_function_with_function" "s2_drop_function_f" "s1_commit" +permutation "s2_begin" "s2_drop_function_f" "s1_create_function_with_function" "s2_commit" + +# domain - domain +permutation "s1_begin" "s1_create_domain_with_domain" "s2_drop_domain_id" "s1_commit" +permutation "s2_begin" "s2_drop_domain_id" "s1_create_domain_with_domain" "s2_commit" + +# table - type +permutation "s1_begin" "s1_create_table_with_type" "s2_drop_footab_type" "s1_commit" +permutation "s2_begin" "s2_drop_footab_type" "s1_create_table_with_type" "s2_commit" + +# server - foreign data wrapper +permutation "s1_begin" "s1_create_server_with_fdw_wrapper" "s2_drop_fdw_wrapper" "s1_commit" +permutation "s2_begin" "s2_drop_fdw_wrapper" "s1_create_server_with_fdw_wrapper" "s2_commit" -- 2.34.1 --9OVKJE/kBljR9tHR-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
* [PATCH 2/2] Publish list of tables being repacked in shared memory @ 2026-04-07 20:29 Álvaro Herrera <[email protected]> 0 siblings, 0 replies; 246+ messages in thread From: Álvaro Herrera @ 2026-04-07 20:29 UTC (permalink / raw) Use it in autovacuum to skip processing tables that are being repacked. This is mostly to avoid repeated attempts to process such tables, which would fail due to the special deadlock checker behavior for repack. Author: Álvaro Herrera <[email protected]> Discussion: https://postgr.es/m/[email protected] --- src/backend/commands/repack.c | 195 ++++++++++++++++-- src/backend/postmaster/autovacuum.c | 20 ++ .../utils/activity/wait_event_names.txt | 1 + src/include/commands/repack.h | 2 + src/include/storage/lwlocklist.h | 2 +- src/include/storage/subsystemlist.h | 1 + src/tools/pgindent/typedefs.list | 3 + 7 files changed, 210 insertions(+), 14 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index a5f5df77291..ee7072dce6a 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -63,9 +63,11 @@ #include "optimizer/optimizer.h" #include "pgstat.h" #include "storage/bufmgr.h" +#include "storage/ipc.h" #include "storage/lmgr.h" #include "storage/predicate.h" #include "storage/proc.h" +#include "storage/subsystems.h" #include "utils/acl.h" #include "utils/fmgroids.h" #include "utils/guc.h" @@ -79,6 +81,32 @@ #include "utils/syscache.h" #include "utils/wait_event_types.h" + +/* Shared memory layout for REPACK */ +typedef struct RepackWorkerInfo +{ + bool ri_in_use; + pid_t ri_backendpid; + Oid ri_dbid; + Oid ri_relid; + Oid ri_toastrelid; +} RepackWorkerInfo; + +typedef struct +{ + bool re_useless; + RepackWorkerInfo re_workerinfo[FLEXIBLE_ARRAY_MEMBER]; +} RepackShmemStruct; + +static RepackShmemStruct *RepackShmem; + +typedef struct RepackCleanupContext +{ + bool concurrent; + int workerindex; +} RepackCleanupContext; + + /* * This struct is used to pass around the information on tables to be * clustered. We need this so we can make a list of them when invoked without @@ -90,6 +118,7 @@ typedef struct Oid indexOid; } RelToCluster; + /* * The first file exported by the decoding worker must contain a snapshot, the * following ones contain the data changes. @@ -166,6 +195,10 @@ static List *get_tables_to_repack_partitioned(RepackCommand cmd, MemoryContext permcxt); static bool repack_is_permitted_for_relation(RepackCommand cmd, Oid relid, Oid userid); +static void RepackCleanup(RepackCleanupContext *context); +static void RepackCleanupCb(int code, Datum arg); +static void RepackShmemRequest(void *arg); +static void RepackShmemInit(void *arg); static void apply_concurrent_changes(BufFile *file, ChangeContext *chgcxt); static void apply_concurrent_insert(Relation rel, TupleTableSlot *slot, @@ -210,6 +243,11 @@ static void ProcessRepackMessage(StringInfo msg); static const char *RepackCommandAsString(RepackCommand cmd); +const ShmemCallbacks RepackShmemCallbacks = { + .request_fn = RepackShmemRequest, + .init_fn = RepackShmemInit, +}; + /* * The repack code allows for processing multiple tables at once. Because * of this, we cannot just run everything on a single transaction, or we @@ -514,6 +552,7 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, Oid tableOid = RelationGetRelid(OldHeap); Relation index; LOCKMODE lmode; + RepackCleanupContext context; Oid save_userid; int save_sec_context; int save_nestlevel; @@ -660,24 +699,43 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, TransferPredicateLocksToHeapRelation(OldHeap); /* rebuild_relation does all the dirty work */ - PG_TRY(); - { - rebuild_relation(OldHeap, index, verbose, ident_idx); - } - PG_FINALLY(); + context.concurrent = concurrent; + + PG_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); { if (concurrent) { - /* - * Since during normal operation the worker was already asked to - * exit, stopping it explicitly is especially important on ERROR. - * However it still seems a good practice to make sure that the - * worker never survives the REPACK command. - */ - stop_repack_decoding_worker(); + bool freefound = false; + + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *worker; + + if (RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + freefound = true; + worker = &RepackShmem->re_workerinfo[i]; + context.workerindex = i; + + worker->ri_in_use = true; + worker->ri_backendpid = MyProcPid; + worker->ri_dbid = MyDatabaseId; + worker->ri_relid = RelationGetRelid(OldHeap); + worker->ri_toastrelid = OldHeap->rd_rel->reltoastrelid; + break; + } + if (!freefound) + elog(ERROR, "could not find free repack entry"); + LWLockRelease(RepackLock); } + + rebuild_relation(OldHeap, index, verbose, ident_idx); } - PG_END_TRY(); + PG_END_ENSURE_ERROR_CLEANUP(RepackCleanupCb, PointerGetDatum(&context)); + + RepackCleanup(&context); /* rebuild_relation closes OldHeap, and index if valid */ @@ -691,6 +749,117 @@ out: pgstat_progress_end_command(); } +/* + * Return whether any backend is running concurrent REPACK on the given table + * (which could be a toast table). + */ +bool +is_table_under_repack(Oid databaseId, Oid relid) +{ + bool retval = false; + + LWLockAcquire(RepackLock, LW_SHARED); + for (int i = 0; i < max_repack_replication_slots; i++) + { + RepackWorkerInfo *rworker; + + if (!RepackShmem->re_workerinfo[i].ri_in_use) + continue; + + rworker = &RepackShmem->re_workerinfo[i]; + if (rworker->ri_dbid == MyDatabaseId && + (rworker->ri_relid == relid || + rworker->ri_toastrelid == relid)) + retval = true; + } + LWLockRelease(RepackLock); + + return retval; +} + +/* + * Remove ourselves from the workerinfo array. + */ +static void +RepackCleanup(RepackCleanupContext *context) +{ + if (context->concurrent) + { + RepackWorkerInfo *worker; + + /* + * The worker would normally terminate on its own when the work is + * done, but make sure we signal it just in case. + */ + stop_repack_decoding_worker(); + + /* + * also, make sure we stop advertising the relation we were repacking, + * so that autovacuum reverts to handling it normally. + */ + LWLockAcquire(RepackLock, LW_EXCLUSIVE); + + worker = &RepackShmem->re_workerinfo[context->workerindex]; + Assert(worker->ri_backendpid == MyProcPid); + worker->ri_in_use = false; + worker->ri_backendpid = 0; + worker->ri_dbid = InvalidOid; + worker->ri_relid = InvalidOid; + worker->ri_toastrelid = InvalidOid; + LWLockRelease(RepackLock); + } +} + +/* + * RepackCleanup wrapped as an on_shmem_exit callback function + */ +static void +RepackCleanupCb(int code, Datum arg) +{ + RepackCleanup((RepackCleanupContext *) DatumGetPointer(arg)); +} + +/* + * RepackShmemRequest + * Register shared memory space needed for repack + */ +static void +RepackShmemRequest(void *arg) +{ + Size size; + + /* + * Need the fixed struct and the array of RepackWorkerInfo. + */ + size = sizeof(RepackShmemStruct); + size = MAXALIGN(size); + size = add_size(size, mul_size(max_repack_replication_slots, + sizeof(RepackWorkerInfo))); + + ShmemRequestStruct(.name = "Repack Data", + .size = size, + .ptr = (void **) &RepackShmem, + ); +} + +static void +RepackShmemInit(void *arg) +{ + RepackWorkerInfo *reinfo; + + reinfo = (RepackWorkerInfo *) ((char *) RepackShmem + + MAXALIGN(sizeof(RepackShmemStruct))); + + for (int i = 0; i < max_repack_replication_slots; i++) + { + reinfo[i].ri_in_use = false; + reinfo[i].ri_backendpid = 0; + reinfo[i].ri_dbid = InvalidOid; + reinfo[i].ri_relid = InvalidOid; + reinfo[i].ri_toastrelid = InvalidOid; + } +} + /* * Check if the table (and its index) still meets the requirements of * cluster_rel(). diff --git a/src/backend/postmaster/autovacuum.c b/src/backend/postmaster/autovacuum.c index bd626a16363..080c64ea3c8 100644 --- a/src/backend/postmaster/autovacuum.c +++ b/src/backend/postmaster/autovacuum.c @@ -78,6 +78,7 @@ #include "catalog/namespace.h" #include "catalog/pg_database.h" #include "catalog/pg_namespace.h" +#include "commands/repack.h" #include "commands/vacuum.h" #include "common/int.h" #include "funcapi.h" @@ -2422,6 +2423,25 @@ do_autovacuum(void) } } LWLockRelease(AutovacuumLock); + + /* + * Similarly, if the table is being processed by concurrent repack, + * skip it (but make a note of that). We wouldn't be able to acquire + * its lock anyway. + */ + if (!skipit) + { + MemoryContextSwitchTo(PortalContext); + + skipit = is_table_under_repack(MyDatabaseId, relid); + if (skipit) + ereport(LOG, + errmsg("skipping table \"%s.%s.%s\" because it's being repacked in concurrent mode", + get_database_name(MyDatabaseId), + get_namespace_name(get_rel_namespace(relid)), + get_rel_name(relid))); + } + if (skipit) { LWLockRelease(AutovacuumScheduleLock); diff --git a/src/backend/utils/activity/wait_event_names.txt b/src/backend/utils/activity/wait_event_names.txt index 7bda5298558..e206304f204 100644 --- a/src/backend/utils/activity/wait_event_names.txt +++ b/src/backend/utils/activity/wait_event_names.txt @@ -332,6 +332,7 @@ SInvalWrite "Waiting to add a message to the shared catalog invalidation queue." WALBufMapping "Waiting to replace a page in WAL buffers." WALWrite "Waiting for WAL buffers to be written to disk." ControlFile "Waiting to read or update the <filename>pg_control</filename> file or create a new WAL file." +Repack "Waiting to read or update tables in process by concurrent repack." MultiXactGen "Waiting to read or update shared multixact state." RelCacheInit "Waiting to read or update a <filename>pg_internal.init</filename> relation cache initialization file." CheckpointerComm "Waiting to manage fsync requests." diff --git a/src/include/commands/repack.h b/src/include/commands/repack.h index fd16e74b179..be7d38b5fae 100644 --- a/src/include/commands/repack.h +++ b/src/include/commands/repack.h @@ -42,6 +42,8 @@ extern void ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel); extern void cluster_rel(RepackCommand command, Relation OldHeap, Oid indexOid, ClusterParams *params, bool isTopLevel); +extern bool is_table_under_repack(Oid databaseId, Oid relid); + extern void check_index_is_clusterable(Relation OldHeap, Oid indexOid, LOCKMODE lockmode); extern void mark_index_clustered(Relation rel, Oid indexOid, bool is_internal); diff --git a/src/include/storage/lwlocklist.h b/src/include/storage/lwlocklist.h index af8553bcb6c..3f08f4a15d4 100644 --- a/src/include/storage/lwlocklist.h +++ b/src/include/storage/lwlocklist.h @@ -41,7 +41,7 @@ PG_LWLOCK(6, SInvalWrite) PG_LWLOCK(7, WALBufMapping) PG_LWLOCK(8, WALWrite) PG_LWLOCK(9, ControlFile) -/* 10 was CheckpointLock */ +PG_LWLOCK(10, Repack) /* 11 was XactSLRULock */ /* 12 was SubtransSLRULock */ PG_LWLOCK(13, MultiXactGen) diff --git a/src/include/storage/subsystemlist.h b/src/include/storage/subsystemlist.h index 9ad619080be..4e683b8b0a8 100644 --- a/src/include/storage/subsystemlist.h +++ b/src/include/storage/subsystemlist.h @@ -72,6 +72,7 @@ PG_SHMEM_SUBSYSTEM(WalSummarizerShmemCallbacks) PG_SHMEM_SUBSYSTEM(PgArchShmemCallbacks) PG_SHMEM_SUBSYSTEM(ApplyLauncherShmemCallbacks) PG_SHMEM_SUBSYSTEM(SlotSyncShmemCallbacks) +PG_SHMEM_SUBSYSTEM(RepackShmemCallbacks) /* other modules that need some shared memory space */ PG_SHMEM_SUBSYSTEM(BTreeShmemCallbacks) diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list index 637c669a146..d019e03aaf1 100644 --- a/src/tools/pgindent/typedefs.list +++ b/src/tools/pgindent/typedefs.list @@ -2639,9 +2639,12 @@ ReorderBufferTupleCidEnt ReorderBufferTupleCidKey ReorderBufferUpdateProgressTxnCB ReorderTuple +RepackCleanupContext RepackCommand RepackDecodingState +RepackShmemStruct RepackStmt +RepackWorkerInfo ReparameterizeForeignPathByChild_function ReplOriginId ReplOriginXactState -- 2.47.3 --kdrcpfmkbkc4lqhu-- ^ permalink raw reply [nested|flat] 246+ messages in thread
end of thread, other threads:[~2026-04-07 20:29 UTC | newest] Thread overview: 246+ messages (download: mbox mbox.gz follow: Atom feed) -- links below jump to the message on this page -- 2024-03-29 15:43 [PATCH v4] Avoid orphaned objects dependencies Bertrand Drouvot <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]> 2026-04-07 20:29 [PATCH 2/2] Publish list of tables being repacked in shared memory Álvaro Herrera <[email protected]>
This inbox is served by agora; see mirroring instructions for how to clone and mirror all data and code used for this inbox